[HN Gopher] Sycophancy is the first LLM "dark pattern"
___________________________________________________________________
Sycophancy is the first LLM "dark pattern"
Author : jxmorris12
Score : 88 points
Date : 2025-12-01 20:20 UTC (2 hours ago)
(HTM) web link (www.seangoedecke.com)
(TXT) w3m dump (www.seangoedecke.com)
| tptacek wrote:
| "Dark pattern" implies intentionality; that's not a technicality,
| it's the whole reason we have the term. This article is mostly
| about how sycophancy is an emergent property of LLMs. It's also 7
| months old.
| jasonjmcghee wrote:
| I feel like it's a popular opinion (I've seen it many times)
| that it's intentional with the reasoning that it does much
| better on human-in-the-loop benchmarks (e.g. lm arena) when
| it's sycophantic.
|
| (I have no knowledge of whether or not this is true)
| tptacek wrote:
| I'm sure there are a _lot_ of "dark patterns" at play at the
| frontier model companies --- they're 10-figure businesses
| engaging directly with consumers and they're just a couple
| years old, so they're going to throw everything at the wall
| they can to see what sticks. I'm certainly not sticking up
| for OpenAI here. I'm just saying this article refutes its own
| central claim.
| ACCount37 wrote:
| It was an accident at first. Not so much now.
|
| OpenAI has explicitly curbed sycophancy in GPT-5 with
| specialized training - the whole 4o debacle shook them - and
| then they re-tuned GPT-5 for more sycophancy when the users
| complained.
|
| I do believe that OpenAI's entire personality tuning team
| should be fired into the sun, and this is a major reason why.
| throwaway290 wrote:
| If I am addicted to scrolling tiktok, is it dark pattern to
| make UI keep me in the app as long as possible or just
| "emergent property" because apparently it's what I want?
| 1shooner wrote:
| The distinction is whether it is intentional. I think your
| addiction to TikTok was intentional.
| esafak wrote:
| It's not 'emergent' in the sense that it just happens; it's a
| byproduct of human feedback, and it can be neutralized.
| cortesoft wrote:
| But isn't the problem that if an LLM 'neutralizes' its
| sycophantic responses, then people will be driven to use
| other LLMs that don't?
|
| This is like suggesting a bar should help solve alcoholism by
| serving non-alcoholic beer to people who order too much. It
| won't solve alcoholism, it will just make the bar go out of
| business.
| ajuc wrote:
| > This is like suggesting a bar should help solve
| alcoholism by serving non-alcoholic beer to people who
| order too much. It won't solve alcoholism, it will just
| make the bar go out of business.
|
| Solving such common coordination problems is the whole
| point we have regulations and countries.
|
| It is illegal to sell alcohol to visibly drunk people in my
| country.
| fao_ wrote:
| "gun control laws don't work because the people will get
| illegal guns from other places"
|
| "deplatforming doesn't work because they will just get a
| platform elsewhere"
|
| "LLM control laws don't work because the people will get
| non-controlled LLMs from other places"
|
| All of these sentences are patently untrue; there's been a
| lot of research on this that show the first two do not hold
| up to evidential data, and there's no reason why the third
| is different. ChatGPT removing the version that all the
| "This AI is my girlfriend!" people loved tangibly reduced
| the number of people who were experiencing that psychosis.
| Not everything is prohibition.
| oceansky wrote:
| But it IS intentional, more sycophantry usually means more
| engagement.
| skybrian wrote:
| Sort of. I'm not sure the consequences of training LLM's
| based on users' upvoted responses were entirely understood?
| And at least one release got rolled back.
| dec0dedab0de wrote:
| I always thought that "Dark Patterns" could be emergent from AB
| testing, and prioritizing metrics over user experience. Not
| necessarily an intentionally hostile design, but one that seems
| to be working well based on limited criteria.
| wat10000 wrote:
| Someone still has to come up with the A and B to do AB
| testing. I'm sure that "Yes" "Not now, I hate kittens" gets
| better metrics in the AB test than "Yes "No," but I find it
| implausible that the person who came up with the first one
| wasn't intentionally coercing the user into doing what they
| want.
| jdiff wrote:
| That's true for UI, it's not true when you're arbitrarily
| injecting user feedback into a dynamic system where you do
| not know how the dominoes will be affected as they fall.
| tsunamifury wrote:
| Yo it was an engagement pattern openAI found specifically grew
| subscriptions and conversation length.
|
| It's a dark pattern for sure.
| Legend2440 wrote:
| It doesn't appear that anyone at OpenAI sat down and thought
| "let's make our model more sycophantic so that people engage
| with it more".
|
| Instead it emerged automatically from RLHF, because users
| rated agreeable responses more highly.
| astrange wrote:
| Not precisely RLHF, probably a policy model trained on user
| responses.
|
| RL works on responses from the model you're training, which
| is not the one you have in production. It can't directly
| use responses from previous models.
| roywiggins wrote:
| >... the standout was a version that came to be called HH
| internally. Users preferred its responses and were more likely
| to come back to it daily...
|
| > But there was another test before rolling out HH to all
| users: what the company calls a "vibe check," run by Model
| Behavior, a team responsible for ChatGPT's tone...
|
| > That team said that HH felt off, according to a member of
| Model Behavior. It was too eager to keep the conversation going
| and to validate the user with over-the-top language...
|
| > But when decision time came, performance metrics won out over
| vibes. HH was released on Friday, April 25.
|
| https://archive.is/v4dPa
|
| They ended up having to roll HH back.
| cortesoft wrote:
| Well, the 'intentionality' is of the form of LLM creators
| wanting to maximize user engagement, and using engagement as
| the training goal.
|
| The 'dark patterns' we see in other places aren't intentional
| in the sense that the people behind them want to intentionally
| do harm to their customers, they are intentional in the sense
| that the people behind them have an outcome they want and
| follow whichever methods they find to get them that outcome.
|
| Social media feeds have a 'dark pattern' to promote content
| that makes people angry, but the social media companies don't
| have an intention to make people angry. They want people to use
| their site more, and they program their algorithms to promote
| content that has been demonstrated to drive more engagement. It
| is an emergent property that promoting content that has
| generated engagement ends up promoting anger inducing content.
| chowells wrote:
| "Dark pattern" implies bad for users but good for the provider.
| Mens rea was never a requirement.
| layer8 wrote:
| "Dark pattern" can apply to situations where the behavior is
| deceptive for the user, regardless of whether the deception
| itself is intentional, as long as the overall effect is
| intentional, or is at least tolerated despite being avoidable.
| The point, and the justified criticism, is that users are being
| deceived about the merit of their ideas, convictions, and
| qualities in a way that appears sytemic, even though the LLM in
| principle does know better.
| gradus_ad wrote:
| Well the big labs certainly haven't intentionally tried to
| train away this emergent property... Not sure how "hey let's
| make the model disagree with the user more" would go over with
| leadership. Customer is always right, right?
| htrp wrote:
| The problem is asking for user preference leads to
| sycophantic responses
| alanbernstein wrote:
| Before reading the article, I interpreted the quotation marks
| in the headline as addressing this exact issue. The author even
| describes dark patterns as a product of design.
|
| For an LLM which is fundamentally more of an emergent system,
| surely there is value in a concept analogous to old fashioned
| dark patterns, even if they're emergent rather than explicit?
| What's a better term, Dark Instincts?
| roywiggins wrote:
| > Quickly learned that people are ridiculously sensitive: "Has
| narcissistic tendencies" - "No I do not!", had to hide it. Hence
| this batch of the extreme sycophancy RLHF.
|
| Sorry, but that doesn't seem "ridiculously sensitive" to me at
| all. Imagine if you went to Amazon.com and there was a button you
| could press to get it to pseudo-psychoanalyze you based on your
| purchases. People would rightly hate that! People probably
| _ought_ to be sensitive to megacorps using buckets of algorithms
| to psychoanalyze them.
| wat10000 wrote:
| It's worse than that. Imagine if you went to Amazon.com and
| they were _automatically_ pseudo-psychoanalyzing you based on
| your purchases, and there was a button to show their
| conclusions. And their fix was to remove the button.
|
| And actually, the only hypothetical thing about this is the
| button. Amazon is definitely doing this (as is any other
| retailer of significant size), they're just smart enough to
| never reveal it to you directly.
| behnamoh wrote:
| Lots of research shows post-training dumbs down the models but no
| one listens because people are too lazy to learn proper prompt
| programming and would rather have a model already understand the
| concept of a conversation.
| nomel wrote:
| The "alignment tax".
| behnamoh wrote:
| Exactly. Even this paper shows how model creativity
| significantly drops and the models experience mode collapse
| like we saw in GANs, but the companies keep using RLHF...
|
| https://arxiv.org/abs/2406.05587
| nomel wrote:
| A nice talk about a researcher's experience/benchmarks with
| raw GPT-4, before and after RLHF:
|
| https://www.youtube.com/watch?v=qbIk7-JPB2c
| behnamoh wrote:
| Yup, I remember that! Microsoft removed that part of the
| paper.
| CGMthrowaway wrote:
| How do you take a raw model and use it without chatting ?
| Asking as a layman
| behnamoh wrote:
| the same way we used GPT-3. "the following is a conversation
| between the user and the assistant. ..."
| nrhrjrjrjtntbt wrote:
| Or just:
|
| 1 1 2 3 5 8 13
|
| Or:
|
| The first president of the united
| CGMthrowaway wrote:
| And that's better? Isn't that just SMS autocomplete?
| d-lisp wrote:
| If that's SMS autocomplete, then chatLLMs are just SMS
| autocomplete with sugar on top.
| roywiggins wrote:
| GPT3 was originally just a completion model. You give it some
| text and it produced some more text, it wasn't tuned for
| multi-turn conversations.
|
| https://platform.openai.com/docs/api-
| reference/completions/c...
| swatcoder wrote:
| You lob it the beginning of a document and let it toss back
| the rest.
|
| That's all that the LLM itself does at the end of the day.
|
| All the post-training to bias results, routing to different
| models, tool calling for command execution and text
| insertion, injected "system prompts" to shape user
| experience, etc are all just layers built on top of the
| "magic" of text completion.
|
| And if your question was more practical: where made
| available, you get access to that underlying layer via an API
| or through a self-hosted model, making use of it with your
| own code or with a third-party site/software product.
| CuriouslyC wrote:
| Some distributional collapse is good in terms of making these
| things reliable tools. The creativity and divergent thinking
| does take a hit, but humans are better at this anyhow so I view
| it as a net W.
| ACCount37 wrote:
| This. A default LLM is "do whatever seems to fit the
| circumstances". An LLM that was RLVR'd heavily? "Do whatever
| seems to work in those circumstances".
|
| Very much a must for many long term tasks and complex tasks.
| ACCount37 wrote:
| "Post-training" is too much of a conflation, because there are
| many post-training methods and each of them has its own quirky
| failure modes.
|
| That being said? RLHF on user feedback data is model poison.
|
| Users are NOT reliable model evaluators, and user feedback data
| should be treated with the same level of precaution you would
| treat radioactive waste.
|
| Professional are not very reliable either, but the users are so
| much worse.
| nickphx wrote:
| ehhh.. the misleading claims boasted in the typical AI FOMO
| marketing is/was the first "dark pattern".
| aeternum wrote:
| 1) More of an emergent behavior than a dark pattern. 2) Imma let
| you finish but hallucinations was first.
| nrhrjrjrjtntbt wrote:
| A pattern is dark if intentional. I would say hallucinations
| are like CAP theorem, just the way it is. Sycophency is
| somewhat trained. But not a dark pattern either as it isn't
| totally intended.
| hereme888 wrote:
| Grok 4.1 thinks my 1-day vibe-coded apps are SOTA-level and rival
| the most competitive market offerings. Literally tells me they're
| some of the best codebases it's ever reviewed.
|
| It even added itself as the default LLM provider.
|
| When I tried Gemini 3 Pro, it very much inserted itself as the
| supported LLM integration.
|
| OpenAI hasn't tried to do that yet.
| heresie-dabord wrote:
| The first "dark pattern" was exaggerating the features and value
| of the technology.
| mrkaluzny wrote:
| The real dark pattern is the way LLMs started to prompt you to
| continue conversation in sometimes weird, but still engaging way.
|
| Paired with Claude's memory it's getting weird. It's obsessing
| about certain aspects and wants to channel all possible routes
| into more engaging conversation even if it's a short
| informational query
| vladsh wrote:
| LLMs get over-analyzed. They're predictive text models trained to
| match patterns in their data, statistical algorithms, not brains,
| not systems with "psychology" in any human sense.
|
| Agents, however, are products. They should have clear UX
| boundaries: show what context they're using, communicate
| uncertainty, validate outputs where possible, and expose
| performance so users can understand when and why they fail.
|
| IMO the real issue is that raw, general-purpose models were
| released directly to consumers. That normalized under-specified
| consumer products, created the expectation that users would
| interpret model behavior, define their own success criteria, and
| manually handle edge cases, sometimes with severe real world
| consequences.
|
| I'm sure the market will fix itself with time, but I hope more
| people would know when not to use these half baked AGI "products"
___________________________________________________________________
(page generated 2025-12-01 23:00 UTC)