[HN Gopher] Sycophancy is the first LLM "dark pattern"
       ___________________________________________________________________
        
       Sycophancy is the first LLM "dark pattern"
        
       Author : jxmorris12
       Score  : 88 points
       Date   : 2025-12-01 20:20 UTC (2 hours ago)
        
 (HTM) web link (www.seangoedecke.com)
 (TXT) w3m dump (www.seangoedecke.com)
        
       | tptacek wrote:
       | "Dark pattern" implies intentionality; that's not a technicality,
       | it's the whole reason we have the term. This article is mostly
       | about how sycophancy is an emergent property of LLMs. It's also 7
       | months old.
        
         | jasonjmcghee wrote:
         | I feel like it's a popular opinion (I've seen it many times)
         | that it's intentional with the reasoning that it does much
         | better on human-in-the-loop benchmarks (e.g. lm arena) when
         | it's sycophantic.
         | 
         | (I have no knowledge of whether or not this is true)
        
           | tptacek wrote:
           | I'm sure there are a _lot_ of  "dark patterns" at play at the
           | frontier model companies --- they're 10-figure businesses
           | engaging directly with consumers and they're just a couple
           | years old, so they're going to throw everything at the wall
           | they can to see what sticks. I'm certainly not sticking up
           | for OpenAI here. I'm just saying this article refutes its own
           | central claim.
        
           | ACCount37 wrote:
           | It was an accident at first. Not so much now.
           | 
           | OpenAI has explicitly curbed sycophancy in GPT-5 with
           | specialized training - the whole 4o debacle shook them - and
           | then they re-tuned GPT-5 for more sycophancy when the users
           | complained.
           | 
           | I do believe that OpenAI's entire personality tuning team
           | should be fired into the sun, and this is a major reason why.
        
         | throwaway290 wrote:
         | If I am addicted to scrolling tiktok, is it dark pattern to
         | make UI keep me in the app as long as possible or just
         | "emergent property" because apparently it's what I want?
        
           | 1shooner wrote:
           | The distinction is whether it is intentional. I think your
           | addiction to TikTok was intentional.
        
         | esafak wrote:
         | It's not 'emergent' in the sense that it just happens; it's a
         | byproduct of human feedback, and it can be neutralized.
        
           | cortesoft wrote:
           | But isn't the problem that if an LLM 'neutralizes' its
           | sycophantic responses, then people will be driven to use
           | other LLMs that don't?
           | 
           | This is like suggesting a bar should help solve alcoholism by
           | serving non-alcoholic beer to people who order too much. It
           | won't solve alcoholism, it will just make the bar go out of
           | business.
        
             | ajuc wrote:
             | > This is like suggesting a bar should help solve
             | alcoholism by serving non-alcoholic beer to people who
             | order too much. It won't solve alcoholism, it will just
             | make the bar go out of business.
             | 
             | Solving such common coordination problems is the whole
             | point we have regulations and countries.
             | 
             | It is illegal to sell alcohol to visibly drunk people in my
             | country.
        
             | fao_ wrote:
             | "gun control laws don't work because the people will get
             | illegal guns from other places"
             | 
             | "deplatforming doesn't work because they will just get a
             | platform elsewhere"
             | 
             | "LLM control laws don't work because the people will get
             | non-controlled LLMs from other places"
             | 
             | All of these sentences are patently untrue; there's been a
             | lot of research on this that show the first two do not hold
             | up to evidential data, and there's no reason why the third
             | is different. ChatGPT removing the version that all the
             | "This AI is my girlfriend!" people loved tangibly reduced
             | the number of people who were experiencing that psychosis.
             | Not everything is prohibition.
        
         | oceansky wrote:
         | But it IS intentional, more sycophantry usually means more
         | engagement.
        
           | skybrian wrote:
           | Sort of. I'm not sure the consequences of training LLM's
           | based on users' upvoted responses were entirely understood?
           | And at least one release got rolled back.
        
         | dec0dedab0de wrote:
         | I always thought that "Dark Patterns" could be emergent from AB
         | testing, and prioritizing metrics over user experience. Not
         | necessarily an intentionally hostile design, but one that seems
         | to be working well based on limited criteria.
        
           | wat10000 wrote:
           | Someone still has to come up with the A and B to do AB
           | testing. I'm sure that "Yes" "Not now, I hate kittens" gets
           | better metrics in the AB test than "Yes "No," but I find it
           | implausible that the person who came up with the first one
           | wasn't intentionally coercing the user into doing what they
           | want.
        
             | jdiff wrote:
             | That's true for UI, it's not true when you're arbitrarily
             | injecting user feedback into a dynamic system where you do
             | not know how the dominoes will be affected as they fall.
        
         | tsunamifury wrote:
         | Yo it was an engagement pattern openAI found specifically grew
         | subscriptions and conversation length.
         | 
         | It's a dark pattern for sure.
        
           | Legend2440 wrote:
           | It doesn't appear that anyone at OpenAI sat down and thought
           | "let's make our model more sycophantic so that people engage
           | with it more".
           | 
           | Instead it emerged automatically from RLHF, because users
           | rated agreeable responses more highly.
        
             | astrange wrote:
             | Not precisely RLHF, probably a policy model trained on user
             | responses.
             | 
             | RL works on responses from the model you're training, which
             | is not the one you have in production. It can't directly
             | use responses from previous models.
        
         | roywiggins wrote:
         | >... the standout was a version that came to be called HH
         | internally. Users preferred its responses and were more likely
         | to come back to it daily...
         | 
         | > But there was another test before rolling out HH to all
         | users: what the company calls a "vibe check," run by Model
         | Behavior, a team responsible for ChatGPT's tone...
         | 
         | > That team said that HH felt off, according to a member of
         | Model Behavior. It was too eager to keep the conversation going
         | and to validate the user with over-the-top language...
         | 
         | > But when decision time came, performance metrics won out over
         | vibes. HH was released on Friday, April 25.
         | 
         | https://archive.is/v4dPa
         | 
         | They ended up having to roll HH back.
        
         | cortesoft wrote:
         | Well, the 'intentionality' is of the form of LLM creators
         | wanting to maximize user engagement, and using engagement as
         | the training goal.
         | 
         | The 'dark patterns' we see in other places aren't intentional
         | in the sense that the people behind them want to intentionally
         | do harm to their customers, they are intentional in the sense
         | that the people behind them have an outcome they want and
         | follow whichever methods they find to get them that outcome.
         | 
         | Social media feeds have a 'dark pattern' to promote content
         | that makes people angry, but the social media companies don't
         | have an intention to make people angry. They want people to use
         | their site more, and they program their algorithms to promote
         | content that has been demonstrated to drive more engagement. It
         | is an emergent property that promoting content that has
         | generated engagement ends up promoting anger inducing content.
        
         | chowells wrote:
         | "Dark pattern" implies bad for users but good for the provider.
         | Mens rea was never a requirement.
        
         | layer8 wrote:
         | "Dark pattern" can apply to situations where the behavior is
         | deceptive for the user, regardless of whether the deception
         | itself is intentional, as long as the overall effect is
         | intentional, or is at least tolerated despite being avoidable.
         | The point, and the justified criticism, is that users are being
         | deceived about the merit of their ideas, convictions, and
         | qualities in a way that appears sytemic, even though the LLM in
         | principle does know better.
        
         | gradus_ad wrote:
         | Well the big labs certainly haven't intentionally tried to
         | train away this emergent property... Not sure how "hey let's
         | make the model disagree with the user more" would go over with
         | leadership. Customer is always right, right?
        
           | htrp wrote:
           | The problem is asking for user preference leads to
           | sycophantic responses
        
         | alanbernstein wrote:
         | Before reading the article, I interpreted the quotation marks
         | in the headline as addressing this exact issue. The author even
         | describes dark patterns as a product of design.
         | 
         | For an LLM which is fundamentally more of an emergent system,
         | surely there is value in a concept analogous to old fashioned
         | dark patterns, even if they're emergent rather than explicit?
         | What's a better term, Dark Instincts?
        
       | roywiggins wrote:
       | > Quickly learned that people are ridiculously sensitive: "Has
       | narcissistic tendencies" - "No I do not!", had to hide it. Hence
       | this batch of the extreme sycophancy RLHF.
       | 
       | Sorry, but that doesn't seem "ridiculously sensitive" to me at
       | all. Imagine if you went to Amazon.com and there was a button you
       | could press to get it to pseudo-psychoanalyze you based on your
       | purchases. People would rightly hate that! People probably
       | _ought_ to be sensitive to megacorps using buckets of algorithms
       | to psychoanalyze them.
        
         | wat10000 wrote:
         | It's worse than that. Imagine if you went to Amazon.com and
         | they were _automatically_ pseudo-psychoanalyzing you based on
         | your purchases, and there was a button to show their
         | conclusions. And their fix was to remove the button.
         | 
         | And actually, the only hypothetical thing about this is the
         | button. Amazon is definitely doing this (as is any other
         | retailer of significant size), they're just smart enough to
         | never reveal it to you directly.
        
       | behnamoh wrote:
       | Lots of research shows post-training dumbs down the models but no
       | one listens because people are too lazy to learn proper prompt
       | programming and would rather have a model already understand the
       | concept of a conversation.
        
         | nomel wrote:
         | The "alignment tax".
        
           | behnamoh wrote:
           | Exactly. Even this paper shows how model creativity
           | significantly drops and the models experience mode collapse
           | like we saw in GANs, but the companies keep using RLHF...
           | 
           | https://arxiv.org/abs/2406.05587
        
             | nomel wrote:
             | A nice talk about a researcher's experience/benchmarks with
             | raw GPT-4, before and after RLHF:
             | 
             | https://www.youtube.com/watch?v=qbIk7-JPB2c
        
               | behnamoh wrote:
               | Yup, I remember that! Microsoft removed that part of the
               | paper.
        
         | CGMthrowaway wrote:
         | How do you take a raw model and use it without chatting ?
         | Asking as a layman
        
           | behnamoh wrote:
           | the same way we used GPT-3. "the following is a conversation
           | between the user and the assistant. ..."
        
             | nrhrjrjrjtntbt wrote:
             | Or just:
             | 
             | 1 1 2 3 5 8 13
             | 
             | Or:
             | 
             | The first president of the united
        
               | CGMthrowaway wrote:
               | And that's better? Isn't that just SMS autocomplete?
        
               | d-lisp wrote:
               | If that's SMS autocomplete, then chatLLMs are just SMS
               | autocomplete with sugar on top.
        
           | roywiggins wrote:
           | GPT3 was originally just a completion model. You give it some
           | text and it produced some more text, it wasn't tuned for
           | multi-turn conversations.
           | 
           | https://platform.openai.com/docs/api-
           | reference/completions/c...
        
           | swatcoder wrote:
           | You lob it the beginning of a document and let it toss back
           | the rest.
           | 
           | That's all that the LLM itself does at the end of the day.
           | 
           | All the post-training to bias results, routing to different
           | models, tool calling for command execution and text
           | insertion, injected "system prompts" to shape user
           | experience, etc are all just layers built on top of the
           | "magic" of text completion.
           | 
           | And if your question was more practical: where made
           | available, you get access to that underlying layer via an API
           | or through a self-hosted model, making use of it with your
           | own code or with a third-party site/software product.
        
         | CuriouslyC wrote:
         | Some distributional collapse is good in terms of making these
         | things reliable tools. The creativity and divergent thinking
         | does take a hit, but humans are better at this anyhow so I view
         | it as a net W.
        
           | ACCount37 wrote:
           | This. A default LLM is "do whatever seems to fit the
           | circumstances". An LLM that was RLVR'd heavily? "Do whatever
           | seems to work in those circumstances".
           | 
           | Very much a must for many long term tasks and complex tasks.
        
         | ACCount37 wrote:
         | "Post-training" is too much of a conflation, because there are
         | many post-training methods and each of them has its own quirky
         | failure modes.
         | 
         | That being said? RLHF on user feedback data is model poison.
         | 
         | Users are NOT reliable model evaluators, and user feedback data
         | should be treated with the same level of precaution you would
         | treat radioactive waste.
         | 
         | Professional are not very reliable either, but the users are so
         | much worse.
        
       | nickphx wrote:
       | ehhh.. the misleading claims boasted in the typical AI FOMO
       | marketing is/was the first "dark pattern".
        
       | aeternum wrote:
       | 1) More of an emergent behavior than a dark pattern. 2) Imma let
       | you finish but hallucinations was first.
        
         | nrhrjrjrjtntbt wrote:
         | A pattern is dark if intentional. I would say hallucinations
         | are like CAP theorem, just the way it is. Sycophency is
         | somewhat trained. But not a dark pattern either as it isn't
         | totally intended.
        
       | hereme888 wrote:
       | Grok 4.1 thinks my 1-day vibe-coded apps are SOTA-level and rival
       | the most competitive market offerings. Literally tells me they're
       | some of the best codebases it's ever reviewed.
       | 
       | It even added itself as the default LLM provider.
       | 
       | When I tried Gemini 3 Pro, it very much inserted itself as the
       | supported LLM integration.
       | 
       | OpenAI hasn't tried to do that yet.
        
       | heresie-dabord wrote:
       | The first "dark pattern" was exaggerating the features and value
       | of the technology.
        
       | mrkaluzny wrote:
       | The real dark pattern is the way LLMs started to prompt you to
       | continue conversation in sometimes weird, but still engaging way.
       | 
       | Paired with Claude's memory it's getting weird. It's obsessing
       | about certain aspects and wants to channel all possible routes
       | into more engaging conversation even if it's a short
       | informational query
        
       | vladsh wrote:
       | LLMs get over-analyzed. They're predictive text models trained to
       | match patterns in their data, statistical algorithms, not brains,
       | not systems with "psychology" in any human sense.
       | 
       | Agents, however, are products. They should have clear UX
       | boundaries: show what context they're using, communicate
       | uncertainty, validate outputs where possible, and expose
       | performance so users can understand when and why they fail.
       | 
       | IMO the real issue is that raw, general-purpose models were
       | released directly to consumers. That normalized under-specified
       | consumer products, created the expectation that users would
       | interpret model behavior, define their own success criteria, and
       | manually handle edge cases, sometimes with severe real world
       | consequences.
       | 
       | I'm sure the market will fix itself with time, but I hope more
       | people would know when not to use these half baked AGI "products"
        
       ___________________________________________________________________
       (page generated 2025-12-01 23:00 UTC)