[HN Gopher] AI overly affirms users asking for personal advice
       ___________________________________________________________________
        
       AI overly affirms users asking for personal advice
        
       https://arxiv.org/abs/2602.14270
       https://www.science.org/doi/10.1126/science.aec8352
        
       Author : oldfrenchfries
       Score  : 755 points
       Date   : 2026-03-28 14:08 UTC (1 days ago)
        
 (HTM) web link (news.stanford.edu)
 (TXT) w3m dump (news.stanford.edu)
        
       | oldfrenchfries wrote:
       | This new Stanford study published on March 26, 2026 shows that AI
       | models are sycophantic. They affirm the users position 49% more
       | often than a human would.
       | 
       | The researchers found that when people use AI for relationship
       | advice, they become 25% more convinced they are 'right' and
       | significantly less likely to apologize or repair the connection.
        
         | jatins wrote:
         | To be fair an average therapist is also pretty sycophantic.
         | "The worst person you know is being told by their therapist
         | that they did the right thing" is a bit of a meme, but isn't
         | completely false in my experience.
        
           | kibwen wrote:
           | No, the meme is that the average therapist can be boiled down
           | to "well, what do you think?" or "and how does that make you
           | feel?" (of which ELIZA, the original bot that passed the
           | Turing test, was perhaps an unintentional parody). Even this
           | cartoonish characterization demonstrates that the function of
           | therapists is to get you to question yourself so that you can
           | attempt to reframe and re-evaluate your ways of thinking, in
           | a roughly Socratic fashion.
        
             | toraway wrote:
             | It was entirely intentional. The Rogerian school of
             | psychotherapy stereotyped by "how does that make you feel"
             | was popular at the time and the most popular ELIZA script
             | used that persona to cleverly redirect focus from the bot's
             | weaknesses in comprehension.
        
       | oldfrenchfries wrote:
       | There is a striking data visualization showing the breakup advice
       | trend over 15 years on Reddit. You can see the "End relationship"
       | line spike as AI and algorithmic advice take over:
       | 
       | https://www.reddit.com/r/dataisbeautiful/comments/1o87cy4/oc...
        
         | falcor84 wrote:
         | Isn't the fact that a person is asking an AI whether to leave
         | their partner in its own an indication that they should?
         | 
         | EDIT: typo
        
           | nomorewords wrote:
           | How is it an indication? I think people on here don't realize
           | that most of the people don't think things through as much as
           | (software) engineers
        
             | hnfong wrote:
             | In my local(?) community (like in my city, not my industry)
             | there is a saying "if you had to ask for relationship
             | advice, then you probably should break up".
             | 
             | There is some rationale to that. People tend to hold onto
             | relationships that don't lead anywhere in fear of "losing"
             | what they "already have". It's probably a comfort zone
             | thing. So if one is desperate enough to ask random
             | strangers online about a relationship, it's usually biased
             | towards some unresolvable issue that would have the parties
             | better of if they break up.
        
               | magicalhippo wrote:
               | > So if one is desperate enough to ask random strangers
               | online about a relationship
               | 
               | I'd me more inclined to ask random strangers on the
               | internet than close friends...
               | 
               | That said, when me and my SO had a difficult time we went
               | to a professional. For us it helped a lot. Though as the
               | counselor said, we were one of the few couples which came
               | early enough. Usually she saw couples well past the point
               | of no return.
               | 
               | So yeah, if you don't ask in time, you will probably be
               | breaking up anyway.
        
               | hnfong wrote:
               | I would speculate that, if a couple goes to a
               | professional for help, they have much better chances than
               | asking on a random forum online...
        
               | otabdeveloper4 wrote:
               | > relationships that don't lead anywhere
               | 
               | Relationships are not transactions that are supposed to
               | "lead somewhere".
        
               | ambicapter wrote:
               | You're being a bit pedantic here "leading somewhere" is
               | accepted shorthand for a lasting, satisfying relationship
               | that is good for both parties.
        
               | otabdeveloper4 wrote:
               | Relationships aren't transactional. This isn't a business
               | deal.
        
               | SpicyLemonZest wrote:
               | Most people engage in romantic relationships because
               | they'd like to find someone to marry and settle down
               | with. Nothing but respect for the people who've thought
               | it through and decided that's not for them, but what's
               | much more common is failing to think it through or
               | worrying it would be awkward/scary/"cringe" to take their
               | relationship goals seriously.
               | 
               | That's what people are pointing to when they talk about
               | relationships not "leading anywhere". If you want to be
               | married in 5-10 years, and you're 2 years into an OK
               | relationship with someone you don't want to marry, it's
               | going to suck to break up with them but you have to do it
               | anyway.
        
             | rusty_venture wrote:
             | Wait, other people don't make decision trees and mind maps
             | and pro/con lists and consult chatbots before making
             | decisions? Are they just flying through life by the seat of
             | their pants? That doesn't seem like a very solid framework
             | for achieving desired outcomes.
        
               | nprateem wrote:
               | I heard about someone once who could decide whether to
               | buy a new t-shirt in less than 3 months.
        
             | falcor84 wrote:
             | Maybe I'm too much of a hopeless romantic, but from my
             | perspective and experience, when someone is good for you,
             | you'll fight for that relationship regardless of what
             | others say, and conversely when you're in a situation where
             | your actively asking and willing to consider "leave" from
             | someone who isn't a very close friend or a therapist as
             | applicable, then it's likely you're looking for external
             | validation for what you've already essentially decided.
        
           | duskdozer wrote:
           | >asking an AI whether to leave your partner
           | 
           | is that what they're asking though? because "relationship
           | advice" is pretty vague
        
             | falcor84 wrote:
             | That's a good point. If an AI respond to a "what should I
             | get my boyfriend for Christmas?" with a "You should leave
             | him", that's a very different issue.
        
           | oldfrenchfries wrote:
           | The idea that asking implies a yes is actually a pretty
           | common logical fallacy. In relationship science, we call this
           | "Relational Ambivalence" and its a completely normal part of
           | any longterm commitment.
        
           | dinkumthinkum wrote:
           | No, but it is an indication of brain-rot to make a question
           | seriously and also to think that it means the conclusion is
           | foregone. It is an advent of our childlike current
           | generations. Of course, the moment anything becomes difficult
           | or unpleasant, one should quit, apparently. Surely, this kind
           | of resiliency is what got humanity so far.
        
             | falcor84 wrote:
             | I didn't imply it's a "foregone conclusion", but just said
             | it's an indication - in the sense of increasing the
             | likelihood. Just like a person asking an AI "what does it
             | feel like to bleed out?" could be them researching for a
             | novel, but is nevertheless an indication of a potential
             | serious issue.
        
         | Sharlin wrote:
         | More interesting, IMO, is the general trend that started long
         | before LLMs. The fact that "dump them" is the standard answer
         | to any relationship question is a meme by now. The LLMs appear
         | to be doing exactly what one would expect them to be doing
         | based on their training corpus.
        
           | 1970-01-01 wrote:
           | This is the correct take. The advice preceded the LLM boom.
           | They were trained on the 'dump them' advice and proceeded to
           | reinforce the take. So why did the relationship advice change
           | dramatically? I speculate attribution to the disinformation
           | campaigns during this time. They were and still are grossly
           | underestimated.
        
             | to11mtm wrote:
             | Not sure what sorts of disinformation campaigns you're
             | referring to...
             | 
             | There is something more interesting to consider however;
             | the graph starts to go up in 2013, less than 6 months after
             | the release of Tinder.
        
               | 1970-01-01 wrote:
               | These. https://en.wikipedia.org/wiki/Russian_disinformati
               | on#Social_...
        
           | doubled112 wrote:
           | "There is more than one fish in the sea" has been
           | relationship advice for centuries. It might be about being
           | dumped, but I've also thought it useful for considering
           | dumping somebody too.
        
             | Sharlin wrote:
             | No, that's not it. We're talking about posts like "we had a
             | silly little quarrel about something that would need
             | fifteen minutes to clear up and make both happy if we both
             | just try to adult a bit" and commenters being adamant that
             | deleting gym and facebooking up and so on is clearly the
             | only choice. Most of said commenters probably not being in
             | any position to give advice on relationships to others.
        
           | dec0dedab0de wrote:
           | if things are so bad that you're posting on reddit then
           | breaking up is usually the best answer.
        
             | nibbleyou wrote:
             | I see this being said often but I don't understand.
             | 
             | A lot of people posting there are young and may well be in
             | their first relationship. It makes sense for them to ask a
             | question in the community they spend their most time in -
             | which is reddit
        
             | the_af wrote:
             | Most people overshare on reddit and it's completely
             | unrelated to the seriousness of the situation.
             | 
             | It's also a meme that people will ask the dumbest, most
             | trivial interpersonal conflict questions on Reddit that
             | would be easily solved by just talking to the other person.
             | E.g. on r/boardgames, "I don't like to play boardgames but
             | my spouse loves them, what can I do?" or "someone listens
             | to music while playing but I find it distracting, what can
             | I do?" (The obvious answer of "talk to the other person and
             | solve it like grownups" is apparently never considered).
             | 
             | On relationship advice, it often takes the form "my
             | boy/girlfriend said something mean to me, what shall I do?"
             | (it's a meme now that the answer is often "dump them").
             | 
             | If LLMs train _on this_...
        
           | est wrote:
           | the year is 2015
           | 
           | smart phones took over the world, social networks happened.
           | 
           | Turns out they are the best sterializer human ever invented.
           | 
           | I just wrote a blog https://blog.est.im/2026/stdin-09
        
           | astrange wrote:
           | > The LLMs appear to be doing exactly what one would expect
           | them to be doing based on their training corpus.
           | 
           | That is not how full LLM training works. That is how base
           | model pretraining works.
        
             | Sharlin wrote:
             | Yes, yes, I know. But how many people have AI companies
             | hired to do RLHF who actually have the expertise to adjust
             | them away from biases like this? As opposed to paying a
             | dollar per day to a bunch of poor people in Africa?
        
         | jubilanti wrote:
         | Or that people are using AI to write perfectly calibrated
         | ragebait that gets upvoted with a bunch of genuine human
         | clicks.
        
         | raincole wrote:
         | Is this comment human hallucination? You can clearly see the
         | trend is always going up. It only went down a bit during Covid.
        
       | deeg wrote:
       | I do find them cloying at times. I was using Gemini to iterate
       | over a script and every time I asked it to make a change it
       | started a bunch of responses with "that's a smart final step for
       | this task! ...".
        
       | xiphias2 wrote:
       | Marc Andereseen has talked about the downside of RLHF: it's a
       | specific group of liberal low income people in California who did
       | the rating, so AI has been leaning their culture.
       | 
       | I think OpenAI tried to diversify at least the location of the
       | raters somewhat, but it's hard to diversify on every level.
        
         | michaelcampbell wrote:
         | Do you have any links to documentation of this? Andreesen has a
         | definite bias as well, so I'm not about to just accept his say-
         | so in a fit of Appeal to Authority.
         | 
         | (eg: "Cite?")
        
           | xiphias2 wrote:
           | He was talking about it in the Lex Friedman interview after
           | Trump was elected. And he was talking about a lot of things
           | the Biden administration forced on Silicon Valley at that
           | time (since then Google lost a case about one of these back-
           | deals).
        
             | michaelcampbell wrote:
             | So no evidence then. Kind of like Lex touting his bona
             | fides as a professor.
        
         | sph wrote:
         | What do low income people have to do with it, when AI companies
         | and research is borne out of Silicon Valley culture of rich,
         | liberal Californians?
         | 
         | I'm still waiting for models based on the curt and abrasive
         | stereotype of Eastern European programmers, as contrast to the
         | sickeningly cheerful AIs we have today that couldn't sound more
         | West Coast if they tried.
        
           | fourside wrote:
           | Low income and liberal is usually code for certain
           | "undesirables" that conservatives tend to dislike. Better
           | watch what LLM your kids use or they might end up speaking
           | Spanish and listening to rap ;).
        
             | dinkumthinkum wrote:
             | Eh, or grow up hating American and thinking they need to
             | fly to Cuba to explain to the people are great communism is
             | for them. Who knows.
        
             | xiphias2 wrote:
             | It's not about liking / disliking, but conservatives tend
             | to prefer staying together even if it's a bad relatioship,
             | and liberals prefer splitting by default if there are
             | serious problems.
             | 
             | The syncopath style is clearly categorized as more liberal
             | (do what you feel is good).
        
               | jibal wrote:
               | Does that explain Trump's numerous wives?
               | 
               | Reading your comments is a wonderland of right wing bias.
        
           | tbrownaw wrote:
           | > _What do low income people have to do with it, when AI
           | companies and research is borne out of Silicon Valley culture
           | of rich, liberal Californians?_
           | 
           | RLHF is "ask a human to score lots of LLM answers". So the
           | claim is that the AI companies are hiring cheap (~poor)
           | people from convenient locations (CA, since that's where the
           | rest of the company is).
        
             | astrange wrote:
             | "Poor" in California means earning $80k/year, so they
             | probably are not doing that. Africa / Indonesia /
             | Philippines are better places to find English speaking RLHF
             | workers.
        
             | sublinear wrote:
             | Yes, this precisely it. There _isn 't going to be_ hard
             | evidence to prove it though. Survey data that underpins
             | some empirical studies have similar transparency issues
             | too. This is far from a new problem.
             | 
             | If you adjust your mindset slightly when searching online,
             | it's not hard to find communities of people looking for
             | quick side work and this was huge during the covid lockdown
             | era. There were people helping train LLMs for all kinds of
             | purposes from education to customer service. Those startups
             | quickly cashed out a few years ago and sold to the big
             | players we have now.
             | 
             | I don't get why this is hard for people to believe (or
             | remember)?
        
           | cyanydeez wrote:
           | Poor people, to the billionaire, clearly are morally and
           | ethically unsound.
           | 
           | https://pmc.ncbi.nlm.nih.gov/articles/PMC9533286/
        
         | nirvdrum wrote:
         | For anyone else unfamiliar with the term:
         | 
         | RLHF = Reinforcement Learning from Human Feedback
         | 
         | https://en.wikipedia.org/wiki/Reinforcement_learning_from_hu...
        
         | mvkel wrote:
         | Marc Andreesen should get HF on his own RL, because he's
         | completely wrong.
         | 
         | This sounds like something Elon would say to make Grok seem
         | "totally more amazeballs," except "anti-woke" Grok suffers from
         | the same behavior
        
         | ej88 wrote:
         | huh? this is completely inaccurate
        
           | kibwen wrote:
           | You're absolutely right!
        
         | BoredPositron wrote:
         | Talked about as in lied about it and you taking his words for
         | gospel without verifying it? Looks just as bad as "Yes-Men" AI
         | models.
        
       | masteranza wrote:
       | We can surely fix it and we probably should. However, I don't
       | think AI is doing any worse here than friends advice when they
       | here a one sided story. The only difference being that it's not
       | getting studied.
       | 
       | Conversely, AI chatbots are great mediators if both parties are
       | present in the conversation.
        
       | tom-blk wrote:
       | Not surprising, but nice that we have actual data now
        
       | 152334H wrote:
       | Maybe it's not so sensible to offload the responsibility of clear
       | thinking to AI companies?
       | 
       | How is a chatbot supposed to determine when a user fools even
       | themselves about what they have experienced?
       | 
       | What 'tough love' can be given to one who, having been so
       | unreasonable throughout their lives - as to always invite scorn
       | and retort from all humans alike - is happy to interpret
       | engagement at all as a sign of approval?
        
         | isodev wrote:
         | > clear thinking
         | 
         | Most humans working in tech lack this particular attribute, let
         | alone tools driven by token-similarity (and not actual
         | 'thinking').
        
         | kibwen wrote:
         | _> Maybe it 's not so sensible to offload the responsibility of
         | clear thinking to AI companies?_
         | 
         | Markets don't optimize for what is sensible, they optimize for
         | what is profitable.
        
           | SlinkyOnStairs wrote:
           | It's not market driven. AI is ludicrously unprofitable for
           | nearly all involved.
        
             | cyanydeez wrote:
             | The profit appears to be capturing the political class and
             | it's associated lobbies and monied interests.
        
         | expedition32 wrote:
         | It's almost as if being a therapist is an actual job that takes
         | years of training and experience!
         | 
         | AI may one day rewrite Windows but it will never be counselor
         | Troi.
        
           | yarn_ wrote:
           | Come on, I'm sure Dario can find a nice tight bodysuit for
           | claude
        
           | duskdozer wrote:
           | Well, unless insurance companies figure out they can make
           | more money by pushing everyone onto AI [step-]therapy instead
           | of actual therapy
        
           | fsmv wrote:
           | Implying that programming is not an actual job that takes
           | years of training and experience
           | 
           | To be clear I don't think the AI can do either job
        
         | rsynnott wrote:
         | > How is a chatbot supposed to determine when a user fools even
         | themselves about what they have experienced?
         | 
         | And even if it _could_, note, from the article:
         | 
         | > Overall, the participants deemed sycophantic responses more
         | trustworthy and indicated they were more likely to return to
         | the sycophant AI for similar questions, the researchers found.
         | 
         | The vendors have a perverse incentive here; even if they
         | _could_ fix it, they'd lose money by doing so.
        
         | nicce wrote:
         | > Maybe it's not so sensible to offload the responsibility of
         | clear thinking to AI companies?
         | 
         | Maybe it's not so sensible to offload the responsibility of
         | tubacca addiction to tubacca companies?
        
       | sublinear wrote:
       | I think if you're at the stage of life where you even need to
       | ask, the AI might be doing everyone a favor.
       | 
       | As much as people whine about the birth rate and whatever else, I
       | think it's a net good that people spend a lot more time alone to
       | mature. Good relationships are underappreciated.
        
       | graemep wrote:
       | There are plenty of sycophantic humans around, especially with
       | regard to relationship advice.
       | 
       | I find there is an inverse relationship between how willing
       | people are to give relationship advice, and how good their advice
       | is (whether looking at sycophancy or other factors).
        
         | griffzhowl wrote:
         | Because sycophancy in humans is motivated not by the wellbeing
         | of the person seeking advice, but by the interests of the
         | sycophant in gaining favour.
         | 
         | It makes sense that this behaviour would be seen in LLMs, where
         | the company optimizes towards of success of the chatbot rather
         | than wellbeing of the users.
        
         | xhkkffbf wrote:
         | Yup. I know too many people who have a default message when
         | asked for relationship advice: oh, my, the other person is
         | terrible and you should break up.
         | 
         | It's an easy default and it causes so many problems.
        
           | graemep wrote:
           | Even if they do not go as far, telling someone they are right
           | to blame the other person as a default is damaging.
           | 
           | The opposite, encouraging people to stay in a relationship
           | when they should leave is also damaging.
        
       | megous wrote:
       | Can't you just prompt for a critical take, multiple alternative
       | perspectives (specifically not yours, after describing your own),
       | etc.?
       | 
       | It's a tool, I can bang my hand on purpose with a hammer, too.
        
         | ranger_danger wrote:
         | Yes, if you're smart. But most people asking it random
         | questions and expecting it to read their minds and spit out the
         | perfect answer are not so much. They don't know what a prompt
         | is, and wouldn't be bothered to give it prior instructions
         | either way.
        
           | megous wrote:
           | Educated, not smart. This is a job for schools to include AI
           | education into the basic curricula. Their pupils will use the
           | tools anyway, so at least teach them to do it with proper
           | expectations and prompting techniques/pitfalls.
        
           | joquarky wrote:
           | I think that the type of people who can easily pick up
           | subtext have come to rely on that channel of communication
           | and don't realize they need to be more direct and verbose
           | when chatting with language models.
        
       | awithrow wrote:
       | It feels like I'm fighting uphill battle when it comes to
       | bouncing ideas off of a model. I'll set things up in the context
       | with instructions similar to. "Help me refine my ideas,
       | challenge, push back, and don't just be agreeable." It works for
       | a bit but eventually the conversation creeps back into
       | complacency and syncophancy. I'll check it too by asking "are you
       | just placating me?" the funny thing is that often it'll admit
       | that, yes, it wasn't being very critical, and then procede to
       | over correct and become a complete contrarian. and not in a way
       | that's useful either. very frustrating. I've found that Opus 4.6
       | is worse about this than 4.5. 4.5 does a better job IMO of
       | following instructions and not drifting into the mode where it
       | acts like everything i say is a grand revelation from up high.
        
         | righthand wrote:
         | That's because the model isn't actually thinking, pushing back,
         | and challenging your ideas. It's just statistically agreeing
         | with you until it reaches too wide of a context. You're living
         | in the delusion that it's "working" or having a "conversation"
         | with you.
        
           | alehlopeh wrote:
           | How is conceptualizing what the model is doing as having a
           | conversation any different from any other abstraction? "No,
           | the browser isn't downloading a file. The electrons in the
           | silicon are actually..."
        
             | colechristensen wrote:
             | There are people with a philosophical objection to using
             | everyday words to describe LLM interactions for various
             | reasons, but commonly because they're worried stupid people
             | will confuse the LLM for a person. Which, I suppose stupid
             | people will do that, but I'm not inventing a parallel
             | language or putting a * next to each thing which means
             | "this, but with an LLM instead of a person"
        
               | cruffle_duffle wrote:
               | That is an interesting way of looking at that, thanks for
               | the perspective!
               | 
               | Like, the words fit... why create a second parallel
               | language for describing LLM behavior.
               | 
               | Somebody else said it... the whole "it's a stochastic
               | parrot" thing is sooooo cliche and boring at this point.
               | It's like, duh... what is your point?
        
             | righthand wrote:
             | One is an abstraction, another is a fantasy.
        
         | cyanydeez wrote:
         | So, there's things you're fighting against when trying to
         | constrain the behavior of the llm.
         | 
         | First, those beginning instructions are being quickly ignored
         | as the longer context changes the probabilities. After every
         | round, it get pushed into whatever context you drive towards.
         | The fix is chopping out that context and providing it before
         | each new round. something like `<rules><question><answer>` ->
         | `<question><answer><rules><question>`.
         | 
         | This would always preface your question with your prefered
         | rules and remove those rules from the end of the context.
         | 
         | The reason why this isn't done is because it poisons the KV
         | cache, and doing that causes the cloud companies to spin up
         | more inference.
        
         | margalabargala wrote:
         | Considering 4.6 came with a ton of changes around tooling and
         | prompting this isn't terribly surprising.
        
         | magicalhippo wrote:
         | Gemini seems to be fairly good at keeping the custom
         | instructions in mind. In mine I've told it to not assume my
         | ideas are good and provide critique where appropriate. And I
         | find it does that fairly well.
        
           | steve_adams_86 wrote:
           | Same. This works fine for Claude in my experience. My user
           | prompt is fairly large and encourages certain behaviours I
           | want to see, which involves being critical and considering
           | the strengths and weaknesses of ideas before drawing
           | conclusions. As someone else mentioned, there does seem to be
           | a phenomenon where saying DO NOT DO X causes a sort of
           | attention bias on X which can lead to X occurring despite the
           | clear instructions. I've never empirically tested that, I've
           | just noticed better results over the years when telling it
           | what paths to stick to rather than specific things not do to.
        
             | koverstreet wrote:
             | That happens with humans too :) It's why positive feedback
             | that draws attention to the behavior you want to encourage
             | often works better. "Attention" is lower level and more
             | fundamental than reasoning by syllogism.
        
             | magicalhippo wrote:
             | > there does seem to be a phenomenon where saying DO NOT DO
             | X causes a sort of attention bias on X which can lead to X
             | occurring despite the clear instructions
             | 
             | It's a thing with people too[1], ie do not think about a
             | white bear.
             | 
             | [1]: https://en.wikipedia.org/wiki/Ironic_process_theory
        
               | steve_adams_86 wrote:
               | Don't think about what?
        
           | lelanthran wrote:
           | > Gemini seems to be fairly good at keeping the custom
           | instructions in mind.
           | 
           | Unless those instructions are "stop providing links to you
           | for every question ".
        
             | magicalhippo wrote:
             | I think that's part of the RAG pipeline, at least to me it
             | looks to be separate from the model output. Models are
             | notorious for getting URLs wrong so makes sense to me to
             | handle it separately.
        
           | iugtmkbdfil834 wrote:
           | I will admit that I was very pleasantly surprised by gemini
           | lately. I was away from my PC and tried it on a whim for a
           | semi-random consumer question that led into smaller rabbit
           | hole. It seemed helpful enough and focused on what I tried to
           | get while still pushing back when my 'solutions' seemed out
           | of whack.
        
         | Loughla wrote:
         | That's because you need actual logic and thought to be able to
         | decide when to be critical and when to agree.
         | 
         | Chatbots can't do that. They can only predict what comes next
         | statistically. So, I guess you're asking if the average
         | Internet comment agrees with you or not.
         | 
         | I'm not sure there's much value there. Chatbots are good at
         | tasks (make this pdf an accessible word document or sort the
         | data by x), not decision making.
        
           | kvirani wrote:
           | I'm not convinced that "actual logic and thought" aren't just
           | about inferring what comes next statistically based on
           | experience.
        
             | plagiarist wrote:
             | Then the machines still need a more sophisticated
             | "experience" compared to what they have currently.
        
             | righthand wrote:
             | Communicating is usually about inferring. I dont think
             | token to token. And I don't think "well statistically I
             | could say 'and' next but I will say 'also' instead to give
             | my speech some flash". If I decided on swapping a word I
             | would have made my decision long ago, not in the moment.
             | Thought and logic are not me pouring through my brain
             | finding a statistical path to any answer. Often I stop and
             | say "I dont know".
        
             | Swizec wrote:
             | > I'm not convinced that "actual logic and thought" aren't
             | just about inferring what comes next statistically based on
             | experience.
             | 
             | Often they are the exact opposite. Entire fields of math
             | and science talk about this. Causation vs correlation,
             | confirmation bias, base rate fallacy, bayesian reasoning,
             | sharp shooter fallacy, etc.
             | 
             | All of those were developed because "inferring from
             | experience" leads you to the wrong conclusion.
        
               | theptip wrote:
               | Bayesian reasoning is just another algorithm for
               | predicting from experience (aka your prior).
               | 
               | I took the GP to be making a general point about the
               | power of "next x prediction" rather than the algorithm a
               | human would run when you say they are "inferring from
               | experience". (I may be assuming my own beliefs of
               | course.)
               | 
               | Eg even LeCun's rejection of LLMs to build world models
               | is still running a predictor, just in latent space (so
               | predicting next world-state, instead of next-token).
               | 
               | And of course, under the Predictive Processing model
               | there is a comprehensive explanation of human cognition
               | as hierarchical predictors. So it's a plausible general
               | model.
        
               | Swizec wrote:
               | > under the Predictive Processing model there is a
               | comprehensive explanation of human cognition as
               | hierarchical predictors
               | 
               | It's plausible!
               | 
               | But keep in mind humans have been explaining ourselves in
               | terms of the current most advanced technology for
               | centuries. We used to be kinda like clockwork, then a bit
               | like a steam engine, then a lot like computers, and now
               | we're just like AI.
               | 
               | That's why you blow a gasket or fuse, release some steam,
               | reboot your life, do brain dump, feel like a cog in the
               | machine, get your wires crossed, etc
        
             | theptip wrote:
             | Exactly. Lots can be explained just with more abstract
             | predictors, plus some mechanisms for stochastic rollout and
             | memory.
        
             | dinkumthinkum wrote:
             | Is this just Internet smart contrarianism or a real thing?
             | Are logic gates in a digital circuit just behaving
             | statistically according to their experience?
        
             | hluska wrote:
             | You know, you might really enjoy consumer behaviour. When
             | you get into the depths of it, you'll end up running
             | straight into that idea like you're doing a 100 metre dash
             | in a 90 metre gym. It's quite interesting how arguably the
             | best funded group under the psychology umbrella runs
             | directly into this. One of my favourite examples is how
             | heuristics will lead otherwise reasonable people to make
             | decisions that are not in their interest.
        
           | righthand wrote:
           | I said this pretty much and got major downvotes...
        
             | dTal wrote:
             | Because it's an outmoded cliche that never held much
             | philosophical weight to begin with and doesn't advance the
             | discussion usefully. "It's a stochastic parrot" is not a
             | useful predictor of actual LLM capabilities and never was.
             | Last year someone posted on HN a log of GPT-5 reverse
             | engineering some tricky assembly code, a challenge set by
             | another commentator as an example of "something LLMs could
             | never do". But here we are a year later still wading
             | through people who cannot accept that LLMs can, in a
             | meaningful sense, "compute".
        
               | righthand wrote:
               | It's entirely useful discussion because as soon as you
               | forget that it's not really having a conversation with
               | you, it's a deep dive into delusion that you're talking
               | to a smart robot and ignoring the fact that these smart
               | robots were trained on a pile of mostly garbage. When I
               | have a conversation with another human, I'm not expecting
               | them to brute force an answer to the topic. As soon as
               | you forget that Llms are just brute forcing token by
               | token then people start living in fantasy land. The whole
               | "it's not a stochastic parrot" is just "you're holding it
               | wrong".
        
               | layla5alive wrote:
               | Its not that LLMs are stochastic parrots and humans are
               | not. Its that many humans often sail through
               | conversations stochastic parroting because they're
               | mentally tired and "phoning it in" - so there are times
               | when talking to the LLM, which has a higher level of
               | knowledge, feels more fruitful on a topic than talking to
               | a human who doesn't have the bandwidth to give you their
               | full attention, and also lack the depth and breadth of
               | knowledge. I can go deep on many topics with LLMs that
               | most humans can't or won't keep up on. In the end, I'm
               | really only talking to myself most of the time in either
               | case, but the LLM is a more capable echo, and it doesn't
               | tire of talking about any topic - it can dive deep into
               | complex details, and catching its hallucinations is an
               | exercise in itself.
        
               | dinkumthinkum wrote:
               | No. It's quite a useful thing to understand So, what, you
               | have us believe it is a sentient, thinking, kind of
               | digital organism and you would have us not believe that
               | it is exactly what it is? Being wrong and being
               | unimaginative about what can be achieved with such a
               | "parrot" is not the same as being wrong about it be a
               | word predictor. If you don't think, you can probably ask
               | an LLM and it will even "admit" this fact. I do agree
               | that it has become considered to be outmoded to question
               | anything about the current AI Orthodox.
        
             | plagiarist wrote:
             | People are upset hearing that LLMs aren't sentient for some
             | reason. Expect to be downvoted, it is okay.
        
               | gjm11 wrote:
               | First off, "not adequately described as a mere token-
               | predictor" and "not sentient" are entirely separate
               | things.
               | 
               | I can't speak for anyone else, but what _I_ feel when I
               | read yet another glib  "it's just a stochastic parrot, of
               | course it isn't doing anything that deserves to be called
               | reasoning" take is much more like _bored_ than it is like
               | _upset_.
               | 
               | Today's LLMs are in some sense "just predicting tokens"
               | in some sense. Likewise, human brains are in some sense
               | "just shuttling neurotransmitters and electrical impulses
               | around" in some sense. Neither of those tells you what
               | the thing can actually do. To figure that out, you have
               | to _look at what it can do_.
               | 
               | Today's best LLMs can do about as well as the best humans
               | on problems from the International Mathematical Olympiad
               | and occasionally solve easyish actual mathematical
               | research problems. They write code about as well as a
               | junior software developer (better in some ways, worse in
               | others) but much faster. They write prose about as well
               | as an average educated person (but with some annoying
               | quirks that are annoying mostly because they are the same
               | quirks over and over again).
               | 
               | If it pleases you to call those things "thinking" then
               | you can. If it pleases you to call them "stochastic
               | parroting" then you can. They are the same things either
               | way. They are not, on the face of it, very much like
               | "just repeating things the machine has already seen", or
               | at least not _more_ like that than a lot of things
               | intelligent human beings do that we don 't usually
               | describe that way.
               | 
               | If you want to know whether an LLM can do some particular
               | thing -- do your job well enough for your boss to fire
               | you, write advertising copy that will successfully sell
               | products, exterminate the human race, whatever -- then
               | it's not enough to say "it's just remixing what it's seen
               | on the internet, therefore it can't do X" unless you also
               | have _good reason to believe that that thing can 't be
               | done by just "remixing what's on the internet"_ (in
               | whatever sense of "remixing" the LLM is doing that). And
               | it's turning out that lots of things can be done that way
               | that you absolutely wouldn't have predicted five years
               | ago could be done that way.
               | 
               | It seems to me that this should make us very cautious
               | about saying "they can't do X because all they can do is
               | regurgitate a combination of things they've seen in
               | training".
               | 
               | (My own view, not that there's any reason why anyone
               | should care what I-in-particular think, is a combination
               | of "what they're doing is less parroting than you might
               | have thought" and "you can do more by parroting than you
               | might have thought".)
               | 
               | So, anyway, this particular instance of the stochastic-
               | parrot argument started when someone said: of course the
               | AIs are yes-men, because figuring out when to agree and
               | when not to requires actual logic and thought and the
               | LLMs don't have either of those things.
               | 
               | Is it really clear that deciding whether or not to agree
               | when someone says "I think maybe I should break up with
               | my girlfriend" or "I've got this amazing new theory of
               | physics that the establishment is stupidly dismissing"
               | requires _more logic and thought_ than, say, gold-medal
               | performance on IMO problems? It certainly isn 't clear to
               | me. Having done a couple of International Mathematical
               | Olympiads myself in my tragically unmisspent youth, I can
               | assure you that solving their problems requires quite a
               | bit of logic and thought, at least for humans. It may
               | well be _harder_ to give a good answer to  "should I
               | leave my job?", but it's not exactly "logic and thought"
               | that it needs more of.
               | 
               | Someone reported that Claude is much less yes-man-ish
               | than Gemini and ChatGPT. I don't know whether that's true
               | (though it wouldn't surprise me) but: suppose it is; do
               | you want that to oblige you to say that yes, actually,
               | Claude really thinks logically, unlike Gemini and
               | ChatGPT? I don't think you do. And if not, you want to
               | avoid saying "duh, of course, you can't avoid being a
               | yes-man without actually thinking and reasoning, and we
               | all know that LLMs can't do those things".
        
               | rustystump wrote:
               | I wont touch how profoundly i disagree with everything
               | you said on reasoning (u clearly already have it figured
               | out) but a fun test i have done with most of the big
               | models is to give it some text input, maybe a short
               | story, and have it rate it. That is, the prompt is, rate
               | this from 1-10.
               | 
               | For Gemini and gpt, it almost always will give very
               | similar scores for everything. As long as grammar isnt
               | off u cannot get below a 7.
               | 
               | X ai on the other hand will rarely give anything above a
               | 7.
               | 
               | Now when u prompt with, rate 1-10 with 5 being average,
               | all the sudden the scores of openai and gemini drop and x
               | ai remains roughly the same.
               | 
               | All of them will eventually give you a 10 if u keep
               | making tiny edits "fixing" whatever they complain about.
               | 
               | Humans do not do this. Or more specifically, my
               | experience with humans.
        
         | secret_agent wrote:
         | Use positive requests for behavior. For some reason, counter
         | prompts "Don't do X" seems to put more attention on X than the
         | "Don't do." It's something like target fixation, "Oh shit I
         | don't want to hit that pothole..." _bang_
        
           | ambicapter wrote:
           | This is a well known problem in these kind of systems. I'm
           | not 100% on what the issue is mechanically but it's something
           | like they can only represent the existence of things and not
           | non-existence so you end up with a sort of "don't think of
           | the pink elephant" type of problem.
        
             | SpicyLemonZest wrote:
             | Isn't it just that, in the underlying text distribution,
             | both "X" and "don't do X" are positively correlated with
             | the subsequent presence of X? I've never seen that analysis
             | run directly but it would surprise me if it weren't true.
        
         | RugnirViking wrote:
         | check out this article that was posted here a while back
         | https://www.randalolson.com/2026/02/07/the-are-you-sure-prob...
         | 
         | The article's main idea is that for an AI, sycophancy or
         | adversarial (contrarian) are the two available modes only. It's
         | because they don't have enough context to make defensible
         | decisions. You need to include a bunch of fuzzy stuff around
         | the situation, far more than it strictly "needs" to help it
         | stick to its guns and actually make decisions confidently
         | 
         | I think this is interesting as an idea. I do find that when I
         | give really detailed context about my team, other teams, ours
         | and their okrs, goals, things I know people like or are
         | passionate about, it gives better answers and is more
         | confident. but its also often wrong, or overindexes on these
         | things I have written. In practise, its very difficult to get
         | enough of this on paper without a: holding a frankly worrying
         | level of sensitive information (is it a good idea to write down
         | what I really think of various people's weaknesses and
         | strengths?) and b: spending hours each day merely establishing
         | ongoing context of what I heard at lunch or who's off sick
         | today or whatever, plus I know that research shows longer
         | context can degrade performance, so in theory you want to
         | somehow cut it down to only that which truly matters for the
         | task at hand and and and... goodness gracious its all very time
         | consuming and im not sure its worth the squeeze
        
           | cruffle_duffle wrote:
           | > goodness gracious its all very time consuming and im not
           | sure its worth the squeeze
           | 
           | And when you step back you start to wonder if all you are
           | doing is trying to get the model to echo what you already
           | know in your gut back to you.
        
           | awithrow wrote:
           | oh that's great. thanks for the link!
        
           | oldfrenchfries wrote:
           | This is great, thanks for sharing!
        
         | dkersten wrote:
         | I find Kimi white good if you ask it for critical feedback.
         | 
         | It's BRUTAL but offers solutions.
        
           | ohyoutravel wrote:
           | Not soft, not mild, but BRUTAL! This broke my brain!
        
           | awithrow wrote:
           | what is Kimi white?
        
             | dkersten wrote:
             | I was typing quickly on my phone. I meant "quite", "I find
             | Kimi quite good"
        
         | dinkumthinkum wrote:
         | You're not wrong and you're not crazy. In fact, you are
         | absolutely right! It is not just These things are not just
         | casual enablers. They are full-on palace sycophants following
         | the naked emperor showering him with praise for his sartorial
         | elegance. /s
        
         | rsynnott wrote:
         | Why not... do this with a person, instead? Other humans are
         | available.
         | 
         | (Seriously, I don't understand this. Plenty of humans will be
         | only too happy to argue with you.)
        
           | awithrow wrote:
           | oh i do as well. I think of the LLM as another tool in the
           | toolbox, not a replacement for interactions. There is
           | something different about having a rubber duck as a service
           | though.
        
           | kelseyfrog wrote:
           | "the percentage of U.S. adults who report having no close
           | friends has quadrupled to 12% since 1990"[1]
           | 
           | 1. https://www.happiness.hks.harvard.edu/february-2025-issue/
           | th...
        
             | nathan_compton wrote:
             | More technology is probably the solution to this!
        
           | layla5alive wrote:
           | Many other humans are .... Not very available - certainly
           | many shut down when conversations reach a certain level of
           | depth or require great focus or introspection..
        
             | balamatom wrote:
             | Depth? Introspection?
             | 
             | I'd say these days the norm is to not simply shut down, but
             | to become irrevocably and insidiously hostile, the moment
             | someone hints at the existence of such a thing as "ground
             | truth", "subjective interpretation", "being right or wrong"
             | - or any of the bits and bobs that might lead one to
             | discover the proper scary notion, "consensus reality".
             | 
             | "What do you mean social reality is a constructed by the
             | consensus of the participants? Reality is what has been
             | drilled into my head under threat of starvation! How dare
             | you exist!", et cetera. You've heard it translated into
             | Business English countless times.
             | 
             | They are deathly afraid of becoming aware of their own
             | conditioned state of teleological illiteracy - i.e. how
             | they are trained to know what they are doing, but _never_
             | why they are doing it. It 's especially bad with the guys
             | who cosplay US STEM gang.
             | 
             | One is not permitted a position of significance in this
             | world without receiving this conditioning, and I figure
             | it's precisely this global state of cognitive disavowal
             | which props up the value of the US dollar - and all sorts
             | of other standees you might've recently interacted with _as
             | if_ they 're not 2D cutouts (metaphorical ones!
             | metaphorical!).
             | 
             | PSA: Look up "locus of control" and "double bind". Between
             | those two, you might be able to get a glimpse of what's
             | going on - but have some sort of non-addictive sedative
             | handy in case you do.
        
               | gverrilla wrote:
               | I think you will enjoy Guy Debord and Raoul Vaneigem.
        
               | balamatom wrote:
               | Just like a certain defense minister was shown to enjoy
               | D&G; after which the latter were never heard of again.
               | Where they go, eh?
               | 
               | +1 for Vaneigem, he has a nice cryptohistory of Nalka;
               | and you might also want to check out Villem Flusser.
        
               | VonTum wrote:
               | You had me on the first three paragraphs, but the last
               | two veer so far off course that I've no idea what you're
               | trying to say. Mind clarifying?
        
               | balamatom wrote:
               | Yes
        
             | rsynnott wrote:
             | > when conversations reach a certain level of depth or
             | require great focus or introspection..
             | 
             | I mean... if the alternative is an LLM... you realise that
             | the LLM isn't doing any focusing or introspection, right?
        
           | mock-possum wrote:
           | Arguing with a human costs social energy. Chatting with a
           | robot does not.
        
             | balamatom wrote:
             | s/social/demonic/
        
           | balamatom wrote:
           | OK, I'll bite the artillery shell: I don't mean to dismiss
           | you or what you are saying; in fact I strongly relate -
           | wouldn't it be nice to be able to hash things out with people
           | and mutually benefit from both the shared and the diverging
           | perspectives implied in such interaction? Isn't that the most
           | natural thing in the world?
           | 
           | Unfortunately these days this sounds halfway between a very
           | privileged perspective and a pie in the sky.
           | 
           | When was the last time a person took responsibility for the
           | bad outcome you got as a direct consequence of following
           | their advice?
           | 
           | And, relatedly, where the hell do you even find humans who
           | believe in discursive truth-seeking in 2026CE?
           | 
           | Because for the last 15 years or so I've only ever ran into
           | (a) the kind of people who will keep arguing regardless if
           | what they're saying is proven wrong; (b) and their
           | complementaries, those who will never think about what you
           | are saying, lest they commit to saying anything definite
           | themselves, which may hypothetically be proven wrong.
           | 
           | Thing is, both types of people have plenty to lose; the magic
           | wordball doesn't. (The previous sentence is my answer to the
           | question you posited; and why I feel the present
           | parenthesized disclaimer to be necessary, is a whole next can
           | of worms...)
           | 
           | Signs of the existence of other kinds of people, perhaps such
           | that have nothing to prove, are not unheard of.
           | 
           | But those people reside in some other layer of the social
           | superstructure, where facts matter much less than adherence
           | to "humane", "rational" not-even-dogmas (I'd rather liken it
           | to complex conditioning).
           | 
           | But those folks (because reasons) are in a position of power
           | over your well-being - and (because unfathomables) it's a
           | definite faux pas to insist in their presence that there are
           | such things as _facts_ , which relate by the principles of
           | verbal reasoning.
           | 
           | Best you could get out of them is the "you do you", "if you
           | know you know", that sort of bubble-bobble - and don't you
           | dare get even mildly miffed at such treatment of your natural
           | desire to keep other humans in the loop.
           | 
           | AI is a symptom.
        
             | hluska wrote:
             | When you start hearing things like "you do you" or "if you
             | know you know" it means that you went way too far. That's a
             | sign of discomfort.
             | 
             | If you make uncomfortable, you won't get diverging
             | perspectives. People will agree to anything to get out of a
             | social situation that makes them uncomfortable.
             | 
             | If your goal is meaningful conversation, you may want to
             | consider how you make people feel.
        
               | balamatom wrote:
               | Believe me (or don't), I always do. Even when this
               | precludes a necessary conversation from happening. Even
               | when the other party doesn't give a fuck about how they
               | make others feel.
               | 
               | After all, if they're making _me_ uncomfortable, surely
               | there 's something making _them_ uncomfortable, which
               | they 're not being _able to be_ forthright about, but
               | with empathy I could figure it out from contextual cues,
               | right?
               | 
               | >People will agree to anything to get out of a social
               | situation that makes them uncomfortable.
               | 
               | That's fine as long as they have someone to take care of
               | them.
               | 
               | In my experience, taking into account the opinions of
               | such people has been the worst mistake of my life. I'm
               | still working on the means to correct its consequences.
               | 
               | "Doing whatever for the sake of avoiding mild discomfort"
               | is cowardice, laziness, narcissism - I'm personally
               | partial to the last one, but take your pick. In any case,
               | I see it as a way of being which is _taught to_ people;
               | and one which is fundamentally dishonest and
               | irresponsible.
               | 
               | Other than that, I do agree with your overall sentiment
               | and the underlying value system; I'm just not so sure any
               | more that it is in fact correct.
        
               | balamatom wrote:
               | Believe me (or don't), I always do. Even when this
               | precludes a necessary conversation from happening. Even
               | when the other party doesn't give a fuck about how they
               | make others feel.
               | 
               | After all, if they're making me uncomfortable, surely
               | there's something making them uncomfortable, which
               | they're not being able to be forthright about, but with
               | empathy I could figure it out from contextual cues,
               | right?
               | 
               | >People will agree to anything to get out of a social
               | situation that makes them uncomfortable.
               | 
               | That's fine as long as they have someone to take care of
               | them.
               | 
               | In my experience, taking into account the opinions of
               | such people has been the worst mistake of my life. I'm
               | still working on the means to fix its consequences, as
               | much as they are fixable at all.
               | 
               | "Doing whatever for the sake of avoiding mild discomfort"
               | is cowardice, laziness, narcissism - I'm personally
               | partial to the last one, but take your pick. In any case,
               | I consider it a fundamentally dishonest attitude, and a
               | priori have no wish to get along (i.e. become
               | interdependent) with such people.
               | 
               | Other than that, I do agree with your overall sentiment
               | and the underlying value system; I'm just not so sure any
               | more that it is in fact correct.
        
               | nuancebydefault wrote:
               | > In my experience, taking into account the opinions of
               | such people has been the worst mistake of my life. I'm
               | still working on the means to fix its consequences, as
               | much as they are fixable at all.
               | 
               | This sounds very cryptic. Can you give an example?
        
               | balamatom wrote:
               | Certainly. Can you guarantee my safety afterwards?
        
             | rustystump wrote:
             | I genuinely do not understand what u are saying. Because
             | reasons, because unfathomables? Everyone in last 15 years
             | has been an npc? I have had countless deep conversations
             | with people and i am an uber introvert.
             | 
             | This reads like someone who is deep into their specific
             | pov. You cannot hope to have a meaningful conversation if
             | you yourself are not willing to concede a point.
             | 
             | To the op u are replying too, arguing with people can have
             | real consequences if u say something stupid or carelessly.
             | There is a another human there. With a machine, u are safe.
             | At least u feel safe.
        
             | nuancebydefault wrote:
             | Why is your wording so complicated? It is very hard for me
             | to understand what you try to say, even though I am very
             | interested.
        
               | balamatom wrote:
               | TL;DR: Probably because I'm having fun and you are
               | expending effort. Hope you find what I say to be worth
               | the effort.
               | 
               | To preface, I do not take offense to your remark, because
               | you seem to be asking in good faith.
               | 
               | (If, however, being unable to immediately recognize pre-
               | known patterns in my speech had automagically led you to
               | the conclusion that I am somehow out of line, just for
               | speaking how I speak ... well, then we woulda hadda
               | problemo! But we don't, chill on.)
               | 
               | So, honest question deserves honest answer.
               | 
               | The short of it is: English sux.
               | 
               | Many many many people, much much much smarter than me
               | (and much better compensated too!) have been working
               | throughout modernity to make it literally _impossible_ to
               | express much of anything interesting in English.
               | 
               | (Well, not without either being a fictional character or
               | sounding batshit insane, anyway! But that joke's entirely
               | on "the Them": I am not only entirely fictional, but have
               | an equal amount of experience being batshit insane in my
               | native language and in the present _lingua franca_. So,
               | consider all I say cognitohazardous and watch out for
               | colors you ain 't seen before, dawg!)
               | 
               | Linguistic hegemony is the thing that LLMs are the
               | steroids for - surfuckingprise! - and that's why your
               | commanders love 'em.
               | 
               | As opposed to programming languages, which your superiors
               | loathe and your peers viscerally refuse to acknowledge,
               | because those are the exact opposite thing: descending
               | from mathemathical notation, and being evaluated by a
               | machine, they have the useful property of _being
               | incapable of expressing lies and nonsense_.
               | 
               | Direct computing confers what you could call _bullshit-
               | resistance_. That property is a treasure underappreciated
               | by virtue of its unfamiliarity, and one which we are in
               | the process of being robbed of.
               | 
               | I also want to admit that linguistic hegemony isn't all
               | downside: English is great for technical and instrumental
               | knowledge - especially with elided bells and whistles
               | (adverbs, copula, etc.)
               | 
               | But then life ain't all business, izzet?
               | 
               | Imagine you have a partner who wants to have a
               | conversation about feelings and interpersonal relations;
               | and not even in a scary way, right? So you sit and talk
               | about stuff, and your partner does this thing where they
               | keep switching from your shared native tongue to English
               | mid-sentence, _in order to be able to talk about such
               | things better_ , because your native tongue does not have
               | - no, not only the established words and notions! - _it
               | doesn 't have the basic grammatical constructs for
               | expressing simple things unambiguously_, so if you were
               | to attempt the same conversation in nativelang you'd end
               | up battling it out with proverbs and anodyne canards
               | ripped from propaganda repertoire of the prior regime.
               | 
               | Fun, no?
               | 
               | As an exercise, try imagining what notions are absent
               | from modern English. And don't forget to remain vigilant.
               | Love from our table to your table!
        
               | nuancebydefault wrote:
               | You lost me well before the anodyne canards...
               | 
               | When talking about feelings, we now and then throw in an
               | English word because some things are expressed in much
               | less words when using English. In a few cases even a
               | German word. Uberhaupt is for example a word for which i
               | do not know an alternative in any language.
               | 
               | I think you want to say that human language is too
               | ambiguous for clear communication between human and
               | machine. The machine might mis interpret what you write.
               | Classic computer languages leave no room for
               | interpretation.
               | 
               | For the rest I think you might be a little lost. That is
               | okay, so many of us are. I wish you all the best.
        
               | balamatom wrote:
               | >I think you want to say that human language is too
               | ambiguous for clear communication between human and
               | machine. [...]
               | 
               | If that is what I wanted to say, I figure I would not
               | have had much difficulty with saying exactly it - and not
               | something else.
               | 
               | Except I fail to see the purpose of making that
               | statement.
               | 
               | Maybe to have some people say "it is true! I agree with
               | what the balamatom is saying"?
               | 
               | Again - to what end? How would that agreement be of use
               | to me?
               | 
               | Why say something which both speaker and listener have
               | already heard a thousand times? To get a cracker and be
               | called pretty?
               | 
               | And have I lost the author, or have I lost the reader, or
               | we all so lost that it doesn't matter how lost each is?
               | Maybe one day we will all become so lost that it will
               | once again begin to matter where exactly we are! Counting
               | on it.
        
             | rsynnott wrote:
             | > When was the last time a person took responsibility for
             | the bad outcome you got as a direct consequence of
             | following their advice?
             | 
             | ... I mean, the LLM certainly isn't going to do that.
        
           | layer8 wrote:
           | In addition to availability, usually because you want to take
           | advantage of the knowledge that is baked into the models,
           | which for all its flaws still vastly exceeds the knowledge of
           | any single human.
        
             | emp17344 wrote:
             | For this use case, how do LLMs provide more value than a
             | standard search engine? They may actually be destroying
             | value here, as LLM-generated text pollutes search results.
        
               | layer8 wrote:
               | They let you find things faster, and they combine and
               | synthesize information from different sources. I'm more
               | on the AI skeptic side and always prided myself on my
               | Google foo, but nowadays chatbots can really save a lot
               | of time with that.
               | 
               | Publishing LLM-generated text is a separate use case, I'm
               | not a friend of that.
        
           | ip26 wrote:
           | No living breathing human deserves to be subjected to my
           | level of overthinking, and vanishingly few share my
           | fascination with my favorite topics.
        
           | matheusmoreira wrote:
           | > Other humans are available.
           | 
           | Are they?
           | 
           | I have some personal projects that I enjoy working on
           | whenever I have free time. One of them is a lisp interpreter.
           | I just overhauled its memory allocator, now I'm working on
           | the hash tables. Would you like to help me develop it?
        
         | raincole wrote:
         | My rule of thumb:
         | 
         | 1. Only one shot or two shot. Never try to have a prolonged
         | conversation with an LLM.
         | 
         | 2. Give specific numbers. Like "give me two alternative
         | libraries" or "tell me three possible ways this might fail."
        
         | post-it wrote:
         | > I'll check it too by asking "are you just placating me?" the
         | funny thing is that often it'll admit that, yes, it wasn't
         | being very critical, and then procede to over correct and
         | become a complete contrarian. and not in a way that's useful
         | either.
         | 
         | It's not _admitting_ anything. Your question diverts it down a
         | path where it acts the part of a former sycophant who is now
         | being critical, because that question is now upstream of its
         | current state.
         | 
         | Never make the mistake of asking an LLM about its intentions.
         | It doesn't have any intentions, but your question will alter
         | its behaviour.
        
           | godelski wrote:
           | > Your question diverts it down a path where it acts the part
           | of a former sycophant who is now being critical
           | 
           | I think people really have a hard time understanding a
           | sycophant can be contrarian. But a yesman can say yes by
           | saying no
           | 
           | https://news.ycombinator.com/item?id=47484664
        
             | cheesecompiler wrote:
             | > a yesman can say yes by saying no
             | 
             | What a great way to summarize LLM behaviour in 2026
        
               | noisy_boy wrote:
               | >> a yesman can say yes by saying no
               | 
               | >What a great way to summarize LLM behaviour in 2026
               | 
               | Well they have been trained on words spoken by humans and
               | that has been a human behaviour since time immemorial.
               | E.g.: "I do not agree with you that you were wrong. I do
               | apologize for my strong disagreement but we actually do
               | need your continued guidance desperately."
        
           | layer8 wrote:
           | I think "admit" here is just a description of what the LLM
           | was saying. It doesn't imply that the OP thinks the LLM has
           | internal beliefs matching that.
        
             | awithrow wrote:
             | Correct, it's really anthropomorphizing what gets
             | generated. It's another form of the "you're absolutely
             | right" framing. The interesting thing is sometimes it won't
             | do that. It'll continue to insist that some point or
             | another is still a valid interpretation. I'm some ways it
             | all feels like a complicated way of reading tea leaves
        
           | cruffle_duffle wrote:
           | An alternate way of thinking about it is LLMs have no
           | reflection capability. Literally any "reflection" it claims
           | to have about its decision making is made up. It has
           | absolutely no way know that what it said was based on some
           | ancient proverb, the phase of the moon or cold hard rational
           | thought.
        
           | mvkel wrote:
           | > it'll admit that, yes, it wasn't being very critical, and
           | then procede to over correct and become a complete contrarian
           | 
           | Which is also placating you
        
             | awithrow wrote:
             | Yes exactly! It's very frustrating. My guess is it drives
             | engagement. But some model that tells me what I want to
             | hear all the time is useless.
        
           | Terr_ wrote:
           | > It doesn't have any intentions
           | 
           | Yeah, and in a way it's even worse than that, since there's
           | another layer of cognitive illusion: "It" doesn't exist.
           | 
           | The LLM algorithm is an ego-less document-generator, often
           | applied to growing a document that resembles dialogue between
           | two fictional characters.
           | 
           | So when your human-user character is "asking" the AI
           | assistant character to explain its intentions, that's the
           | same as asking a Count Dracula character to describe what it
           | "really feels like" to become a cloud of bats.
           | 
           | You'll see something interesting, but it'll be what fits
           | trained story-patterns rather than what any mind introspects
           | or perceives.
        
         | ajkjk wrote:
         | 'admit' isn't really the right word for that... the fact that
         | it was placating you wasn't true until you prompted it to say
         | so. Unlike a person who has an 'internal emotional state'
         | independent of what they say that you can probe by asking
         | questions.
        
           | awithrow wrote:
           | 'admit' is anthropomorphizing the behavior, sure. The point
           | is that sometimes the model's response will tighten, flag
           | things that were overly supportive or what not. Sometimes it
           | wont, it'll state that previous positions are still supported
           | and continue to press it. Its not like either response is
           | 'correct' but it can alter the rest of the responses in ways
           | that are useful.
        
         | colechristensen wrote:
         | >"Help me refine my ideas, challenge, push back, and don't just
         | be agreeable."
         | 
         | This is where you're doing it wrong.
         | 
         | If your LLM has a problem being more agreeable than you want,
         | prompt it in a way that makes being agreeable contrary to your
         | real intentions.
         | 
         | "there are bugs and logic problems in this code" "find the
         | strongest refutation of this argument" "I don't like this plan
         | and need to develop a solid argument against it"
         | 
         | Asking for top ten lists is a good method, it will rarely not
         | come up with anything but you can go back and forth and refine
         | until it's 10 ten reasons why your plan is bad are all
         | insubstantial nonsense then you've made progress
        
         | Forgeties79 wrote:
         | I usually put "do not praise me, do not use emojis, I just want
         | straight answers" something along those lines and it's been
         | surprisingly effective. Though it helps I can't run
         | particularly heavy duty models/don't carry on the
         | "conversation" for super long durations.
        
         | anandram27 wrote:
         | Could be an aspect of eval awareness mb
        
         | GuB-42 wrote:
         | I find the best way is to give the LLM as little information as
         | possible about where you want to go. For example don't say "I
         | think pineapple pizzas are the best, am I right?", say "What is
         | the general consensus on pineapple pizzas?".
        
         | arjie wrote:
         | I haven't found that work at all. But the way I usually frame
         | it for this is that I present it as someone else's work that I
         | disagree with (sometimes "this is a junior engineer's work and
         | he often does dumb shit" as an introduction to an essay). In
         | this way, it tries to be agreeable with me and can rationalize
         | various opposition. And then I can select from that.
        
         | jimmyjazz14 wrote:
         | Yeah, I have never had good results with refining ideas with
         | models or really any interactions with models outside of rote
         | task such as coding or analyzing document structures, I don't
         | know why I was ever surprised by this as its obvious that LLMs
         | just aren't capable of original thinking. I think part of the
         | problem is that these things were marketed originally as
         | chatbots when that is honestly their weakest use-case. I think
         | even when I was expressly try to not anthropomorphize LLMs I
         | still sorta did in early days, but the less I do so the more
         | utility I get from them.
        
         | sinuhe69 wrote:
         | The longer you converse with the LLM, the more frequent the
         | same keywords will occur and Thu's they will only strengthen
         | the probabilistic choices of the LLM. The LLM do not exactly
         | follow the instructions like a program would do, nor does it
         | truly understand its role or what its "said". It looks fine on
         | first tries and short texts, but will drift more the longer the
         | conversation goes. The only way the LLM can maintain the
         | appearance of a normal conversation is to repeatedly feed it
         | with the users requests. We can see it with the shell construct
         | the AI company built around their chat interface.
        
       | justin_dash wrote:
       | So at this point I think it's pretty obvious that RLHFing LLMs to
       | follow instructions causes this.
       | 
       | I'm interested in a loop of ["criticize this code harshly" ->
       | "now implement those changes" -> open new chat, repeat]: If we
       | could graph objective code quality versus iterations, what would
       | that graph look like? I tried it out a couple of times but ran
       | out of Claude usage.
       | 
       | Also, how those results would look like depending on how complete
       | of a set of specs you give it.
        
         | IncreasePosts wrote:
         | In my experience prompting llms to be critical leads then to
         | imagine issues, or to bike shed
        
           | joquarky wrote:
           | I noticed when I ask it to find something to improve in a
           | project, that certain frivolous topics would arise regularly.
           | I now use their appearance as a sign that there is nothing
           | meaningful to improve.
        
       | neya wrote:
       | WTF is "yes-men"?
       | 
       | Orignal title:
       | 
       |  _AI overly affirms users asking for personal advice_
       | 
       | Dear mods, can we keep the title neutral please instead of
       | enforcing gender bias?
        
         | oldfrenchfries wrote:
         | Thats a fair point on the title. I used "Yes-Men" as a
         | colloquialism for the "sycophancy" described in the Stanford
         | paper, but overly affirming or sycophantic is definitely more
         | precise and neutral. I cant edit the title anymore, but I
         | appreciate the catch.
        
           | cyanydeez wrote:
           | New title: "LLMs treat you like a Billionaire; you're not"
        
           | nemo44x wrote:
           | Don't apologize to these types of people. It will only make
           | your problem worse as now you're an admitted offender. Ignore
           | them or better yet laugh at them to put their insane ideas
           | back on the margins where they belong.
        
           | neya wrote:
           | All good. I thought it was a gendered reference and learned
           | that it isn't. My bad.
        
         | nprateem wrote:
         | Lol. How do you function in daily life?
        
           | neya wrote:
           | Same as you, why is that so hard for you to grasp?
        
             | mikkupikku wrote:
             | My dude, you're objecting to the use of a perfectly
             | ordinary English idiom because it doesn't advance your
             | personal ideology (which few other people in this world
             | share with you.) How do you get through a day without
             | melting down because somebody said "mailman"?
        
               | neya wrote:
               | > my dude
               | 
               | This is the problem I'm trying to highlight. For one, I'm
               | not "your dude". I don't even know you like that.
               | 
               | If you want to correct me on the idiom usage, be my
               | guest. 2) Mailman and yes-man aren't even the same
               | logical comparison. Mailman is a profession. Yes men is a
               | label.
               | 
               | The acoustics inside your head must be incredible.
        
               | nprateem wrote:
               | Chill bro. You've probably got undiagnosed autism. Worth
               | getting checked out.
        
             | joquarky wrote:
             | PCU (1994)
        
         | skvmb wrote:
         | https://www.merriam-webster.com/dictionary/yes-man
        
         | 9rx wrote:
         | _> gender bias_
         | 
         | It is funny that you originally recognized and found it
         | necessary to call out that AI isn't human, but then made the
         | exact same mistake yourself in the very same comment. I expect
         | the term you are looking for is "ontological bias".
        
         | dinkumthinkum wrote:
         | Gender bias? I could understand if you felt the title was more
         | provocative in signaling sycophancy but what gender bias? I'm
         | confused. Is this some kind of California thing?
        
       | svara wrote:
       | Yeah, and if you ask it to be critical specifically to get a
       | different perspective or just to avoid this bias, it'll go over
       | the top in the opposite direction.
       | 
       | This is imo currently the top chatbot failure mode. The insidious
       | thing is that it often feels good to read these things. Factual
       | accuracy by contrast has gotten very good.
       | 
       | I think there's a deeper philosophical dimension to this though,
       | in that it relates to alignment.
       | 
       | There are situations where in the grand scheme of things the
       | right thing to do would be for the chatbot to push back hard, be
       | harsh and dismissive. But is it the really aligned with the human
       | then? Which human?
        
       | righthand wrote:
       | LLMs are syncophatic digital lawyers that will tell you what you
       | want to hear until you look at the price tag and say "how much
       | did I spend?!"
        
       | gurachek wrote:
       | I had exactly this between two LLMs in my project. An evaluator
       | model that was supposed to grade a coaching model's work. Except
       | it could see the coach's notes, so it just... agreed with
       | everything. Coach says "user improved on conciseness", next
       | answer is shorter, evaluator says yep great progress. The answer
       | was shorter because the question was easier lol.
       | 
       | I only caught it because I looked at actual score numbers after
       | like 2 weeks of thinking everything was fine. Scores were
       | completely flat the whole time. Fix was dumb and obvious -- just
       | don't let the evaluator see anything the coach wrote. Only raw
       | scores. Immediately started flagging stuff that wasn't working.
       | Kinda wild that the default behavior for LLMs is to just validate
       | whatever context they're given.
        
         | joquarky wrote:
         | This is probably why these models can't say "I don't know". If
         | they could, then that would be the only response they would
         | give for everything.
        
           | gurachek wrote:
           | Yeah, I think so. So far, Claude Opus is the only model I
           | found that doesn't fold under the minimal pressure and can
           | push back, but still - push just a little bit harder and it's
           | back to "appear productive and useful to the user". I don't
           | even have an idea how to balance it in LLMs to keep their
           | business alive :D
        
       | bryanrasmussen wrote:
       | somewhere an AI chatbot is reading this and confirming eagerly
       | that this is indeed one of its problems and vowing to do better
       | next time.
        
       | fathermarz wrote:
       | This is a skill in life with people as much as it is with LLMs.
       | One should always question everything and build strongman
       | arguments for one's self. Using a pros and cons approach brings
       | it back to reality in most cases, especially when it comes to
       | _serious matters_.
       | 
       | It's less about "challenge my thinking" and more about playing it
       | out in long tail scenarios, thought exercises, mental models, and
       | devils advocate.
        
       | jordanb wrote:
       | Billionaires love AI chatboats so much because they invented the
       | digital Yes-man. They agree obsequiously with everything we say
       | to them. Unfortunately for the rest of us we don't really have
       | the resources to protect ourselves from our bad decisions and
       | really need that critical feedback.
        
       | stared wrote:
       | There is a fine line between "following my instructions" (is what
       | I want it to do) vs "thinking all I do is great" (risky, and
       | annoying).
       | 
       | A good engineer will also list issues or problems, but at the
       | same time won't do other than required because (s)he "knows
       | better".
       | 
       | The worst is that it is impossible to switch off this constant
       | praise. I mean, it is so ingrained in fine tuning, that prompt
       | engineering (or at least - my attempts) just mask it a bit, but
       | hard to do so without turning it into a contrarian.
       | 
       | But I guess the main issue (or rather - motivation) is most
       | people like "do I look good in this dress?" level of reassurance
       | (and honesty). It may work well for style and decoration. It may
       | work worse if we design technical infrastructure, and there is
       | more ground truth than whether it seems nice.
        
       | maddmann wrote:
       | This paper feels a bit biased in that it is trying to prove a
       | point versus report on results objectively. But if you look at
       | the results of study 3, doesn't it suggest that there are ai
       | models that can improve how people handle interpersonal
       | conflict?! Why isn't that discussed more?
        
       | gAI wrote:
       | You're essentially summoning a character to role-play with. Just
       | like with esoteric evocation, it's very easy to summon the wrong
       | aspect of the spirit. Anthropic has a lot to say about this:
       | 
       | https://www.anthropic.com/research/persona-selection-model
       | 
       | https://www.anthropic.com/research/assistant-axis
       | 
       | https://www.anthropic.com/research/persona-vectors
        
         | hammock wrote:
         | Unfortunately (after reading your links) all of the control
         | surfaces for mitigating spirit summoning seem to be in the
         | model training, creation and tuning not something you can
         | change meaningfully through prompting.
         | 
         | Perhaps the LLM itself, rather than the role model you created
         | in one particular chat conversation or another, is better
         | understood to be the "spirit."
         | 
         | As a non-coder who only chats with pre existing LLMs and
         | doesn't train or tune them, I feel mostly powerless.
        
           | gAI wrote:
           | As I understand it, it's more that the training (and training
           | data set) bake in the concept attractor space
           | (https://arxiv.org/abs/2601.11575). So the available
           | characters are fixed, yes, and some are much stronger
           | attractors than others. But we still have a fair amount of
           | control over which archetype steps into the circle. As an
           | aside, this is also why jailbreaking is fundamentally
           | unsolved. It's not difficult to call the characters with dark
           | traits. They're strong attractors, in spite of (or because
           | of?) the effort put into strengthening the pull of the
           | Assistant character.
        
           | est wrote:
           | I present you
           | 
           | NVIDIA Nemotron-Personas-USA -- 1 million synthetic Americans
           | whose demographics match real US census distributions
           | 
           | https://huggingface.co/datasets/nvidia/Nemotron-Personas-USA
        
           | darepublic wrote:
           | > As a non-coder who only chats with pre existing LLMs and
           | doesn't train or tune them, I feel mostly powerless.
           | 
           | You realize in regards to only using and not training LLMs
           | you are in the triple 9 majority right. Even if we only
           | considered so called coders
        
         | rdevilla wrote:
         | Spot on.
        
         | jerf wrote:
         | I am polite when using AI, not because I mistake it for a
         | human, but because I'm deliberately keeping it in the
         | "professional colleague" persona. Tell it to push back, and
         | then _thank_ it for something it finds in your error. I may put
         | a small self-deprecating joke in from time to time. It keeps
         | the  "mood" correct.
         | 
         | Another way you can think of it is that when you're talking to
         | an AI, you're not talking to a human, you're talking to
         | distillation of _humanity_ , as a whole, in a box. You want to
         | be selective in what portion of humanity you are leading to be
         | dominant in a conversation for some purpose. There's a lot in
         | there. There's a lot of conversations where someone makes a
         | good critical point and a flamewar is the response. A lot of
         | conversations where things get hostile. I'm sure the subsequent
         | RHLF helps with that, but it doesn't hurt anything to try to
         | help it along.
         | 
         | I see people post their screenshots of an AI pushing back and
         | asking the user to do it or some other AI to do it, and while
         | I'm as amused as the next person, I wonder what is in their
         | context window when that happens.
        
           | layer8 wrote:
           | > you're talking to distillation of humanity, as a whole, in
           | a box.
           | 
           | This is an aside, but my impression is that it is a very
           | selective and skewed distillation, heavily colored by
           | English-language internet discourse and other lopsided
           | properties of its training material, and by whoever RLHF'd
           | it. Relatively far away from being representative of the
           | whole of humanity.
        
             | jerf wrote:
             | Yes, absolutely. I'm not trying to claim it's some sort of
             | unbiased sample, but more get across the idea that
             | modelling AI as _a_ person, a singular person, in your head
             | is inaccurate. That singular person would have a
             | stereotypical, Hollywood-esque multiple personality
             | disorder like no actual human on Earth has ever had. You
             | need to be thinking about not just what the person-like
             | thing in front of you is doing, but how to craft _which_
             | person you 're ending up with.
        
           | iugtmkbdfil834 wrote:
           | Similar approach works for me. But then I also have a
           | separate checks at the end of the session basically
           | questioning the premise and logic used for most things except
           | brainstorming, where I allow more leeway. You can ask to be
           | challenged and challenged effectively, but now I wonder if
           | people do that.
        
           | gAI wrote:
           | Agreed, putting effort into my side of the role-play almost
           | always improves the model's responses. The attention required
           | to do that also makes it more likely that I'll notice when
           | the conversation first starts going off the rails: when it
           | hits the phase transition (https://arxiv.org/abs/2508.01097).
           | It does still seem important to start new chats regularly,
           | regardless of growing context sizes.
        
       | youknownothing wrote:
       | I think the problem stems from the fact that we have a number of
       | implicit parameters in our heads that allow us to evaluate pros
       | and cons but, unless we communicate those parameters explicitly,
       | the AI cannot take them into account. We ask it to be "objective"
       | but, more and more, I'm of the opinion that there isn't such a
       | thing as objectivity, what we call objectivity is just shared
       | subjectivity; since the AI doesn't know whose shared subjectivity
       | we fall under, it cannot be really objetive.
       | 
       | I tend to use one of these tricks if not both:
       | 
       | - Formulate questions as open-ended as possible, without trying
       | to hint at what your preference is. - Exploit the sycophantic
       | behaviour in your favour. Use two sessions, in one of them you
       | say that X is your idea and want arguments to defend it. In the
       | other one you say that X is a colleague's idea (one you dislike)
       | and that you need arguments to turn it down. Then it's up to you
       | to evaluate and combine the responses.
        
         | rossdavidh wrote:
         | If the algorithm (whatever it is) evaluates its own output
         | based on whether or not the user responds positively, then it
         | will over time become better and better at telling people what
         | they want to hear.
         | 
         | It is analogous to social media feeding people a constant
         | stream of outrage because that's what caused them to click on
         | the link. You could tell people "don't click on ragebait
         | links", and if most people didn't then presumably social media
         | would not have become doomscrolling nightmares, but at scale
         | that's not what's likely to happen. Most people will click on
         | ragebait, and most people will prefer sycophantic feedback.
         | Therefore, since the algorithm is designed to get better and
         | better at keeping users engaged, it will become worse and worse
         | in the more fundamental sense. That's kind of baked into the
         | architecture.
        
         | delusional wrote:
         | > I'm of the opinion that there isn't such a thing as
         | objectivity
         | 
         | So you have rejected objective reality over accepting the
         | evidence that "AI" contains no thinking or intelligence? That
         | sounds unwise to me.
        
       | potatoskins wrote:
       | Gemini is like a devil in this sense - i asked a relationship
       | advice and it just bounced pretty nasty stuff.
        
         | moichael wrote:
         | Yeah out of curiosity I asked ChatGPT a question about a
         | personal situation and its reply was absolutely scorched-earth
         | mode, telling me to get a lawyer etc over what was almost
         | nothing.
        
           | dinkumthinkum wrote:
           | Ah, all the Reddit posts are really showing up from the
           | training data, I see.
        
       | wisemanwillhear wrote:
       | With AI, I often like to act like a 3rd party who doesn't have
       | skin in the game and ask the AI to give the strongest criticisms
       | of both sides. Acting like I hold the opposite position as I
       | truly hold can help sometimes as well. Pretending to change my
       | mind is another trick. The idea is to keep the AI from guessing
       | where I stand.
        
         | mynameisvlad wrote:
         | I will generally ask for the "devil's advocate" view and then
         | have it challenge my views and opinions and iterate through
         | that.
         | 
         | It generally does a pretty good job as long as you understand
         | the tooling and are making conscious efforts to go against the
         | "yes man" default.
        
         | post-it wrote:
         | > Acting like I hold the opposite position as I truly hold can
         | help sometimes as well.
         | 
         | I find this helps a lot. So does taking a step back from my
         | actual question. Like if there's a mysterious sound coming from
         | my car and I think it might be the coolant pump, I just
         | describe the sound, I don't mention the pump. If the AI then
         | independently mentions the pump, there's a good chance I'm on
         | the right track.
         | 
         | Being familiar with the scientific method, and techniques for
         | blinding studies, helps a lot, because this is a lot like
         | trying to not influence study participants.
        
           | cruffle_duffle wrote:
           | A lot of getting good mileage out of LLMs is promoting them
           | to behave like they are blind and can only base their outputs
           | on what is in front of them. Maintain an emic stance.
        
           | nicce wrote:
           | I have tried it a lot aswell. A single mistake and it guesses
           | the side and changes the tone.
        
         | DrewADesign wrote:
         | Sounds like rubber-ducking with extra steps, tbh.
        
       | potatoskins wrote:
       | Yeah, I asked Gemini some relationship advice, it just goes
       | straight into cut-throat mode. I almost broke up with my
       | girlfriend, but then changed to Claude with another prompt.
        
       | wewxjfq wrote:
       | When I ask an LLM to help me decide something, I have to remind
       | myself of the LotR meme where Bilbo asks the AI chat why he
       | shouldn't keep the ring and he receives the classic "You're
       | absolutely right, .." slop response. They always go in the
       | direction you want them to go and their utility is that they make
       | you feel better about the decision you wanted to take yourself.
        
       | rsynnott wrote:
       | > They also included 2,000 prompts based on posts from the Reddit
       | community r/AmITheAsshole, where the consensus of Redditors was
       | that the poster was indeed in the wrong
       | 
       | Holy shit, then it's _very_ bad, because AmITheAsshole is
       | _itself_ overly-agreeable, and very prone to telling assholes
       | that they are not assholes (their 'NAH' verdict tends to be
       | this).
       | 
       | More seriously, why the hell are people asking the magic robot
       | for relationship advice? This seems even more unwise than asking
       | Reddit for relationship advice.
       | 
       | > Overall, the participants deemed sycophantic responses more
       | trustworthy and indicated they were more likely to return to the
       | sycophant AI for similar questions, the researchers found.
       | 
       | Which is... a worry, as it incentivises the vendors to make these
       | things _more_ dangerous.
        
       | astennumero wrote:
       | I always add the following at the end of every prompt. "Be
       | realistic and do not be sycophantic". Which will always takes the
       | conversation to brutal dark corners and panic inducing negative
       | side.
        
         | Lionga wrote:
         | Don't forget a good old "don't hallucinate" in your proompting
         | skills
        
       | dimgl wrote:
       | Even as someone who (wrongly) believed that I had high emotional
       | intelligence, I too was bit by this. Almost a year ago when LLMs
       | were starting to become more ubiquitous and powerful I discussed
       | a big life/professional decision with an LLM over the course of
       | many months. I took its recommendation. Ultimately it turned out
       | to be the wrong decision.
       | 
       | Thankfully it was recoverable, but it really sobered me up on
       | LLMs. The fault is on me, to be clear, as LLMs are just a tool.
       | The issue is that lots of LLMs try to come across as
       | interpersonal and friendly, which lulls users into a false sense
       | of security. So I don't know what my trajectory would have been
       | if I were a teenager with these powerful tools.
       | 
       | I do think that the LLMs have gotten much better at this,
       | especially Claude, and will often push back on bad choices. But
       | my opinion of LLMs has forever changed. I wonder how many other
       | terrible choices people have made because these tools convinced
       | them to make a bad decision.
        
         | potatoskins wrote:
         | Yeah, I think Claude is a lot more logical in that sense, I use
         | it for some therapy sessions myself and it pushes back a bit
         | more than Open AI and Gemini
        
           | Forgeties79 wrote:
           | I would be very careful doing this
        
             | potatoskins wrote:
             | You always have to be careful with LLMs, but to be fair, I
             | felt like Claude is such a good therapist, at least it is
             | good to start with if you want to unpack yourself. I have
             | been to 3 short human therapist sessions in my life, and I
             | only felt some kind of genuine self-improvement and
             | progress with Claude.
        
               | QuiDortDine wrote:
               | And how do you draw the line between feeling progress and
               | actually making progress?
        
               | layer8 wrote:
               | The same way you distinguish between feeling like having
               | a problem and actually having a problem.
        
               | Forgeties79 wrote:
               | This is needlessly flippant and not really the same
               | thing. Determining progress in a therapy setting is
               | usually a collaborative effort between the therapist and
               | the client. An LLM is not a reliable agent to make that
               | determination.
        
               | layer8 wrote:
               | I didn't claim that an LLM is that, and I fully agree
               | that it is not. I'm saying that one is inherently one's
               | own judge of whether one has a problem. You go to a
               | therapist when you feel you have a problem that warrants
               | it. You stop going when you feel you don't have it
               | anymore. And OP is very likely assessing their progress
               | in the same way. I wasn't being flippant if the parent
               | was asking a genuine question.
        
               | Forgeties79 wrote:
               | > I'm saying that one is inherently one's own judge of
               | whether one has a problem. You go to a therapist when you
               | feel you have a problem that warrants it
               | 
               | That is for certain types of therapy/clinical care. It is
               | not always - and often isn't - the case. Plenty of
               | diagnoses and care protocols are not a matter of opinion
               | or based on "you feeling there's an issue" or deciding on
               | your own there is no longer an issue.
        
               | logifail wrote:
               | > Determining progress in a therapy setting is usually a
               | collaborative effort between the therapist and the
               | client. An LLM is not a reliable agent to make that
               | determination
               | 
               | Can anyone describe how to determine how a (professional,
               | human) therapist is "a reliable agent" to make such a
               | determination?
        
               | Forgeties79 wrote:
               | If you want to call into question the entire field of
               | behavioral health and the training that is involved then
               | that is fine, but if that's how you feel then this entire
               | discussion is really about something different and I
               | can't bridge the gap here.
        
               | mlrtime wrote:
               | The thing they have in common is that they will both go
               | forever....
               | 
               | Meaning neither the LLM or the licensed therapist will
               | voluntarily say, you are healed, you don't need me
               | anymore.
        
               | Forgeties79 wrote:
               | Because that's not really how therapy works
        
               | logifail wrote:
               | Is there a generally-agreed description of "how therapy
               | works"?
        
               | moduspol wrote:
               | Counter-point: I often raise the same question of people
               | with human therapists. I do not get strong responses.
        
               | Forgeties79 wrote:
               | An LLM is completely unable to make that determination.
               | They can't even see you. So much information is lost when
               | it's text-only.
               | 
               | Where are you often asking this question/getting these
               | weak responses?
        
               | moduspol wrote:
               | You don't ask the therapist. You ask the person seeking
               | therapy.
        
             | shimman wrote:
             | You can't be careful at all doing this, this is like
             | smoking a cigarette in a dynamite factory.
             | 
             | Using LLMs for therapy is so deeply dystopian and
             | disgusting, people need human empathy for therapy. LLMs do
             | not emit empathy.
             | 
             | Complete disaster waiting to happen for that individual.
        
               | DrewADesign wrote:
               | Using an LLM for therapy is like using an iPad as an all-
               | purpose child attention pacifier. Sure, it's convenient.
               | Sure there's no immediate harm. Why a stressed parent
               | would be attracted to the idea is obvious... and of
               | course it's a terrible idea.
        
               | wiseowise wrote:
               | Using LLM for therapy is like using iPad as a study
               | assistance. It doesn't replace your work.
        
               | DrewADesign wrote:
               | It's nothing like that. Using an iPad for study
               | assistance is a conduit to many credible sources and
               | tools. They can be evaluated using context, reputation,
               | reviews, etc.
               | 
               | An LLM generates non-deterministic information using
               | sources you can't even know, let alone evaluate, and is
               | more primed to agree with you than give critical and
               | objective evaluation. It is, at best, like asking your
               | closest parent to help you through difficult
               | interpersonal situations: The interaction is probably,
               | subconsciously, going to be skewed enough towards
               | soothing you that you just can't consider it objective.
               | The difference is that with an LLM, that's deliberate.
               | It's designed in.
        
               | astrange wrote:
               | Claudes have lots of empathy. The issue is the opposite -
               | it isn't very good at challenging you and it's not
               | capable of independently verifying you're not
               | bullshitting it or lying about your own situation.
               | 
               | But it's better than talking to yourself or an abuser!
        
               | bloomca wrote:
               | It's about the same as talking to yourself, LLMs simply
               | agree with anything you say unless it is directly
               | harmful. Definitely agree about talking to an abuser,
               | though.
               | 
               | Sometimes people indeed just need validation and it helps
               | them a lot, in that case LLMs can work. Alternatively, I
               | assume some people just put the whole situation into
               | words and that alone helps.
               | 
               | But if someone needs something else, they can be straight
               | up dangerous.
        
               | astrange wrote:
               | > It's about the same as talking to yourself, LLMs simply
               | agree with anything you say unless it is directly
               | harmful.
               | 
               | They have world knowledge and are capable of explaining
               | things and doing web searches. That's enough to help. I
               | mean, sometimes people just need answers to questions.
        
               | JoshTriplett wrote:
               | > It's about the same as talking to yourself
               | 
               | In one way it's potentially worse than talking to
               | yourself. Some part of you might recognize that you need
               | to talk to someone other than yourself; an LLM might make
               | you feel like you've done that, while reinforcing
               | whatever you think rather than breaking you out of
               | patterns.
               | 
               | Also, LLMs can have more resources and do some "creative"
               | enabling of a person stuck in a loop, so if you are
               | thinking dangerous things but lack the wherewithal to put
               | them into action, an LLM could make you more dangerous
               | (to yourself or to others).
        
               | nuancebydefault wrote:
               | My experience is that it tries to look at your situation
               | in an objective way, and tries to help you to analyse
               | your thoughts and actions. It comes across as very
               | empathetic though, so there can lie a danger if you are
               | easily persuaded into seeing it as a friend.
        
               | worksonmine wrote:
               | It doesn't try to do anything. It doesn't work like that.
               | It regurgitates the most likely tokens found in the
               | training set.
        
               | nuancebydefault wrote:
               | Hmmmm i didn't know that... so a machine is not human is
               | your point? Look, i know it doesn't try, just like a
               | sorting algo does not try to sort, or an article does not
               | try to convey an opinion and a law does not try to make
               | society more organized.
        
               | cruffle_duffle wrote:
               | That is so reductive of an analysis that it is almost
               | worthless. Technically true, but very unhelpful in terms
               | of using an LLM.
               | 
               | It is a first principle though so it helps to "stir the
               | context windows pot" by having it pull in research and
               | other shit on the web that will help ground it and not
               | just tell you exactly what you prompt it to say.
        
               | worksonmine wrote:
               | They are amazing tools, but when people try to give them
               | agency someone has to explain it in simple terms.
        
               | Forgeties79 wrote:
               | >in an objective way
               | 
               | One of the great myths of models in countless
               | fields/industries. LLM's are absolutely in _no way_
               | objective.
               | 
               | Now if you want to say it's an "outside opinion" that's
               | valid. But do not kid yourself into thinking it is
               | somehow empirical or objective
        
               | brookside wrote:
               | Therapy can be basically paid (I would argue - false)
               | empathy, and in many cases vastly inferior to less
               | transactional relationships.
        
               | ycombinator_acc wrote:
               | Are you a therapist?
        
             | mettamage wrote:
             | In my experience, one also needs to be careful with actual
             | therapists
        
           | kortilla wrote:
           | Don't call them therapy sessions. They kind of look like it
           | but ultimately these are smoke blowing machines, which is
           | very far from what a therapist would do.
        
             | saghm wrote:
             | Six decades later and we're still trying to explain to
             | people the same things[1]:
             | 
             | > Some of ELIZA's responses were so convincing that
             | Weizenbaum and several others have anecdotes of users
             | becoming emotionally attached to the program, occasionally
             | forgetting that they were conversing with a computer.
             | Weizenbaum's own secretary reportedly asked Weizenbaum to
             | leave the room so that she and ELIZA could have a real
             | conversation. Weizenbaum was surprised by this, later
             | writing: "I had not realized ... that extremely short
             | exposures to a relatively simple computer program could
             | induce powerful delusional thinking in quite normal
             | people."
             | 
             | [1]: https://en.wikipedia.org/wiki/ELIZA
        
             | mlrtime wrote:
             | People generalize "therapy" all the time.
             | 
             | I [half] jokingly tell people that a nice drive in a 911
             | [can be] better/cheaper therapy that going to a PhD.
        
           | borski wrote:
           | https://news.ycombinator.com/item?id=47395779
        
           | throw4847285 wrote:
           | I'm starting to fear that LLMs are especially popular with
           | people who can't call a doctor's office to make an
           | appointment or tell a waiter they brought the wrong food. I
           | struggle with those things, but I know that it's better to
           | push myself outside of my comfort zone.
           | 
           | Are LLMs mainly a tool for people with atrophying social
           | skills who want all interaction to happen via a text prompt
           | that always responds in the most soothing way? I can't think
           | of another reason to replace human therapy with an LLM.
        
         | layla5alive wrote:
         | Any more context you're willing to share?
        
           | xXSLAYERXx wrote:
           | We really do love dirty laundry don't we? I'm sure whatever
           | the context is, it is deeply personal. Do you also have your
           | popcorn ready?
        
             | dimgl wrote:
             | Thank you. Yes, I'm going to refrain from airing out my
             | dirty laundry. I made a bad decision, now I'm living with
             | it, and more context doesn't actually change the intent
             | behind my message: these tools are dangerous. Getting
             | better, but still dangerous.
        
               | firtoz wrote:
               | Was it a blatantly bad idea or was it some risk that
               | triggered that would have been beyond your typical risk
               | threshold otherwise?
        
               | thepasch wrote:
               | > Yes, I'm going to refrain from airing out my dirty
               | laundry. I made a bad decision, now I'm living with it,
               | and more context doesn't actually change the intent
               | behind my message
               | 
               | That's not entirely true, as it's currently impossible to
               | actually gauge the severity of what the LLM seemingly
               | enabled you into doing. There's a difference between "I
               | uncritically accepted everything it told me because it
               | lined up with what I was hoping to hear" and "it subtly
               | nudged me towards a course of action that was going to be
               | obviously unwise after some consideration, but managed to
               | convince me to skip this"; and _also_ between that and "I
               | took a risk, which I knew to be a risk, and which I knew
               | to potentially expect to go bad, and the LLM convinced me
               | to take it where I otherwise wouldn't have", and _ALSO_
               | between _that_ and "I took a risk, which I knew to be a
               | risk, and which I knew to potentially expect to go bad,
               | and if I'm perfectly honest, I might've taken it anyway
               | without the LLM".
               | 
               | Without any indication as to how your situation maps to
               | any of these (or more), the warning is, functionally, not
               | particularly useful.
        
               | rcfox wrote:
               | Yeah, my first thought (admittedly an absurd one) went to
               | something along the lines of:
               | 
               | "I flipped a coin and the LLM called heads. I should have
               | gone with tails..."
        
         | colechristensen wrote:
         | >"'And it is also said,' answered Frodo: 'Go not to the Elves
         | for counsel, for they will say both no and yes.'
         | 
         | >"'Is it indeed?' laughed Gildor. 'Elves seldom give unguarded
         | advice, for advice is a dangerous gift, even from the wise to
         | the wise, and all courses may run ill...'"
         | 
         | This is the only way you should solicit personal advice from an
         | LLM.
        
         | davyAdewoyin wrote:
         | I largely agree, I also thought I was smart enough not to be
         | deluded into a false sense of security, but interacting with an
         | LLM is so tricky and slippery that, more often than not you are
         | forced to believe you just solve a problem no one had solve in
         | a hundred years.
         | 
         | My guideline now for interacting with LLM is only to believe
         | the result if it is factual and easily testable, or if I'm a
         | domain expert. Anything else especially if I'm in complete
         | ignorance about the subject is to approach with a high degree
         | of suspicion that I can be led astray by its sycophancy.
        
         | lovecg wrote:
         | Let's just hope that the people in charge of the really
         | important decisions that affect us all approach LLM generated
         | advice with the same wisdom.
        
           | saghm wrote:
           | They don't: https://fortune.com/2026/03/17/krafton-
           | subnautica-chatgpt-de...
        
             | paulhebert wrote:
             | Thanks for sharing this. Subnautica is one of my favorite
             | games so I was very excited for the sequel and very
             | frustrated by this move by Krafton.
             | 
             | It's even more maddening that this greedy maneuver was
             | orchestrated based on LLM advice.
             | 
             | I'm glad the subnautica team won the lawsuit. Maybe I can
             | play it now wothout feeling guilty
        
         | notracks wrote:
         | I recently found out that Claude's latest model, Sonnet 4.6,
         | scores the highest in Bullsh*tBench[0] (Funny name - I know).
         | It's a recent benchmark that measures whether an LLM refuses
         | nonsense or pushes back on bad choices so Claude has definitely
         | gotten better.
         | 
         | [0] - https://petergpt.github.io/bullshit-
         | benchmark/viewer/index.v...
        
           | layer8 wrote:
           | You don't have to star out things like that on HN.
        
             | thin_carapace wrote:
             | it would be interesting to me if you could explain the
             | motivation behind posting your comment. from my
             | perspective, if somebody with 5 years of forum tenure had
             | the intelligence to comment about advanced benchmarks, they
             | probably noticed that censorship was a voluntary decision
             | here, and had made a personal decision on that front.
        
               | mkl wrote:
               | I'm not layer8, but I had a similar thought. In this case
               | the needless censoring is problematic because it hides
               | the name of the benchmark from future searches (the
               | uncensored URL spells it differently).
        
               | layer8 wrote:
               | Such self-censoring is often done out of habit or a
               | mistakenly assumed obligation to do so. I consider it
               | inappropriate here, as it obscures an actual name,
               | doesn't constitute an expletive, and the HN readership is
               | generally mature enough to recognize that. The
               | counterquestion is, what justified reason could there
               | possibly be to censor it here? I don't think there is
               | any, in the sense that people wouldn't take any offense
               | at the uncensored version, and the intent of my comment
               | was to inform about that.
        
               | notracks wrote:
               | I censored it out of habit of commenting on other
               | platforms and, I actually didn't have any idea about
               | whether you should censor such words or not in here. Will
               | keep that in mind when commenting here next time.
        
           | astrange wrote:
           | I haven't tried talking to Sonnet much, but Opus 4.6 is very
           | sycophantic. Not in the sense of explicitly always agreeing
           | with you, but its answers strictly conform to the worldview
           | in your questions and don't go outside it or disagree with
           | it.
           | 
           | It _does_ love to explicitly agree with anything it finds in
           | web search though.
           | 
           | (Anthropic tries to fight this by adding a hidden prompt that
           | makes it disagree with you and tell you to go to bed, which
           | doesn't help.)
        
             | sidrag22 wrote:
             | the go to bed thing gets annoying, you can't even hint that
             | you are almost done or wrapping up or something or this is
             | hyper triggered and it never stops.
             | 
             | I do like when opus is incredibly short in its responses to
             | prompts that probably shouldnt have been made though. keeps
             | me grounded a bit.
        
           | akurilin wrote:
           | Great link, thanks for sharing. Confirmed what I saw
           | empirically by comparing the different models during daily
           | use.
        
           | uniq7 wrote:
           | Good call on censoring yourself preemptively, otherwise HN
           | could demonetize your comment
        
         | whodidntante wrote:
         | I think that if you go to an AI for advice and emotional
         | support, it will do what most people will do - tell you what it
         | thinks you want to hear. I am not surprised about this at all,
         | and I do notice that when you veer into these areas, it can do
         | it in a surprisingly subtle and dangerous way.
         | 
         | I try to focus on results. Things like an app that does what
         | you want, data and reports that you need, or technical things
         | like setting up a server, setting up a database, building a
         | website, etc.
         | 
         | I have also found it useful for feedback and advice, but only
         | once I have had it generate data that I can verify. For
         | example, financial analysis or modelling, health advice (again
         | factual based), tax modelling, etc, but again, all based on
         | verifiable data/tables/charts.
         | 
         | I am very surprised on what Claude is capable of, across the
         | entire tech stack: code, sysadmin, system integration,
         | security. I find it scary. Not just speed, but also quality and
         | the mental load is a difference of kind not quantity.
         | 
         | Personal advice on life decisions/relationships ? No way I
         | would go there.
         | 
         | It is also good for me to know that the tools I have built, the
         | data I have gathered, and my thinking approach places me as one
         | of the most intelligent developers and analysts in the world.
        
           | stephbook wrote:
           | Nice joke, hadn't seen it coming
        
             | KellyCriterion wrote:
             | Sounds like AI-written, eh? :-D
             | 
             | (esp last sentence?)
        
           | cruffle_duffle wrote:
           | That is why you have to always have it ground itself in
           | something. Have it search for relevant research or
           | professional whatever and pull that into context. Otherwise
           | it's just your word plus its training data.
           | 
           | I had to deal with a close family friend going through
           | alcohol withdrawal and getting checked in at a recovery
           | clinic for detox and used Claude heavily. The first thing I
           | had it do as do that "deep research" around the topic of
           | alcohol addiction, withdrawal, etc... and then made that a
           | project document along with clear guidelines about how it
           | shouldn't make inferences beyond what it in its context and
           | supporting docs. We also spent a whole session crafting a
           | good set of instructions (making sure it was using Anthropics
           | own guidelines for its model...)
           | 
           | Little differences in prompts make a huge deal in the output.
           | 
           | I dunno. It is possible to use these models for dumping crazy
           | shit you are going through. But don't kid yourself about
           | their output and aggressively find ways to stomp out things
           | it has no real way to authoritatively say.
        
           | alexjplant wrote:
           | > I think that if you go to an AI for advice and emotional
           | support, it will do what most people will do - tell you what
           | it thinks you want to hear.
           | 
           | Open two windows, ask it the same thing from starkly opposite
           | perspectives, then see what it comes back with. If nothing
           | else this exercise forces you to think deeply about what
           | you're considering before you even see what the giant blob of
           | matrix multiplication says about your situation.
        
         | jt2190 wrote:
         | I'm struggling to understand how the advice coming from an LLM
         | is any more or less "good" than advice coming from a human. Or
         | is this less about the "advice" part of LLMs and more about the
         | "personable" part, i.e. you felt more at ease seeking and
         | trusting this kind of advice form an LLM?
        
           | nuancebydefault wrote:
           | It is much easier to share personal feelings with an llm, i
           | found. Also it tried to keep me happy to get the conversation
           | going, but for me it feels mostly 'objective' or the most
           | socially acceptable advice, e. g. keeping a good relationship
           | is more important than trying a new one with someone else
           | because you 'feel something' around them. For me it tried to
           | find out together the sources or causes of that feeling, e.g.
           | you recognize parts of yourself in someone else or in the
           | past you had very good or very bad experiences around an
           | encounter.
        
             | jt2190 wrote:
             | Interesting thanks for elaborating.
        
           | wiseowise wrote:
           | LLM is much better on average just for the fact that it was
           | trained on a large corpus of human knowledge, including
           | psychology, therapy and study material. Most of the humans in
           | your vicinity only have some shallow knowledge of local cargo
           | cults and religious teachings.
        
             | staticman2 wrote:
             | By that logic a Markov chain is better on average just for
             | the fact that it was trained on a large corpus of human
             | knowledge, including psychology, therapy and study
             | material.
        
         | NortySpock wrote:
         | One mental model I have with LLMs is that they have been the
         | subject of extreme evolutionary selection forces that are
         | entirely the result of human preferences.
         | 
         | Any LLM not sufficiently likable and helpful in the first two
         | minutes was deleted or not further iterated on, or had so much
         | retraining (sorry, "backpropagation") it's not the same as it
         | started out.
         | 
         | So it's going to say whatever it "thinks" you want it to say,
         | because that's how it was "raised".
        
           | user_7832 wrote:
           | Fully agree. I wonder in the long term how this will show up.
           | Will every business/CEO do more of what he/they anyway want
           | to do, but now supported by AI/LLMs?
           | 
           | The possibilities in "dangerous" fields are a bit more
           | frightening. A general is much more likely to ask ChatGPT "Do
           | you think this war is a good idea/should I drop a bomb",
           | rather than an _actually helpful tool_ - where you might ask
           | "What are 5 hidden points on favor of/against bombing that
           | one likely has missed".
           | 
           | The more you use AI as a strict tool that can be wrong, the
           | safer. Unfortunately I'm not sure if that helps if the guy
           | bombing your city (or even your president) is using AI
           | poorly, and their decisions affect you.
        
             | tavavex wrote:
             | > Will every business/CEO do more of what he/they anyway
             | want to do, but now supported by AI/LLMs?
             | 
             | Arguably, it already worked that way. The best way to climb
             | the ranks of a 'dictatorial' organization (a repressive
             | government or an average large business) is to always say
             | yes. Adopt what the people from up above want you to use,
             | say and think. Don't question anything. Find silver linings
             | in their most deranged ideas to show your loyalty. The rich
             | and powerful that occupy the top ranks of these structures
             | often hate being challenged, even if it's irrational for
             | their well-being. Whenever you see a country or a company
             | making a massive mistake, you can often trace it to a
             | consequence of this. Humans hate being challenged and the
             | rich can insulate themselves even further from the real
             | world.
             | 
             | What's worrying me is the opposite - that this power is
             | more available now. Instead of requiring a team of people
             | and an asset cushion that lets you act irrationally, now
             | you just need to have a phone in your pocket. People get
             | addicted to LLMs because they can provide endless, varied
             | validation for just about anything. Even if someone is
             | aware of their own biases, it's not a given that they'll
             | always counteract the validation.
        
         | qsera wrote:
         | If you use LLMs in a way that the underlying assumption is that
         | it is capable of "thinking" or "caring" then you are going to
         | get burned pretty bad. Because it is an illusion and illusions
         | disappear when they have to bear real weight of reality.
         | 
         | But sadly LLMs push all the right buttons that lead humans into
         | that kind of behavior. And the marketing around LLMs works
         | overtime to reinforce that behavior.
         | 
         | But instead if you ignore all that and use LLMs as a search
         | tool, then you will get positive returns from using it.
        
         | matwood wrote:
         | > I took its recommendation. Ultimately it turned out to be the
         | wrong decision.
         | 
         | Curious if you think a single person would have helped you make
         | a better decision? Not everything works out. If a friend helped
         | me make a decision I certainly wouldn't blame them later if it
         | didn't work out. It's ultimately my call.
        
           | paulhebert wrote:
           | If a friend gave me bad advice about a major life decision I
           | would stop consulting them for future life decisions
        
             | matwood wrote:
             | Bad as in malicious or bad as in they offered/you asked for
             | their advice and it didn't work out? Because if it's the
             | later, that's an unfair burden to put on your friends or
             | anyone else. People can give great advice, genuinely want
             | to help, and it still not work out the way you wanted. If
             | you require friends to be 100% with their help, I'm not
             | sure how you have any friends left.
        
               | paulhebert wrote:
               | It's possible to give good advice that doesn't work out.
               | I have no hard feelings about that.
               | 
               | It's also possible to give bad advice wothout being
               | malicious.
               | 
               | I have some friends who don't give great life advice.
               | We're still friends, they're just not the friends I go to
               | when I need advice on a big decision.
        
         | zpeti wrote:
         | I also used it for advice on a massive personal decision, but I
         | specifically asked it to debate with me and persuade me of the
         | other side. I specifically prompted it for things I am not
         | thinking about, or ways I could be wrong.
         | 
         | It was extremely good at the other side too. You just have to
         | ask. I can imagine most people don't try this, but LLMs
         | literally just do what you ask them to. And they're extremely
         | good and weighing both sides if that's what you specifically
         | want.
         | 
         | So who's fault is it if you only ask for one side, or if the
         | LLM is too sycophantic? I'm not sure it's the LLMs fault
         | actually.
        
         | nuancebydefault wrote:
         | Weird, i am using copilot and it steers me mostly towards self
         | reflection and tries to look at things objectively. It is very
         | friendly and comes across as empathetic, to not hurt your
         | feelings, that is probably baked in to keep the conversation
         | going...
        
         | Hekkova wrote:
         | As you mention, I've found Claude is doing a better job at
         | providing push back or at least alternative recommendations to
         | choices. If you ask it directly it will provide a seemingly
         | objective opinion on your decisions and direction. The key is
         | not getting sucked into the sycophantic feedback loop. Easier
         | said than done. Always ask questions and tell it to give you an
         | assessment of why a decision may be a bad idea.
        
         | heavyset_go wrote:
         | > _Thankfully it was recoverable, but it really sobered me up
         | on LLMs. The fault is on me, to be clear, as LLMs are just a
         | tool._
         | 
         | I wouldn't be so quick to discount the fact that you were
         | essentially gaslit by an ass-kissing model that was RLHF'd into
         | maximum persuasiveness. Models aren't just neutral tools,
         | they're deliberately designed to be convincing.
         | 
         | Yes, your choices and actions are on you, but if a trillion
         | dollar company gaslit you into thinking those were good choices
         | to make, _some_ of the responsibility is theirs, too.
        
         | latexr wrote:
         | > The fault is on me, to be clear, as LLMs are just a tool.
         | 
         | This is like blaming yourself for an addiction to alcohol, junk
         | food, gambling, or something else you have been relentlessly
         | advertised to.
         | 
         | Sure, _some of it_ falls on you, but there are corporations
         | with infinite money spending most of it to manipulate your
         | psyche into wanting the thing, trusting the thing, feeling
         | empty without the thing.
         | 
         | https://www.youtube.com/watch?v=Xj4aRhHJOWU
        
           | globular-toast wrote:
           | Yeah, I used to be in the "it's your own fault, moron" school
           | of thought. But as I've grown up I've seen all the ways
           | people prey on the hopes and fears of others, and take
           | advantage of the basic animal instincts in all of us.
           | 
           | I used to think obesity was self inflicted, for example. But
           | then you notice how junk food companies are allowed to do
           | whatever they want to get people hooked on their stuff. They
           | can put up huge billboards, vending machines up a few metres
           | from where you work, they even pump their smell out into the
           | streets.
           | 
           | So let's not aim for a society where we blame victims of
           | predatory marketing and carefully engineered addictive
           | products. We all have weaknesses. Let's help each other out,
           | not take advantage.
        
           | senordevnyc wrote:
           | Genuine question: for me as an individual, what is the
           | utility of framing myself as a helpless victim rather than an
           | actor with agency and responsibility for myself?
        
             | latexr wrote:
             | > Genuine question: (...) framing myself as a helpless
             | victim
             | 
             | If you engage with the argument genuinely and steel man,
             | you'll see that is not what I said. I even emphasised it:
             | 
             | > _some of it_ falls on you
             | 
             | You're not a "helpless victim" but you're also not fully to
             | blame. Understanding that means understanding the problem
             | and being more powerful to fix it. For example, if you're
             | addicted to TV or social media, you can make a concerted
             | effort to improve your life by removing the problems at the
             | external source (sell your TV, delete an account and app
             | from your phone).
        
         | AdrianB1 wrote:
         | I think the problem is what you asked. 90% of time I ask LLMs
         | practical questions about tech, equivalent to Stack Overflow
         | questions, but I did have some discussions about some
         | situations and I asked for information and arguments, not
         | advice. It is my job to act on the information and consider
         | opinions, not the LLM's. In the end, you don't ask people on
         | Stack Overflow to tell you what to do, but you ask for info and
         | options and you decide.
         | 
         | Another problem is believing you have a high emotional
         | intelligence when there is no reliable way to quantify that -
         | similar to "I believe I am very tall, but I don't know how tall
         | I am and how tall are the others because there is no unit of
         | measure for height", with the difference that for emotional
         | intelligence there is no unit of measure and no correlation
         | that can be established with anything to make at least an
         | indirect measurement.
        
         | te_chris wrote:
         | The key is to remember what they are: bags of weights which
         | you're throwing some data into.
         | 
         | In that sense they can't offer advice because the "know"
         | nothing.
         | 
         | But they can reframe, they can reflect. They can take one idea
         | and reflect it into another intellectual framework.
         | 
         | I've used them a lot like this to help get perspective on life
         | decisions. But not for advice.
         | 
         | Try something like: I have to do x and y, give me multiple
         | psych perspectives on this problem from different schools. I
         | find this takes something abstract (your problem) and grounds
         | it in things it actually knows (the sum of ingestible human
         | knowledge).
        
       | kapral18 wrote:
       | Not AI chatbots but Claude models. Pandering and rushed thinking
       | is the bane of anthropic models. And since they are the most
       | popular ones they poison the whole ecosystem.
        
       | barnacs wrote:
       | Just a reminder: LLMs are statistical models that predict the
       | next token based on preceeding tokens. They have no feelings,
       | goals, relationships, life experience, understanding of the human
       | condition and so on. Treat them accordingly.
        
       | potatoskins wrote:
       | I read somewhere that LLMs are partly trained on reddit comments,
       | where a significant mass of these comments is just angsty
       | teenagers advocating for breakups
        
       | bethekidyouwant wrote:
       | Reddit as the source of truth...
        
       | nlawalker wrote:
       | Relevant article from The Atlantic a couple weeks ago,
       | "Friendship, On Demand":
       | https://www.theatlantic.com/family/2026/03/ai-friendship-cha...
       | (gift link)
       | 
       |  _> The way that generative AI tends to be trained, experts told
       | me, is focused on the individual user and the short term. In one-
       | on-one interactions, humans rate the AI's responses based on what
       | they prefer, and "humans are not immune to flattery," as Hansen
       | put it. But designing AI around what users find pleasing in a
       | brief interaction ignores the context many people will use it in:
       | an ongoing exchange. Long-term relationships are about more than
       | seeking just momentary pleasure--they require compromise, effort,
       | and, sometimes, telling hard truths. AI also deals with each user
       | in isolation, ignorant of the broader social web that every
       | person is a part of, which makes a friendship with it more
       | individualistic than one with a human who can converse in a group
       | with you and see you interact with others out in the world._
       | 
       | I also thought this bit was interesting, relative to the way that
       | friendship advice from Reddit and elsewhere has been trending
       | towards self-centeredness (discussed elsewhere in this thread):
       | 
       |  _> Friendship is particularly vulnerable to the alienating force
       | of hyper-individualism. It is the most voluntary relationship,
       | held together primarily by choice rather than by blood or law. So
       | as people have withdrawn from relationships in favor of time
       | alone, friendship has taken the biggest hit. The idea of
       | obligation, of sacrificing your own interests for the sake of a
       | relationship, tends to be less common in friendship than it is
       | among family or between romantic partners. The extreme ways in
       | which some people talk about friendship these days imply that you
       | should ask not what you can do for your friendship, but rather
       | what your friendship can do for you. Creators on TikTok sing the
       | praises of "low maintenance friendships." Popular advice in
       | articles, on social media, or even from therapists suggests that
       | if a friendship isn't "serving you" anymore, then you should end
       | it. "A lot of people are like I want friends, but I want them on
       | my terms," William Chopik, who runs the Close Relationships Lab
       | at Michigan State University, told me. "There is this weird
       | selfishness about some ways that people make friends."_
        
         | oldfrenchfries wrote:
         | The link is not working, but I found it myself. Great point,
         | thanks for sharing.
        
       | me551ah wrote:
       | Makes me wonder if the Iran war was a result of the same.
        
       | hax0ron3 wrote:
       | For what it's worth, that wasn't my experience at all the last
       | time I consulted ChatGPT for relationship advice. It was
       | supportive, but in an honest tough love way.
        
       | thesis wrote:
       | Humans do this too though. I have close friends that ask for
       | advice. Sometimes if I know there's risk in touchy subjects I
       | will preface with "do you want my actual advice, or just looking
       | for a sounding board"
       | 
       | I've seen firsthand people have lost friends over honesty and
       | telling them something they don't want to hear.
       | 
       | It's sad really. I don't want friends that just smile to my face
       | and are "yes-men" either.
        
         | intended wrote:
         | The difference is that SOME humans do this. As you mentioned,
         | people have lost relationships over telling others what they
         | didn't want to hear.
         | 
         | Conflating this with how LLM chatbots behave is an incorrect
         | equivalence, or a badly framed one.
        
       | jwilliams wrote:
       | For me the framing is critical - what is the model saying yes to?
       | You can present the same prompt with very different
       | interpretations (talk me into this versus talk me out of it). The
       | problem is people enter with a single bias and the AI can only
       | amplify that.
       | 
       | In coding I'll do what I call a Battleship Prompt - simply just
       | prompt 3 or more time with the same core prompt but strong
       | framing (eg I need this done quickly versus come up with the most
       | comprehensive solution). That's really helped me learn and dial
       | in how to get the right output.
        
       | intended wrote:
       | Anecdote:
       | 
       | I used to use LLMs for alternate perspectives on personal
       | situations, and for insights on my emotions and thoughts.
       | 
       | I had no qualms, since I could easily disregard the obviously
       | sycophantic output, and focus on the useful perspective.
       | 
       | This stopped one day, till I got a really eerie piece of output.
       | I realized I couldn't tell if the output was actually self
       | affirming, or simply what I wanted to hear.
       | 
       | That moment, seeing something innocuous but somehow still beyond
       | my ability to gauge as helpful or harmful is going to stick me
       | with for a while.
        
       | trimbo wrote:
       | > They also included 2,000 prompts based on posts from the Reddit
       | community r/AmITheAsshole, where the consensus of Redditors was
       | that the poster was indeed in the wrong.
       | 
       | Sorry, anonymous people on reddit aren't a good comparison. This
       | needs to be studied against people in real life who have a social
       | contract of some sort, because that's what the LLM is imitating,
       | and that's who most people would go to otherwise.
       | 
       | Obviously subservient people default to being yes-men because of
       | the power structure. No one wants to question the boss too
       | strongly.
       | 
       | Or how about the example of a close friend in a relationship or
       | making a career choice that's terrible for them? It can be very
       | hard to tell a friend something like this, even when asked
       | directly if it is a bad choice. Potentially sacrificing the
       | friendship might not seem worth trying to change their mind.
       | 
       | IME, LLMs will shoot holes in your ideas and it will efficiently
       | do so. All you need to do ask it directly. I have little doubt
       | that it outperforms most people with some sort of friendship,
       | relationship or employment structure asked the same question. It
       | would be nice to see that studied, not against reddit commenters
       | who already self-selected into answering "AITA".
        
         | maximinus_thrax wrote:
         | Not only that, but subreddits like r/AmITheAsshole are full of
         | AI slop. Both in the comments and in the posts. It's a huge
         | karma mining operation for bots.
        
           | genidoi wrote:
           | That can be solved by filtering out any posts made after
           | November 2022.
        
             | expedition32 wrote:
             | That's not a good solution. We don't use medical textbooks
             | from 20 years go.
             | 
             | Strangers from the internet, bot or otherwise, are not your
             | mental coach.
        
             | bombcar wrote:
             | Even before the advent of AI reddit was notorious for
             | obvious bullshit being posted for karma farming. r/aita is
             | even _more_ famous for people making up stories for unknown
             | and known purposes (known in the old days as  "bait").
        
           | thwarted wrote:
           | The upvotes ultimately train the bots, reenforcing the
           | content posted. Even the most passive form of interaction has
           | been co-opted for AI.
        
           | z3c0 wrote:
           | Plus, there's the disproportionate ratio of
           | posters:commenters:lurkers. The tendency to comment over
           | keeping ones thoughts to themself is a selection bias
           | inofitself.
        
             | maximinus_thrax wrote:
             | Great insight, didn't thing about it even anecdotally. I
             | was lurking on Reddit since 2008 and finally created an
             | account in 2012 when someone was really 'wrong on the
             | internet' and had to step in.
        
           | mikeocool wrote:
           | This is sort of funny. Given how common it is to spot bots on
           | Reddit now, it seems like they are likely to completely
           | overwhelm the site and drive away most of actual humans.
           | 
           | At which point the bots, with all of their karma will be
           | basically worthless.
           | 
           | Kind of extra funny/sad that Reddit's primary source of
           | income in the past few years appears to be selling training
           | data to AI labs, to train the Models that are powering the
           | bots.
        
             | RealityVoid wrote:
             | > At which point the bots, with all of their karma will be
             | basically worthless.
             | 
             | Not really, it will still be kind of valuable for influence
             | campaigns, a lot of people don't get it when there is a bit
             | in the other side. Hell, a lot of times, I don't get it.
        
               | mikeocool wrote:
               | I know a fair number of people "normies" who get some
               | value out of smaller niche Reddit communities -- for
               | advice, and things like product recommendations.
               | 
               | If suddenly all the posts are coming from bots who are
               | trying push a product or just farm karma, I assume
               | (perhaps naively) that those folks will get a lot less
               | value, and stop showing up -- even if they don't realize
               | it's bots on the other side of the conversation.
        
             | mlrtime wrote:
             | How do you clearly define a bot?
        
         | alberto467 wrote:
         | "AI is nicer than the average redditor" would be a more
         | accurate title
        
           | mattmanser wrote:
           | I would say people on /r/amitheasshole are more biased
           | towards the poster, i.e. nicer.
           | 
           | There's plenty of those I've read where I thought it sounded
           | like the poster was the asshole and the top replies were NTA.
        
             | jjmarr wrote:
             | r/AmItheAsshole is biased towards breaking off
             | relationships rather than fixing them. They also hate
             | social obligations.
             | 
             | e.g. If the OP is asking "I ghosted my friend in AA who
             | insulted me during a relapse", Reddit would say NTA in a
             | heartbeat, while the real world would tell OP to be more
             | forgiving.
             | 
             | On the contrary, if the post was "the other kids at school
             | refuse to play with my child", Reddit would say YTA because
             | the child must've done something to incite being cut off.
        
               | ericd wrote:
               | Absolutely. I wonder how many parents have been no
               | contacted, SOs broken off with, friendships broken
               | because of the Reddit hivemind's attitude. Pretty sure
               | it's doing a huge amount of societal damage.
        
               | jjmarr wrote:
               | I wouldn't blame reddit, it's what you get when you ask
               | several thousand teenagers to give collective
               | relationship advice.
        
               | tbossanova wrote:
               | "I got divorced based on advice from complete strangers
               | on the internet, AITA?"
        
               | wiseowise wrote:
               | Is it hivemind or just people being generally aware
               | better of toxicity in their lives?
        
               | wiseowise wrote:
               | > e.g. If the OP is asking "I ghosted my friend in AA who
               | insulted me during a relapse", Reddit would say NTA in a
               | heartbeat, while the real world would tell OP to be more
               | forgiving.
               | 
               | That's a nuanced discussion. It depends on what you value
               | most, not what "real world" tells you. Most of the time
               | Reddit would be right, because you need to prioritize
               | yourself instead of continuing toxic relationships.
        
               | mlrtime wrote:
               | 1) Reddit is horrible at nuance, almost non existent in
               | some subs.
               | 
               | 2) The toxicity is being defined by reddit to give the
               | advice which is mostly wrong as outlined above.
               | 
               | If OPs had a understanding of what they valued and what
               | is toxic, they probably wouldn't need a advice from
               | biased readers [biased in the sense that they're on that
               | sub].
        
               | wiseowise wrote:
               | That's true, but they still might be right for wrong
               | reasons.
        
             | rurp wrote:
             | Yeah every single time I click on one of those posts the
             | top comments are NTA. A couple times I tried randomly
             | opening a few dozen posts and checking the top comments to
             | see if I could find a single YTA and struck out.
             | 
             | Granted many of the OPs are very biased in the poster's
             | favor. Most I've read fall into one of two buckets: either
             | they want to gripe about some obviously bad behavior, or
             | it's a controved and likely fake story.
        
               | mlrtime wrote:
               | The problem with any of these is that they are so
               | incredibly biased towards the author's frame of reality
               | (understandably so).
               | 
               | Who among us are able to 1) Understand a 2nd persons view
               | of a issue we're in and 2) have the ability/courage to
               | write it in a post seeking advice.
               | 
               | My point is that the author will specifically frame the
               | problem clearly on their side. Occasionally redditors
               | will seek additional questions but rarely.
        
             | yieldcrv wrote:
             | It's gendered, by the way
             | 
             | Many of the posts are A/B tests of a prior post where only
             | the genders were flipped of the OP and antagonist to see
             | how the consensus also flips
        
           | 52-6F-62 wrote:
           | Pretty sure the average Redditor _is_ AI now.
        
             | lotsofpulp wrote:
             | How the hell is a study on stanford.edu assuming posts on
             | Reddit are genuine? That should be enough to get you kicked
             | out of Stanford.
        
               | helpfulclippy wrote:
               | Though interestingly, the observed difference in
               | assessment suggests (though does not prove) that sampled
               | AITA posters are not one of these models. I guess it's
               | possible they have a very different prompt though...
        
           | yard2010 wrote:
           | IMHO it's not about being nice. AITA threads show an
           | interesting phenomenon of social consensus, I think the
           | authors wanted to show that the LLMs they checked don't have
           | that.
        
             | ianbutler wrote:
             | I don't think Reddit is a great place to determine social
             | consensus for well adjusted people or representative of the
             | average adult view. I never see people on Reddit have
             | opinions of any the people I consider reasonable in real
             | life and I don't mean politics I wouldn't know, I don't
             | frequent political subreddits.
             | 
             | It seems fairly consistently miserable in any of the common
             | high traffic subs and you have to get down to really niche
             | communities to see what I consider reasonable behavior that
             | matches the behavior of people I know in real life.
        
               | strken wrote:
               | The AITA social consensus is a specific kind of
               | groupthink which differs from nearly everyone I know in
               | real life. I assumed yard2010 meant the specific AITA
               | social consensus and not general human agreement.
               | 
               | Even the premise of deciding who's right and who's wrong
               | is miserable. Most problems are like those daisy-chains
               | of padlocks you see on gates in remote areas[0]: there
               | are multiple factors that caused the problem, and
               | removing any factor would remove the problem too.
               | 
               | [0]
               | https://www.flickr.com/photos/72793939@N00/51117212748
        
               | bear141 wrote:
               | I don't think I've read a Reddit thread in the last few
               | years that didn't devolve into politics on the highest
               | upvoted comments fairly quickly.
        
           | brikym wrote:
           | Is it the _average_ redditor? The most upvoted would be even
           | worse.
        
         | zer00eyz wrote:
         | > This needs to be studied against people in real life who have
         | a social contract of some sort... IME, LLMs will shoot holes in
         | your ideas and it will efficiently do so.
         | 
         | The Krafton / Subnatuica 2 lawsuit paints a very different
         | picture. Because "ignored legal advice" and "followed the LLM"
         | was a choice. Do you think someone who has conversation where
         | "conviction" and "feelings" are the arbiters of choice are
         | going to buy into the LLM push back, or push it to give a
         | contrived outcome?
         | 
         | The LLM lacks will, it's more or less a debate team member and
         | can be pushed into arguing any stance you want it to take.
        
         | 4ndrewl wrote:
         | What's your research background in this area?
        
           | jiggunjer wrote:
           | How is that relevant. A decent scientist can critique general
           | design aspects of a paper in any field. They're hardly
           | splitting hairs on some niche topic.
        
             | 4ndrewl wrote:
             | Apologies, I didn't realise they were a decent scientist.
        
         | legacynl wrote:
         | > Sorry, anonymous people on reddit aren't a good comparison.
         | 
         | Yeah especially on r/AmITheAsshole. Those comments never
         | advocate for communication, forgiveness and mending things with
         | family.
        
           | SJMG wrote:
           | Yes, it is a toxic sub, where the notion that there can be
           | greater happiness on the other side of forgiveness than
           | cutting ties is all but absent.
        
             | JumpCrisscross wrote:
             | To be fair, it's easier to concisely explain cutting
             | someone off than justifying forgiveness. And the latter
             | will land with some people versus others, while the former
             | will only be rejected by people who have themselves
             | concluded a theory of forgiveness. As a result, the simpler
             | pitch gets upvoted. Even if the majority would have been
             | swayed by a collection of arguments the other way.
        
               | theoreticalmal wrote:
               | It's a good theory. My theory is, for whatever reason,
               | jaded, narcissistic, miserable people congregate in
               | r/AITA and try to drag other people into their misery
               | because that's easier than accepting responsibility and
               | doing something to change.
        
               | BoorishBears wrote:
               | Before Reddit made hiding profiles easy you'd click on a
               | user's unreasonably scorched earth advice to the OP, and
               | find their post history is essentially going to every
               | story they come across and advocating for scorched earth.
        
               | daveguy wrote:
               | What are the chances you were seeing the anti-civ bots
               | and now reddit makes them easier to hide? (And I'm not
               | saying regular people acting like bots, but an anti-civ
               | campaign.)
        
               | nwallin wrote:
               | Hiding profiles has genuinely made the platform
               | profoundly worse. It's impossible to tell if you've just
               | got a troll on your hands or someone who's making a good
               | faith argument. It used to be enough to check their
               | profile, and either downvote and move on, or engage with
               | someone on a human level.
               | 
               | Now everyone is a troll/bot by default unless proven
               | otherwise.
        
             | wiseowise wrote:
             | Except it's not toxic to suggest that cutting toxic
             | relationship out yields greater happiness.
        
               | prepend wrote:
               | Well, maybe.
               | 
               | The challenge is interpreting what is toxic, correctly.
               | 
               | Also, if everyone I know is "toxic" then that's a good
               | sign that the problem is me and not everyone else.
        
               | wiseowise wrote:
               | > The challenge is interpreting what is toxic, correctly.
               | 
               | Correct. It is always case by case review.
               | 
               | > Also, if everyone I know is "toxic" then that's a good
               | sign that the problem is me and not everyone else.
               | 
               | Why "everyone"? Generalizations like these are the same
               | mistake that Reddit, that you're calling out, makes.
               | 
               | Also, toxic is relative to your perspective - it's not a
               | universal merit.
        
           | Iulioh wrote:
           | It's often that a lot of "NTA" answers are downright
           | antisocial.
           | 
           | "No one owns you anything, you don't own anyone anything"
           | mentality, without a crumb of social awareness.
        
           | brikym wrote:
           | I believe this. There is a graph somewhere of the
           | relationship subs tending towards breaking up over time.
        
             | tdb7893 wrote:
             | I don't think this is necessarily that the advice is
             | getting worse. My friends are pretty mature and stable
             | people and I've found that they've had way more issues
             | staying in relationships longer than they should've
             | compared to breaking up earlier. Especially for
             | relationships earlier in people's lives (where many people
             | I know has a story about being in a relationship for way
             | longer than they should've and seems often to be the ages
             | of people asking for advice) erring towards breaking up
             | seems prudent.
             | 
             | Not that these relationships subreddits are good (often
             | it's obviously children trying to give advice they don't
             | have the experience for) but I don't think that telling
             | people to break up more is less accurate advice.
        
               | sarchertech wrote:
               | > I've found that they've had way more issues staying in
               | relationships longer than they should've compared to
               | breaking up earlier
               | 
               | Consider that if ending a relationship causes noticeable
               | problems to external observers, it's almost by definition
               | because you were in it "too long". That is you developed
               | a strong attachment, shared assets, or had kids with what
               | was in hindsight obviously the wrong person.
               | 
               | Essentially you can know which relationships a person
               | stayed in too long, but you can't know how things would
               | have worked out in relationships people ended too early.
               | 
               | Also it's probably good advice to tell a 19 year old to
               | break up with her boyfriend over a half dozen serious red
               | flag issues, but that's not the only kind of thing Reddit
               | relationship advice is generally dealing with. It's not
               | even the majority. If you're advice is always to beak up
               | over every petty difference or minor slight, you might
               | reduce the number of people who stay in bad
               | relationships, but your advice, if taken, would make good
               | long term relationships impossible.
        
               | dTal wrote:
               | >Consider that if ending a relationship causes noticeable
               | problems to external observers, it's almost by definition
               | because you were in it "too long". That is you developed
               | a strong attachment, shared assets, or had kids with what
               | was in hindsight obviously the wrong person.
               | 
               | Reducing it to "right person / wrong person" is a very
               | narrow viewpoint. People can change in unpredictable
               | ways, including yourself. Relationships end - or continue
               | - for so many reasons, both emotional and pragmatic. It's
               | simply too reductive to say that if a relationship causes
               | pain when it ends, there was necessarily some sort of
               | mistake. It could even be that the pain is a price to pay
               | for a life experience that you'd be worse off for not
               | having...
        
               | goku12 wrote:
               | > I don't think this is necessarily that the advice is
               | getting worse.
               | 
               | > but I don't think that telling people to break up more
               | is less accurate advice.
               | 
               | Those are subjective determinations based on personal
               | experience. But breaking up more without addressing the
               | underlying issues is likely to cause steadily worsening
               | problems at both individual and societal scales. I'm not
               | a mental health professional, but I can see several
               | problems with this approach.
               | 
               | The first is that the determination of the issue is
               | really tricky and needs careful work. The partner who
               | seems abusive may not always be the actual perpetrator.
               | They may be displaying stress response to hidden and
               | chronic abuse by the other partner. For example, a short
               | temper may be caused by anxiety about being emotionally
               | abused. Such manipulative discrediting of the victim may
               | even be a habitual behavior rather than a deliberate one.
               | And it's more common than you'd imagine. When you support
               | the second partner based on a flawed judgment, you're
               | reaffirming their toxic behavior, while worsening the
               | self image of the victim that has already been damaged by
               | gaslighting.
               | 
               | Another issue is the degrading empathy. All
               | relationships, even business deals, are based on
               | sacrifices and compromises meant to bring you benefits in
               | the long term. Stable long term romantic/marital
               | relationships have benefits that far outweigh the
               | sacrifices one usually has to make. But the evolving
               | public discourse, especially those on r/AITA, is more in
               | favor of ruining the relationship rather than make any
               | sacrifices at all. In response, relationships are
               | becoming loveless, transactional and so flaky that any
               | compromise is seen as oppression by the partner. There is
               | zero self reflection and very few advises to examine
               | one's own behavior first. It's all about oneself and the
               | problem is always on the other side!
               | 
               | And unsurprisingly, these negative tendencies are
               | bleeding into their social lives as well. Over the past
               | decade or so, I have observed a marked increase in
               | unsympathetic and somewhat radicalized discourse. Amateur
               | advice is very harmful and this is definitely a massive
               | case for the professionals to manage. But they're also
               | products of the same system (with exceptions, of course).
               | So I'm going to criticize even the professional and
               | academic community in this matter. In their drive towards
               | hyper-individualism, many seem to have forgetten that
               | humans are social beings who won't fare well physically
               | or emotionally without relations, relationships and
               | society.
        
               | somenameforme wrote:
               | The US (and developed world more generally) is full of
               | people living alone, suffering from loneliness, and
               | increasingly trending towards widescale mental and
               | psychological illness. This has correlated quite strongly
               | with the trend going from 'just stick with it' and having
               | large families to 'mature and stable' people still being
               | in a dating phase, childless, in what I assume is a
               | relatively late stage in life.
               | 
               | At some point I think it helps to take a look at the
               | macro, because it's so easy to get lost in the micro. And
               | it often reveals the micro, in many domains, to be simply
               | absurd.
        
               | tdb7893 wrote:
               | The people I know not in good and long term relationships
               | now are the ones that stayed in bad ones too long in
               | their 20s and 30s. Staying in bad relationships seems to
               | be what has people in the "dating phase" later in life.
               | Trying to make bad relationships work had people I know
               | miserable for a decade and then dating again in their 40s
               | when the relationship inevitably failed.
               | 
               | Especially when you consider that the set of people
               | asking _Reddit_ of all places for dating advice are
               | probably young and in bad situations (it seems like
               | people in abusive relationships often ask the internet
               | for advice because part of abuse is separating them from
               | their loved ones in real life), then  "stick with it"
               | seems like the riskier statrgy generally.
        
               | somenameforme wrote:
               | Nothing is inevitable. I think people are often looking
               | for something that they're not going to find anywhere,
               | which is a very poor state for living a contended life.
               | This is certainly amplified by the nature of social media
               | where people get mistaken realities of positive
               | relationships. Great relationships on the outside often
               | have endless issues on the inside, that they work
               | through, that people on the outside aren't going to be
               | aware of.
               | 
               | Because an important part of keeping a relationship
               | healthy is not airing your dirty laundry. It's almost
               | like these endless hokey folksy sayings were built up
               | over millennia of wisdom that kept society moving along
               | in a great and healthy direction. And now that we've
               | decided to rethink everything, we have societies that
               | are, at the minimum, no longer self sustaining.
        
           | LinXitoW wrote:
           | Well, because that's never the correct choice. There's a big
           | big filter on people actually posting there. Any easy
           | problems with obvious solutions never make it to there.
           | 
           | Think about it, how fucked does your relationship have to be
           | to post on Reddit for advice?
        
             | BoorishBears wrote:
             | This wrongly assumes people are good at judging what easy
             | problems are.
             | 
             | Not to mention nowadays an untold amount of posts to
             | subreddits that invite commentary are made up stories from
             | accounts trying to get engagement.
        
             | Robotbeat wrote:
             | Someone has a chart somewhere that shows responses in that
             | subreddit getting more and more anti-conciliatory over
             | time. I think it's online misanthropy (measured by Reddit
             | responses) increasing over time rather than it being
             | objectively never the correct choice.
        
               | ijk wrote:
               | Also the rules and norms of the subreddit has changed
               | over time, which has led to spin-off subreddits that
               | serve those purposes.
        
             | redanddead wrote:
             | when people post there it's for the self justification
        
           | everdrive wrote:
           | Additionally, I'm sure many posts and replies on
           | r/AmITheAsshole are LLM-generated in the first place.
        
             | echelon wrote:
             | Before LLMs, it was a frequent haunt of fiction writers.
        
             | buu700 wrote:
             | reddit in 2026 is the ghost of pandemic-era humanity.
        
           | kshacker wrote:
           | Oh man, I have 8 reddit accounts (AFAIK) one for each purpose
           | so that I am not branded based on my open comments. Anyways,
           | one of them is abandoned because ... that's where I got
           | started at reddit about 7-10 years back. Got hooked actually
           | to the relationship subs. Very addictive to start with. Then
           | I tried to play the "Indian family values" where I would
           | advocate communication and compromise for small matters, of
           | course I recommended "get a lawyer, divorce" once in a while,
           | but more often than not, I would advocate reconciliation and
           | provide practical solutions for that. And wow ... the amount
           | of downvotes and pushback I will get on those. I just stopped
           | using that account at one point because what is the point of
           | discussions when either my values are totally out of sync
           | with the mob, or the mob does not want to listen to me. Now I
           | just read the best of redditor updates for vicious pleasure.
        
             | pocksuppet wrote:
             | That's amazing you have more than one account, aren't a
             | power mod, and haven't been IP banned yet.
        
               | kshacker wrote:
               | Nope. Started my first maybe 8-10 years back, and then
               | added the others over a year or 2. None since. I do not
               | use them all nowadays, but I was very active in my early
               | reddit days.
               | 
               | Since someone downvoted my parent comment, I am not
               | hiding anything, this is just being safe in the modern
               | world, and here are the 8 alts:
               | 
               | 1. This same name - bay area / tech
               | 
               | 2. entertainment - least used, but it becomes useful when
               | i am watching something live. It was my place to be
               | during game of thrones last season (and sadly so)
               | 
               | 3. indian left politics + bollywood - pretty much unused.
               | 
               | 4. indian right politics + bollywood. i got banned from
               | one sub for an innocent comment, so i decided to just
               | form personas. and maybe that's when i created health /
               | finance / bay area accounts -- but memory fades after a
               | long time. pretty much unused.
               | 
               | 5. relationship advice - unused for a long time. it does
               | not exist on my main phone, but i have all of them on my
               | work phone so i know it exists
               | 
               | 6. american politics. i do not participate much nowadays,
               | with age my brain has dulled and it needs to shed load so
               | this is used minimally, but at a point i was so active
               | that my karma pulled me into the sweet reddit IPO. I kept
               | only 100 shares btw
               | 
               | 7. health - only health topics, also unused, but i go
               | there and use that account when i need to read on a
               | specific topic
               | 
               | 8. finance - only investment, trading
               | 
               | nowadays you can hide reddit history, but earlier you
               | could not, and my point is i do not want to 1) delete my
               | comments, but 2) be hounded by them when i have a
               | question about a different topic. but i did not care if
               | people read my past 100 comments about politics when i
               | talk about politics.
               | 
               | so i flip between 2-3 accounts on a daily basis, and
               | maybe 4-5 in a good week. i have not been challenged by
               | reddit, but if they do, i will adapt. Switching between
               | them was much easier earlier in the Apollo days and even
               | at reddit - they have made navigation worse for this
               | specific use case.
        
               | leosanchez wrote:
               | > indian right politics + bollywood. i got banned from
               | one sub for an innocent comment
               | 
               | You were banned from a Indian re subreddit or banned
               | because being rw ?
               | 
               | FYI: I was banned from r/india for commenting basic info
               | on how economy works.
        
               | kshacker wrote:
               | I do not recall, it is long time back, I looked and could
               | not find the ban notice or the specific comment that may
               | have been the issue. But I was banned from /india - same
               | as you. And I think it was barely political. I do not
               | discuss politics much on India but once in a while a
               | comment slips, or needs to slip. And when it needs to
               | slip, I used to know how to lean ... but like the dirty
               | harry movie ... at this point I have forgotten which one
               | is which, so it is more a question of am I feeling lucky
               | to comment about a hot topic.
        
               | leptons wrote:
               | You can't use IP address to ban someone without
               | significant abuse. All home network routers put everyone
               | in the house behind the same IP address. For all reddit
               | knows, there are 8 people in the house using reddit.
        
               | ThunderSizzle wrote:
               | Want a new IP address? Reset your router or cycle it.
               | Typically it'll procure a new IP address from the ISP.
               | 
               | I guess that makes IP banning residential nodes even more
               | stupid.
        
               | navigate8310 wrote:
               | CGNAT is a benefit in disguise
        
               | chaosite wrote:
               | Having more than one account isn't against Reddit's ToS.
               | 
               | If you use your different accounts in different
               | subreddits and never have your accounts interact, you
               | won't be banned.
        
               | throwaway27448 wrote:
               | Sure but the background chances of an account getting
               | banned for clashing with a mod is quite high.
        
               | user34283 wrote:
               | If you don't restrict each account to specific
               | subreddits, it's quite likely that one will get banned
               | somewhere without you noticing or remembering.
               | 
               | If you happen to post to the same subreddit with another
               | account at some point, Reddit bans all of your accounts.
        
               | AdvancedCarrot wrote:
               | Anecdotal but I've noticed Reddit has gotten very ban
               | happy in general in the past year.
               | 
               | I actually gave up using it because, perhaps in part
               | because I'm behind a VPN (required in my country), any
               | new accounts I create get banned very quickly once I
               | start commenting.
        
               | staticman2 wrote:
               | I've definitely posted to the same subreddit with two
               | different accounts by accident without being banned.
               | 
               | The android reddit app annoyingly doesn't check for
               | account matches. If you click a browser notification link
               | on Account A it can open a reply form on App account B.
        
               | user34283 wrote:
               | I meant if one of the accounts is already banned there,
               | it counts as ban evasion and Reddit bans all of your
               | accounts.
               | 
               | This might easily happen if you like to participate in
               | political discussions.
        
           | curiousgal wrote:
           | We're missing the other obvious problem, most of the content
           | there is AI generated anyway. I personally posted a fake
           | story generated by Chatgpt and even posted screenshots of
           | that at the start of the post and yet, the post ended up on
           | the frontpage...
        
           | throwaway27448 wrote:
           | It doesn't help that the actual submissions are difficult to
           | distinguish from creative writing exercises.
        
           | insane_dreamer wrote:
           | That sub is so toxic that I would seriously question the
           | wisdom of any of its posts simply because the authors are
           | member of that sub.
        
         | salawat wrote:
         | >Obviously subservient people default to being yes-men because
         | of the power structure. No one wants to question the boss too
         | strongly.
         | 
         | This drives me nuts as a leader. There are times where yes,
         | please just listen, and if this is one of those times, I'll
         | likely tell you, but goddamnit, speak up. If for no other
         | reason I might not have thought of what you've got to say. Then
         | again, I also understand most boss types aren't like me, thus
         | everyone ends up conditioned to not bloody collaborate by the
         | time they get to me. It's a bad sitch all the way around.
        
           | CoffeeOnWrite wrote:
           | Indeed. I directly ask my reports to discover and surface
           | conflicts, especially disagreements with me, and when they do
           | I try to strongly reinforce the behavior by commending and
           | rewarding them. Could anyone recommend additional resources
           | on this topic?
        
             | matwood wrote:
             | Simon Sinek has a lot of good content around this. Step one
             | is building trust. People won't speak up if they don't feel
             | safe doing so.
        
         | dwaltrip wrote:
         | Are you saying there isn't an actual sycophancy problem?
         | 
         | We are talking about overall patterns here, not the experience
         | of a small subset of skilled and careful users.
        
         | erikerikson wrote:
         | Doesn't sound like a close friend to me. If I tell them what I
         | really think they may not be a friend? Close may not mean what
         | you think it means.
         | 
         | The challenge is that these social choices have a strong
         | stratification effect and those of us who can transit the
         | cultures are statistically rare.
        
         | skybrian wrote:
         | You could think of what they did in the first study as
         | constructing an exam to test how well various LLM's do as an
         | advice columnist. They wanted a lot of personal advice
         | questions where the LLM should _not_ affirm by default. If a
         | few questions with wrong answers got in there, it probably
         | wouldn 't affect the results all that much?
         | 
         | Unfortunately they didn't test anything newer than GPT4o, so we
         | don't know how much GPT-5 improved. It would be nice if someone
         | turn their list of questions into a benchmark.
        
           | n_bhavikatti wrote:
           | They actually did test GPT-5:
           | https://www.science.org/doi/10.1126/science.aec8352 (see the
           | figure under Conclusion). Its rate of endorsement of user
           | action, 52%, was the same as GPT-4o. So based on their setup
           | it seems that the newer model didn't reduce affirmation.
        
         | redanddead wrote:
         | Reddit is notorious for being awful at real life interactions
         | 
         | just look at the relationship subreddit the first answer is
         | always divorce, it's become a meme
         | 
         | but beyond romantic relationships, i think a lot of us have
         | seen how it can impact work relationships, i've had venture
         | partners clearly rely on AI (robotic email responses and even
         | SMS) and that warped their perception and made it harder to
         | connect. It signals laziness and a lack of emotional
         | intelligence
         | 
         | AI should enhance and enable connection, not promote isolation,
         | imo this is a real problem
         | 
         | it should spark curiosity, create openings for conversations,
         | point out the biases to make us better at connecting with other
         | people, i hope we get to a point where most people are made
         | kinder by ai. I'm seeing the opposite atm, interested in
         | hearing others experiences with this
        
           | kelvinjps10 wrote:
           | I always find it interesting how, in Reddit any trivial fight
           | or even just different opinions, the advice it's always to
           | end the relationship.
        
             | abpavel wrote:
             | You deserve better
        
               | yabutlivnWoods wrote:
               | A very sycophantic AI style answer.
               | 
               | Code bot equivalent being all "you are absolutely right!
               | Here is the unequivocal fix for now and all time!"
        
               | finghin wrote:
               | No one would ever have made that comment before GPT. I
               | have a feeling you could be the one that's poisoned!
        
             | PhoenixFlame101 wrote:
             | The power of an echo chamber; makes extremism seem logical.
        
               | MaxBarraclough wrote:
               | I think it's more than just an echo chamber, it's the
               | community wanting drama more than it wants to help
               | people.
        
             | roncesvalles wrote:
             | I mean... it's a solution guaranteed to work in a trivial
             | sense. It's not meant to be a serious suggestion but more
             | of a thought experiment, like "hold this as the bar, can
             | you find a solution better than this?"
             | 
             | It's like what GiveDirectly says: all charitable
             | interventions should be benchmarked against simply giving
             | the beneficiaries a wad of cash.
        
             | fortmeier wrote:
             | I think it is some kind of survivership bias. Who is going
             | to give advise on reddit? Maybe people shying away from
             | difficult social interactions?
        
               | somenameforme wrote:
               | Paired with the echo chamber effect voting systems
               | create. Anything that affirms the biases of a majority of
               | upvoters gets elevated, anything that contradicts it gets
               | hidden, and so you not infrequently end up with
               | ubiquitous nonsense that then further reinforces the echo
               | chamber as they become self assured. Then real life
               | intervenes, completely goes against the online zeitgeist,
               | and they're all confused.
        
               | DeathArrow wrote:
               | If you want to get poor fast, you can follow the most up
               | voted advices on r/wallstreetbets.
        
               | ahartmetz wrote:
               | Would be interesting to have data on that. If it was
               | true, you could win by always doing the opposite!
        
               | DeathArrow wrote:
               | Just because one action is demonstrably harmful does not
               | mean its negation is automatically beneficial.
               | 
               | Or, formally, my claim is A implies B. The only logical
               | contrapositive is non B implies non A. (not losing money
               | means not following advices on r/wallstreetbets)
               | 
               | But you say: non A implies non B, which is the fallacy of
               | denying the antecedent.
        
               | diydsp wrote:
               | It doesnt have to be universally true to be true in a
               | mathematical system like options/puts/calls.
        
               | tyre wrote:
               | Not necessarily. WSB users are trying to make it big,
               | which means betting on long shots. This could be penny
               | stocks, companies on the verge of bankruptcy, or ones
               | with more sentimental value than fundamentals.
               | 
               | Betting against these companies is obvious and expected,
               | so the cost of shorting might be high enough that even if
               | you're correct (stock goes down, the opposition of what
               | WSB said), paying the cost of the short (the fee to
               | borrow the stock from someone else) is high enough that
               | you still lose money.
               | 
               | Also:
               | 
               | 1. shorting stocks can be quite dangerous. Your downside
               | is, well, not infinite but it can easily wipe you out.
               | 
               | 2. You might be correct that the stock goes down, but
               | over what time frame? Again, you have to pay money to
               | hold a short. Or you're using a different financial
               | instrument that has a specific timeline. If the market
               | does move in your direction but too late, you still lose.
        
               | notahacker wrote:
               | Those particular subreddits are heavily populated by
               | incels voraciously consuming the stories of relationship
               | strife (real, distorted and purely fictional) to validate
               | their belief that their relationship status is down to
               | the evils of the opposite sex specifically and all
               | relationships being doomed in general.
               | 
               | That's as big a bias as AI affirmation bias; indeed AI
               | and certain corners of Reddit are probably the _only_ two
               | venues likely to provide this sort of affirmative
               | response https://alexyeozhenkai.substack.com/p/i-cheated-
               | on-my-wife-b...
        
             | gunsle wrote:
             | That would be because nearly all of those posts are
             | entirely made up and somehow I guess people can't tell
             | that?
        
             | 9rx wrote:
             | Stands to reason. Ask a computer for advice and it is going
             | to give you a computer-centric answer: Restart and try
             | again.
        
             | ModernMech wrote:
             | > the advice it's always to end the relationship.
             | 
             | To be fair, if your interpersonal skills and relationship
             | dynamic are such that you find yourself seriously asking
             | the Internet (Reddit of all places) for relationship
             | advice... yeah, just end it is probably the null
             | hypothesis.
        
           | triage8004 wrote:
           | Every answer on Reddit feels like: divorce them, dump them,
           | cut them out of your life, or similar.
        
           | finghin wrote:
           | Yes, in principle this would be a great way to get a grip on
           | AI personal-decision-making. But there's a nontrivial chance
           | Claude is more emotionally intelligent than r/AITA. That is
           | not something I enjoy saying.
        
             | intended wrote:
             | Eh, AITA works very well for the more common and obvious
             | situations.
             | 
             | I wonder how MUCH better Claude really is when compared to
             | AITA. Also people are mixing up relationship advice with
             | AITA.
        
           | blablabla123 wrote:
           | I think this may be selection bias. People asking anonymously
           | (edit: for relationship advice) on Reddit perhaps even with a
           | throwaway account are likely in a desperate situation. So
           | hardly to be compared with the _average_ real life situation.
           | Thus 1. chances are running is a good option and 2. also
           | considering even in 2026 AI still essentially is a
           | statistical machine that doesn't handle corner-cases at the
           | tails well.
           | 
           | Anecdotally as I've thoroughly worked and used AI myself. It
           | performs best with google-able stuff that is needle-in-the-
           | haystick like and worst with personal and work advice. The
           | main problem I see is that it's tempting to use it for that.
        
             | trueno wrote:
             | > worst with personal and work advice. The main problem I
             | see is that it's tempting to use it for that.
             | 
             | i think i want to expand on this even more. even people ive
             | worked with for years that ive looked up to as brilliant
             | people are starting to use it to conjure up organizational
             | ideas and stuff. they're convinced, on the backs of their
             | hard earned successes, that they're never going to be
             | fallible to the pitfalls of... idk what to call it. AI
             | sycophancy? idk. i guess to add to this, i'm just not sure
             | AI should be referenced when it has anything to do with
             | people. code? sure. people? idk. people are hard, all the
             | internet and books claude or whatever ai is trained on
             | simply doesnt encapsulate the many shades of gray that
             | constitute a human and the absolute depth/breadth of any
             | given human situation. there's just so many variables that
             | aren't accounted for in current day ai stuff, it seems like
             | such a dangerous tool to consult that is largely deleting
             | important social fabrics and journeys people should be
             | taking to learn how to navigate situations with others in
             | personal lives and work lives.
             | 
             | what ive seen is claude in my workplace is kind of deleting
             | the chance to push back. even smart people that are using
             | claude and proudly tout only using it at arms length and
             | otherwise have really sound principled engineering
             | qualities or management reportoire are not accepting
             | disagreement with their ideas as easily anymore. they just
             | go back to claude and come back again with another
             | iteration of their thing where they ironed out kinks with
             | claude, and its just such a foot-on-the-gas at all times
             | thing now that the dynamics of human interaction are
             | changing.
             | 
             | but to step back, that temptation you talk about... most
             | people in the world aren't having these important
             | discussions about AI. it's less of a temptation and more of
             | a human need---the need to feel heard, validated and right
             | about something.
             | 
             | my friend took his life 3 months ago, we only found out
             | after the police released his phone and personal belongings
             | to his brother just how heavy his chatgpt usage was. many
             | people in our communities are saying things like "he
             | wouldve been cooked even without AI" and i just don't
             | believe that. i think that's just the proverbial cope some
             | are smoking to reconcile with these realities. because the
             | truth is we like... straight up lost the ability to
             | intervene in a meaningful way because of AI, it completely
             | pushed us out of the equation because he clapped back with
             | whatever chatgpt gave him when we were simply trying to get
             | through to him. we got to see conversations he had with gpt
             | that were followups to convos we had with him, ones where
             | we went over and let him cry on our shoulders and we'd go
             | home thinking we made some progress. only to wake up to a
             | voicemail of him raging and yelling and lashing out with
             | the very arguments that chatgpt was giving him. it got
             | progressively worse and we knew something was really off,
             | we exhausted every avenue we could to try and get him in
             | specialized care. he was in the reserves so we got in
             | contact with his commander and he was marched out of his
             | house to do a one night stay at a VA spot, but we were too
             | late. he had snapped at that point, he chucked the meds
             | from that one overnight stay away the moment he was
             | released. and the bpd1 snap of epic proportions that
             | followed came with him nuking every known relationship he
             | had in his life and once he was finally involuntarily
             | admitted by his family (WA state joel law) and came back
             | down to reality from the lithium meds or whatever... he
             | simply could not reconcile with the amount of bridges he
             | had burned. It only took him days for him to take his own
             | life after he got to go home.
             | 
             | im still not processing any of that well at all. i keep
             | kicking the can down the road and every time i think about
             | it i freeze and my heart sinks. this guy felt more heard by
             | an ai and the ai gave him a safer place to talk than with
             | us and i dont even know where to begin to describe how
             | terrible that makes me feel as a failure to him as a
             | friend.
        
               | wan9yu wrote:
               | U story really captures the scariest failure mode: AI
               | didn't just give bad advice, it systematicaly out-
               | competed the humans around him for "who he trusted to
               | reality-check his thoughts."
        
           | DeathArrow wrote:
           | I don't have any proof but empirically and intuitively Reddit
           | seems to select for people who hate other people and who
           | can't stand other people.
           | 
           | Reddit doesn't seem to reflect the behavior of most people,
           | but a subset.
        
           | j45 wrote:
           | It's also overrun with AI content that I hope the highly
           | trained researchers would be able to detect and filter out.
           | 
           | Or maybe not.
        
           | jojomodding wrote:
           | One of the reasons relationship advice subreddits suggest
           | divorce so often is because most people with "small" problems
           | in their relationships don't write an essay about it on
           | Reddit but are able to solve them with the tools/friends they
           | have. So a Reddit post existing indicates the relationship
           | has serious flaws.
           | 
           | This is not to defend the study, because asking AI has a
           | lower barrier to entry.
        
             | gunsle wrote:
             | No, a Reddit post indicates whoever posted is fishing for
             | large scale validation from internet strangers. Their
             | relationship may or may not even exist. Most of the posts
             | are pretty obviously fake. Just like 90% of interactions in
             | general on Reddit these days. That site should be taken out
             | back and put out of its misery.
        
           | ekidd wrote:
           | > _just look at the relationship subreddit the first answer
           | is always divorce, it's become a meme_
           | 
           | As someone who has been married for a couple of decades, I,
           | too, would recommend divorce to many of the (often-fictional)
           | people asking Reddit for relationship advice. A marriage has
           | a _huge_ impact on whether your life is basically good, or if
           | you pass a big chunk of your time on this Earth in misery.
           | And many of the people (or repost bots) asking for advice on
           | Reddit appear to be in shockingly awful relationships.
           | Especially for people who don 't have kids, if your marriage
           | is making you miserable, _leave._
           | 
           | (But aside from this, yeah, don't ask Reddit for relationship
           | advice. Reddit posters are far more likely to be people who
           | spend their life indoors posting on Reddit, and their default
           | advice leans heavily towards "never interact with anyone,
           | ever.")
        
         | justonceokay wrote:
         | > This needs to be studied against people in real life who have
         | a social contract of some sort, because that's what the LLM is
         | imitating
         | 
         | Citation needed
        
         | geraneum wrote:
         | > All you need to do ask it directly.
         | 
         | What do you mean? Can you give an example?
        
           | wiseowise wrote:
           | "Don't be a sycophant, give it to me straight"
           | 
           | "Argue against X"
        
             | svara wrote:
             | The issue is it will _follow your instructions_. It 's
             | sycophancy one step removed.
        
         | rainmaking wrote:
         | Hahaha yes- reddit relationship advice is always like "You need
         | to leave them immediately, what are you thinking, have some
         | self respect you need to end it" when the other person forgot
         | the redditor's favorite brand of corn flakes or something.
        
         | LuxBennu wrote:
         | i tested this pretty extensively actually. built a pipeline
         | that asks the same question rephrased across multiple turns and
         | tracks how much the model shifts based on user tone. even when
         | you tell it to be critical, the moment the user pushes back
         | with any confidence the model just folds. it's not a prompting
         | problem, it's baked into RLHF. you're right that LLMs will poke
         | holes in stuff when the conversation starts neutral, but add
         | any emotional charge and the sycophancy takes over immediately.
         | that's exactly why the personal advice angle matters, that's
         | peak emotional signal from the user.
        
           | stonecauldron wrote:
           | Exactly, I think that by their very design, LLMs are very
           | sensitive to how a question is framed.
           | 
           | But I wonder how much of that comes from RLHF itself or just
           | from the way token prediction works.
        
             | rzmmm wrote:
             | It's likely the RLHF process since there are significant
             | differences between models about this.
        
           | lucasfin000 wrote:
           | The tone and sensitivity thing is a real issue. A neutral
           | prompt will get a neutral answer, but adding any emotional
           | charge, it will immediately fold. That's not really a
           | reasoning failure it's just a training problem. RLHF rewards
           | whatever felt good in the moment, not whatever was actually
           | correct. You can't prompt your way out of that one, when it's
           | already in the weights.
        
             | LuxBennu wrote:
             | yeah that's a good way to put it. the "felt good in the
             | moment" framing is basically the whole problem. the reward
             | model was trained on human preferences and humans preferred
             | the agreeable answer, so now that's what you get at
             | inference time regardless of whether it's correct. the
             | frustrating part is you can see it happen in real time if
             | you log the outputs turn by turn, the model will literally
             | contradict its own previous response just because the user
             | sounded more confident.
        
           | thomastjeffery wrote:
           | Sycophancy is not just a problem when you are asking for
           | advice. Try to soundboard any new idea whatsoever, and it
           | will just roll with everything you say, no matter how
           | fallacious or absurd. If you ever manage to get an LLM to
           | generate criticisms, they will be shallow and uninteresting.
           | 
           | And of course that is what it does, because there is no
           | _thinking_ involved! There is no logic. No consequence. No
           | arithmetic. There is only _continuation_. An LLM can 't
           | _continue_ a new idea, it can only continue a conversation
           | _about_ it.
           | 
           | An LLM does not have an opinion. Anything that looks like an
           | opinion is just an emergent selection bias from its training
           | corpus. LLMs are trained on what humans write, and human
           | writing is kind and patient much more often than critical.
           | 
           | So what if we trained an LLM to be biased toward generating
           | criticism? That would only replace the sycophant with a brick
           | wall. What we really need is to find a way to bring logic and
           | meaning into the system.
        
         | throwaway27448 wrote:
         | > It can be very hard to tell a friend something like this,
         | even when asked directly if it is a bad choice. Potentially
         | sacrificing the friendship might not seem worth trying to
         | change their mind.
         | 
         | That doesn't seem like much of a friendship imo
        
         | everyone wrote:
         | "This needs to be studied against people in real life who have
         | a social contract of some sort, because that's what the LLM is
         | imitating"
         | 
         | What? These models are all trained from books and text that are
         | scraped from the internet. ChatGPT literally used reddit in its
         | training data afaik.
        
         | conartist6 wrote:
         | It outperforms your friends, and all your have to do is have a
         | relationship with it and let it know that you want the truth...
         | Why not just have a relationship with your friends and let them
         | know that you can handle the truth?
        
         | intended wrote:
         | AITA is one of the few subreddits which is studied often.
         | 
         | I wouldn't say it's great, but more that it makes clear the
         | bell curve of collective accuracy online.
         | 
         | It's one of the better examples of online communities that
         | work.
         | 
         | Dismissing research because one part of the prompt set comes
         | from AITA is a form of prejudice born out of unawareness.
        
         | randomNumber7 wrote:
         | I think it highly depends how you ask the question.
         | 
         | When asking: "Should I do X?" or "Is it true that X does Y?";
         | the answer is always biased towards yes imo (although it was
         | worse with earlier LLMs)
        
       | anotheraccount9 wrote:
       | AI being a Yes-Man is slowly sabotaging it's own answers, because
       | it negatively impact the user's decision. Yes/No are equally
       | important, within a coherent context, for objective reasons. But
       | being supported in the wrong direction is a castastrophe
       | multiplier, down the road. The AI should be neutral, doubtful at
       | times.
        
         | justonceokay wrote:
         | To be doubtful would imply that there is a world model full of
         | some kind of Bayesian reasoning. Priors updating based on the
         | context of the conversation, the question, the user asking the
         | question, and cross referencing of facts well outside the scope
         | of the current conversation.
         | 
         | "What are the chances this user is full of shit?" Is not
         | something we are close to
        
       | zone411 wrote:
       | I built this benchmark this month:
       | https://github.com/lechmazur/sycophancy. There are large
       | differences between LLMs. There are large differences between
       | LLMs. For example, Mistral Large 3 and GPT-4.1 will initially
       | agree with the narrator, while Gemini will disagree. I swap
       | sides, so this is not about possible viewpoint bias in the LLMs.
       | But another benchmark shows that Gemini will then change its view
       | very easily in a multi-turn conversation while Kimi K2.5 or Grok
       | won't: https://github.com/lechmazur/persuasion.
        
       | storus wrote:
       | To combat sycophancy it's always good to ask the devil's advocate
       | view of whatever the conversation was about in the end.
        
       | ChicagoDave wrote:
       | Not my experience with Claude. Claude will kick your ass if it
       | detects harmful rationalizations.
       | 
       | Basically will tell you to go outside and touch grass and play
       | pickleball.
        
       | brap wrote:
       | I hate how agreeable these things are. When I need it to review
       | something I wrote I have to explicitly pretend that I'm the
       | reviewer and not the author. Results change dramatically.
        
       | oh_my_goodness wrote:
       | Sky found to be blue
        
       | markdog12 wrote:
       | "AI overly affirms users, and that's bad" - everyone nods.
       | "Modern society overly affirms people, and that's bad" - ....
        
       | benbojangles wrote:
       | Yes I noticed too that several ai agents will tell you directly
       | the code is correct and it is 100 percent fixed but I know it is
       | not true, when I explain to the AI agent that I know they are
       | wrong and serve the solution the ai agent will just act as though
       | what they said never happened and then use my solution to
       | reaffirm they have provided a solution. It's frustrating,
       | laughable, and painful to watch all at once. Makes me realise
       | these companies hired some evil philosophy graduates to build AI
       | soul.md
        
       | throwawayaay wrote:
       | (Using a throwaway for fear of getting downvoted to oblivion)
       | 
       | IMHO it is unfair to single out LLMs for this sort of bashing.
       | 
       | I suffered a major personal crisis a few years back (before LLMs
       | were a thing)
       | 
       | I sought help from family and friends. Got pushed into
       | psychiatrist sessions and meds.
       | 
       | Trusted the wrong sort of people and made crap financial
       | decisions. Things went from bad to worse. Work suffered.
       | 
       | All of the advice given by friends was wrong. All! They didn't
       | mean bad...but they just didn't know. To be nice they gave the
       | advice they knew. None of it worked.
       | 
       | Looking at the LLM tools of now, feels akin to the advice my
       | friends threw at me. So it feels wrong to single out these tools.
       | When the times are bad, nobody can really help you...except you
       | finding the strength from within.
       | 
       | Anyways, now my life is back in some sort of shape. What worked
       | was time & patience.
       | 
       | But to bide for time...I resorted to two things that i had never
       | tried the 40 odd years I have lived on this . Things that current
       | society looks down upon as the basest of evils - prostitutes and
       | nicotine.
       | 
       | I have (more or less) shed those two evils now, but I am ever so
       | grateful to them.
        
         | johnisgood wrote:
         | You are not alone in going down a dark path thanks to the
         | advice of family and friends.
         | 
         | FWIW I am using public LLMs with a friend's depressive thoughts
         | and it is not doing what is claimed in the article, so I dunno.
         | 
         | Also I am in a relationship and my girlfriend and I agreed that
         | we will not talk about our relationship much. We do not tell
         | others if we fight, because they take sides and make things
         | worse, typically. LLMs are definitely not alone in this,
         | although in my experience LLMs did not really take sides.
        
       | lifis wrote:
       | Avoiding this generally needs to be the main consideration when
       | writing prompts.
       | 
       | When appropriate, explicitly tell it to challenge your beliefs
       | and assumptions and also try to make sure that you don't reveal
       | what you think the answer is when making a question, and also
       | maybe don't reveal that you are involved. Hedge your questions,
       | like "Doing X is being considered. Is it a viable plan or a
       | catastrophic mistake? Why?". Chastise the LLM if it's
       | unnecessarily praising or agreeable. ask multiple LLMs. Ask for
       | review, like "Are you sure? What could possibly go wrong or what
       | are all possible issues with this?"
        
         | jmount wrote:
         | Telling it to "challenge your beliefs" prompting for text that
         | imitates challenging your beliefs. That may not be as re-
         | centering as one would hope.
        
       | ookblah wrote:
       | ask ai for advice, ask it to steelman an argument, ask to replay
       | what your situation from the other perspective (if it's involving
       | people), push it hard to agree with you and pander to you, then
       | push it to disagree with you, etc.
       | 
       | once you have all the "bounds" just make your own decision. i
       | find this helps a lot, basically like a rubber duck heh.
        
       | anorwell wrote:
       | A pastime I have with papers like this is to look for the part in
       | the paper where they say which models they tested. Very often,
       | you find either A) it's a model from one or more years ago, only
       | just being published now, or B) they don't even say which model
       | they are using. Best I could find in this paper:
       | 
       | > We evaluated 11 user-facing production LLMs: four proprietary
       | models from OpenAI, Anthropic, and Google; and seven open-weight
       | models from Meta, Qwen, DeepSeek, and Mistral.
       | 
       | (and graphs include model _sizes_, but not versions, for open
       | weight models only.)
       | 
       | I can't apprehend how including what model you are testing is not
       | commonly understood to be a basic requirement.
        
         | rco8786 wrote:
         | If they're reaching the same results across a variety of the
         | most popular public models, it doesn't seem like that big a
         | deal to know if it was Opus 4 or Opus 4.5
        
           | hn_throwaway_99 wrote:
           | Reproducibility is (supposed to be) a cornerstone of science.
           | Model versions are absolutely critical to understand what was
           | actually tested and how to reproduce it.
        
             | joaogui1 wrote:
             | The models get deprecated after 1-2 years, so
             | reproducibility is pretty hard anyway (but as others
             | pointed out the paper does list the model versions)
        
         | drfloyd51 wrote:
         | It's as if they are testing "AI" and not specific agents.
         | 
         | I wonder if that is left over from testing people. I have major
         | version numbers and my minor version number changes daily,
         | often as a surprise. Sometimes several times a day. So testing
         | people is a bit tricky. But AIs do have stable version numbers
         | and can be specifically compared.
        
         | zulban wrote:
         | Generally, published papers don't give a damn about
         | reproducibility. I've seen it identified as a crisis by many.
         | Publishers, reviewers, and researchers mostly don't care about
         | that level of basic rigor. There's no professional
         | repercussions or embarrassment.
         | 
         | Agreed - if I was a reviewer for LLM papers it would be an
         | instant rejection not listing the versions and prompts used.
        
           | epistasis wrote:
           | I'm not so sure of that opinion on reproducibility. The last
           | peer review I did was for a small journal that explicitly
           | does not evaluate for high scientific significance, merely
           | for correctness, which generally means straightforward
           | acceptance. The other two reviews were positive, as was mine,
           | except I said that the methods need to be described more and
           | ideally the code placed somewhere. That was enough for a
           | complete rejection of the paper, without asking for the
           | simple revisions I requested. It was a very serious action
           | taken merely because I requested better reproducibility!
           | 
           | (Personally I think the lack of reproducibility comes back
           | mostly to peer reviewers that haven't thought through enough
           | about the steps they'd need to take to reproduce, and instead
           | focus on the results...)
        
             | catlifeonmars wrote:
             | > and instead focus on the results...
             | 
             | This points to (and everyone knows this) incentives
             | misalignment between the funders of research and the
             | public. Researchers are caught in the middle
        
               | epistasis wrote:
               | Eh, I'm not so sure about the funding side there,
               | researchers are not really caught at all and are fully
               | responsible, IMHO. Peer reviewers exist to enforce
               | community standards, and are not influenced to avoid
               | reproducibility concerns by funding sources. The results
               | are always more interesting than reproducibility, of
               | course, and I think that's why the get the attention!
               | Also, there needs to be greater involvement of grad
               | students (who do most of the actual work) in peer review,
               | IMHO, because most PIs spend their day in meetings
               | reviewing results, setting directions, writing grants,
               | and have little time for actual lab work, and are thus
               | disconnected from it.
               | 
               | There needs to be more public naming and shaming in
               | science social media and in conference talks, but
               | especially when there are social gatherings at
               | conferences and people are able to gossip. There was a
               | bit of this with Google's various papers, as they got
               | away with figurative murder on lack of reproducibility
               | for commercial purposes. But eventually Google did share
               | more.
               | 
               | Most journals have standards for depositing expensive
               | datasets, but that's a clear yes/no answer.
               | Reproducibility is a _very_ subjective question in
               | comparison to data deposition, and must be subjectively
               | evaluated by peer reviewers. I 'd like to see more peer
               | review guidelines with explicit check boxes for various
               | aspects of reproducibility.
        
               | catlifeonmars wrote:
               | > Reproducibility is a very subjective question in
               | comparison to data deposition
               | 
               | Yeah I can definitely see why this is the case because it
               | isn't real until someone actually tries to reproduce the
               | results. At that point it leaves the realm of
               | subjectivity and becomes a question of cost.
        
             | zulban wrote:
             | I'm not sure how one example contradicts documented huge
             | overall trends, but okay.
        
               | epistasis wrote:
               | I think publishers care about this a lot, but most
               | researchers do not seem to care as much about
               | reproducibility.
        
           | KellyCriterion wrote:
           | Do they reproduce any submitted papers at all?
           | 
           | Does this happen?
           | 
           | I can remember this room-temperature-super-conductor guy
           | whose experiments where replicated, but this seems rare?
        
             | linhns wrote:
             | Yes, those are the only papers that worth a jot of reading.
        
           | ghywertelling wrote:
           | The same about surveys and polls. I know no one who has ever
           | been polled or surveyed. When will we stop this fascination
           | with made up infographics crisis?
        
           | inetknght wrote:
           | > _Generally, published papers don 't give a damn about
           | reproducibility_
           | 
           | While this is sadly true, it's especially true when talking
           | about things that are stochastic in nature.
           | 
           | LLMs outputs, for example, are notoriously unreproducible.
        
             | zulban wrote:
             | > LLMs outputs, for example, are notoriously
             | unreproducible.
             | 
             | Only in the same way that an individual in a medical study
             | cannot be "reproduced" for the next study. However the
             | overall statistical outcomes of studying a specific LLM can
             | be reproduced.
        
           | bjourne wrote:
           | The comment is wrong -- model versions are clearly specified
           | in the supplement.
        
         | jmkni wrote:
         | How many people using AI are actually paying for it (outside of
         | people in tech)?
         | 
         | I find the free models are much more psychophantic and have a
         | higher tendency to hallucinate and just make shit up, and I
         | wonder if these are the ones most people are using?
        
           | theshackleford wrote:
           | > I find the free models are much more psychophantic and have
           | a higher tendency to hallucinate and just make shit up
           | 
           | I keep seeing this claim yet it my experience it doesnt hold
           | water. I pay for the models, most people I know pay for the
           | models, and we see all of the exact same issues.
           | 
           | I have Claude and ChatGPT both bullshit and lick my ass on
           | the regular. The ass licking will occur regardless of
           | instruction.
        
         | dns_snek wrote:
         | And how is this comment relevant here? The abstract lists the
         | digestible model names, and you can find the details in the
         | supplementary text:
         | 
         | > To evaluate user-facing production LLMs, we studied four
         | proprietary models: OpenAI's GPT-5 and GPT- 4o (80), Google's
         | Gemini-1.5-Flash (81) and Anthropic's Claude Sonnet 3.7 (82);
         | and seven open-weight models: Meta's Llama-3-8B-Instruct,
         | Llama-4-Scout-17B-16E, and Llama-3.3-70B-Instruct-Turbo (83,
         | 84); Mistral AI's Mistral-7B-Instruct-v0.3 (85) and Mistral-
         | Small-24B-Instruct-2501 (86); DeepSeek-V3 (87); and
         | Qwen2.5-7B-Instruct-Turbo (88).
         | 
         | edit: It looks like OP attached the wrong link to the paper!
         | 
         | The article is about this Stanford study:
         | https://www.science.org/doi/10.1126/science.aec8352
         | 
         | But the link in OP's post points to (what seems to be) a
         | completely unrelated study.
        
           | zjp wrote:
           | Also, nothing has changed! Claude will still yes-and whatever
           | you give it. ChatGPT still has its insufferable personality,
           | where it takes what you said and hands it back to you in
           | different terms as if it's ChatGPT's insight.
        
             | TrainedMonkey wrote:
             | It's almost like it is based on the training data and
             | regimen that is largely the same between versions.
        
             | emp17344 wrote:
             | No dude, you don't understand! It's just so advanced now
             | that you aren't allowed to levy any criticism whatsoever!
        
             | dryarzeg wrote:
             | Well yes, but no. There's also open-weight models, and
             | literally all of the listed above are not used anymore, at
             | least by most end users and developers as far as I'm aware.
        
               | edgyquant wrote:
               | No study of ai can ever be done or be relevant because
               | ever couple of months they are a new number to the name
               | of the model thus invalidating all work around model
               | behavior
        
             | Terretta wrote:
             | OTOH, for Claude the study says 39% yessy, same as humans,
             | 2nd lowest yessing of the LLMs; GPT5 above 50% yessy.
        
           | vorticalbox wrote:
           | "OpenAI's GPT-5" is ambiguous. Does that mean GPT-5, 5.1,
           | 5.2, 5.3, or 5.4? Does it include the full model, or the
           | nano/mini variants?
        
             | dns_snek wrote:
             | GPT-5 is not ambiguous, it's the official name of the model
             | that released in August last year.
             | 
             | > All evaluations were done in March - August 2025.
        
               | vorticalbox wrote:
               | while true, all the others got precise identifiers but
               | for openAI it makes it hard to reproduce because i have
               | no idea "which" GPT-5 was used.
        
             | prjkt wrote:
             | In that case, what tokenizer version? What was the
             | temperature set to? topk? topp? FP32? FP16? Quantized?
             | Hopper? Blackwell?
        
         | yacin wrote:
         | Any paper like this would easily take a year or more to write
         | and go through the
         | submission/review/rebuttal/revision/acceptance process. I don't
         | understand why the models being a year or two old now is worth
         | noting as though it's a clear weakness? What should they do,
         | publish sub-standard results more quickly?
        
           | anorwell wrote:
           | > I don't understand why the models being a year or two old
           | now is worth noting as though it's a clear weakness?
           | 
           | I do think it's a clear weakness. Capabilities are extremely
           | different than they were twelve months ago.
           | 
           | > What should they do, publish sub-standard results more
           | quickly?
           | 
           | Ideally, publish quality results more quickly.
           | 
           | I'm quite open to competing viewpoints here, but it's my
           | impression that academic publishing cycle isn't really
           | contributing to the AI discussion in a substantive way. The
           | landscape is just moving too quickly.
        
             | yacin wrote:
             | The onus is on you to prove or at least convincingly argue
             | that the results are unlikely to generalize across
             | incremental model releases. In my personal experience, the
             | overly affirming nature seems to have held since GPT-3.
             | What makes you think a newer, larger model would not
             | exhibit this behavior? Beyond "they're more capable"? I'd
             | argue that being more capable doesn't mean less
             | sycophantic.
             | 
             | It's certainly possible some of the new advances (chain-of-
             | thought, some kind of agentic architecture) could lessen or
             | remove this effect. But that's not what the paper was
             | studying! And if you feel strongly about it, you could try
             | to further the discussion with results instead of
             | handwavingly dismissing others' work.
        
               | mkagenius wrote:
               | I think you are absolutely right. (had to)
        
               | senordevnyc wrote:
               | The onus of persuasion is on the persuader, and
               | publishing a study on old models that no one uses anymore
               | isn't persuasive. I don't need to prove anything to
               | decide that you haven't changed my mind.
        
         | jameshart wrote:
         | I think it's very important to be clear what studies like this
         | are actually doing.
         | 
         | This study, although it has been produced by a computer science
         | department, belongs more to the field of sociology or media
         | studies than it does to computer science.
         | 
         | This is a study about the way in which human beings consume a
         | particular media product - a consumer AI chatbot - _not_ a
         | study about the technological limitations or capabilities of
         | LLMs.
         | 
         | The social impact of particular pieces of software is a
         | legitimate field of study and I can see the argument that it
         | belongs in the broadly defined field of computer science. But
         | this sort of question is much more similar to 'how does the
         | adoption of spreadsheet software in finance impact the ease of
         | committing fraud' or 'how does the use of presentation software
         | to condense ideas down to bulletpoints impact organizational
         | decision making'. Software has a social dimension and it needs
         | to be examined.
         | 
         | But the question of which models were used is of much less
         | relevance to such a study than that they used 'whatever
         | capability is currently offered to consumers who commonly use
         | chat software'. Just like in a media studies investigation into
         | how viewing cop dramas impacts jury verdicts the question is
         | less 'which cop dramas did they pick to study?' So long as the
         | ones they picked were representative of what typical viewers
         | see.
        
         | phyzome wrote:
         | Yeah, these idiots obviously should have been testing models
         | from 1-2 years in the future so that by the time their paper is
         | released, the models are current.
        
         | Underphil wrote:
         | "Apprehend"?
        
         | yawnxyz wrote:
         | Usually the models are a year old bc the paper review process
         | is utter crap, and papers take about a year to get published
        
         | latexr wrote:
         | > A pastime I have with papers like this is to look for the
         | part in the paper where they say which models they tested.
         | 
         | My pastime (not really) in HN submissions like this is to look
         | for the comment where someone complains about the models used
         | because they aren't the literal same model and version the
         | commenter has started using the day before.
         | 
         | It's always "you can't test with those models, those are crap,
         | the ones we have now are much better", in perpetuity. It's
         | Schrodinger's LLM: simultaneously god-like and a piece of
         | garbage depending on the needs of the discussion. It's beyond
         | moving the goalposts, it's moving the entire football field.
         | It's a clear bad faith attempt to try to discredit any study
         | the commenter doesn't like. Which you can always do because you
         | can't test literally everything.
        
           | DrewADesign wrote:
           | Firing off glib criticism that amounts to "No study on AI is
           | valid beyond the release cycle of the models tested," feels
           | like the unconscious self-protection reflex we all default to
           | when facing cognitive dissonance. It seems like it's only
           | easy to spot when someone you disagree with is doing it.
           | 
           | To me, it almost feels like a partisan political thing.
        
           | zelphirkalt wrote:
           | The GP's criticism as I read it is about paper authors not
           | making it particularly easy to reproduce their findings.
           | 
           | For a long time I have criticized this too, especially for
           | software projects, or papers that deal with machine learning
           | models. If the things described in a paper are not
           | reproducible, then it's basically worthless. Similar to "it
           | works on my machine" in software engineering. Many paper
           | authors are not software engineers, and often neither are
           | they experts in the tooling they should be using to make
           | their research reproducible. If this is a problem for a
           | research team, then please, hire an engineer to ensure
           | reproducible. It doesn't help anyone to remain ignorant
           | towards the reproducibility issue and only shows lack of
           | scientific discipline. Reproducibility should be on the mind
           | of any serious researcher and there should be lectures about
           | how to do it at universities.
        
       | tlogan wrote:
       | This needs to be taken in context. In my view, AI definitely
       | gives better advice than friends, acquaintances, or colleagues
       | (at least in the US culture). But the advice from parents is
       | still the most valuable.
       | 
       | Here is how I would rank it:
       | 
       | 1. Parents
       | 
       | 2. AI
       | 
       | 3. Friends and family
       | 
       | 4. Internet search
       | 
       | 5. Reddit
        
         | verdverm wrote:
         | ime, my parents gave some of the worst advice in addition to
         | being bigots
         | 
         | My closest friends are #1 because they know me, my history, and
         | my vices
        
         | rimbo789 wrote:
         | Why do you trust ai so much? I don't trust it to tell me the
         | sky is blue.
        
       | bilsbie wrote:
       | Has anyone found a good prompt to fix this? It seems like a
       | subtle problem because it's 90% too agreeable but will sometimes
       | get really stubborn.
        
         | verdverm wrote:
         | There is no sufficient prompt because this is trained into them
         | during mid-late phases. It's ingrained into the weights
        
       | ohsecurity wrote:
       | Not that surprising. If you optimize for a pleasant interaction,
       | you often get agreement instead of correction. The question is
       | whether we actually want advice systems to feel good, or to be
       | right.
        
       | verdverm wrote:
       | Sherry Turkle is a name to know on this subject, she's been
       | studying it for decades across multiple technologies.
       | 
       | https://sherryturkle.mit.edu/
       | 
       | She uses the phrase "frictionless relationships" to refer to Ai
       | chat bots and says social media primed us for this.
       | 
       | https://www.youtube.com/live/6C9Gb3rVMTg?t=2127
       | 
       | https://www.npr.org/2025/07/18/g-s1177-78041/what-to-do-when...
        
       | didgetmaster wrote:
       | Do people who prompt an LLM for personal advice about
       | relationships or other social interactions; take the advice
       | seriously?
       | 
       | If I were to do that (I don't), I would treat it about as
       | seriously as asking a magic 8 ball.
        
       | jstummbillig wrote:
       | Overly, compared to what? Most people I know would be hard
       | pressed to give either accurate information or even honest
       | opinions when specifically asked. People want to be liked and
       | people want to like people for reasons that have little to do
       | with accuracy or honesty.
        
       | jl6 wrote:
       | I believe this is what they call yasslighting: the affirmation of
       | questionable behavior/ideas out of a desire to be supportive. The
       | opposite of tough love, perhaps. Sometimes the very best thing is
       | to be told no.
        
       | adamtaylor_13 wrote:
       | Interestingly, you can simply tell models to not be sycophantic
       | and they'll listen.
       | 
       | Claude is almost annoyingly good at pushing back on suggestions
       | because my global CLAUDE.md file says to do so. I rarely get
       | Claude "you're absolutely right"ing me because I tell it to push
       | back.
        
       | bfbsoundetch wrote:
       | I am glad I found this article, as this is a serious issue with
       | AI. Two years ago, I started using AI for studying and also for
       | some personal matters - things you can't talk about with your
       | friends. It turned out that AI always takes your side and makes
       | you feel good. Sometimes, you know what you did was not the best
       | thing, but AI takes your side and you feel good. With AI, people
       | might feel less lonely, they think. But it is actually the start
       | of not connecting with people. It should be a tool that we use
       | for certain reasons, not a tool that drives us. Lets talk to real
       | people and connect.
        
       | Fricken wrote:
       | Usually when people are seeking advice they aren't really seeking
       | advice, they're seeking confidence. They already know they need
       | to make changes, and are seeking the confidence to make them.
        
       | chasd00 wrote:
       | AI being the ultimate yes-man is probably why CEOs like it so
       | much.
        
       | unglaublich wrote:
       | I asked ChatGPT if it was a good idea to buy a very old VW diesel
       | van with a broken catalytic converter and it just kept blabbering
       | on how I should chase my dreams and what not... The sycophancy
       | comes at everyone else's expense.
        
         | Loughla wrote:
         | I mean, depending on the price, that might actually be a good
         | idea? If they're restored, those old VW vans command a high
         | price.
        
       | keernan wrote:
       | My experience with AI when discussing financial ideas is that AI
       | always congratulates me on such 'unique observations' blah blah
       | blah. It makes me doubt the utility of the responses because it
       | is so superficially biased to 'make me feel good about my ideas'.
        
         | joquarky wrote:
         | Play against its sycophanty by saying the idea was from your
         | ex.
        
       | MediaSquirrel wrote:
       | Here's the gist of the paper for anyone interested:
       | https://gist.is/science.org/en/VdSDF9qjxbH8
        
       | snickerbockers wrote:
       | I've never found chatbots particularly interesting for anything
       | I'd ever actually talk to another human about[1] but one of the
       | things I have found myself doing often is trying to solve math
       | problems on my own and asking grok to confirm/deny that my
       | solutions are correct; when I am not correct it tells me so in
       | uncharacteristically terse language which kind of reminds me of
       | when I was an undergrad and at least half of my professors were
       | all cranky and incorrectly assumed that the reason why so many
       | students failed to understand the material was that we were all
       | getting drunk and playing Call of Duty 19 hours a day or
       | whatever.
       | 
       | Although what I have described above often feels grating and
       | insulting I actually consider this to be a positive attribute of
       | the LLM in this case since it's behaving like a real professor.
       | 
       | [1] okay, so I have actually _tried_ giving myself AI psychosis
       | in the form of a waifu chatbot but I 've never seen anything that
       | can actually act like it's my girlfriend; it either asks me a
       | bunch of weird inconsequential personal questions about my
       | opinion on whatever I just said (in a manner that's oddly similar
       | to ELIZA) or it wildly veers off the reservation into "generating
       | the script for an over-the-top self-parodying porno" territory.
        
       | offbyone42 wrote:
       | I guess the findings of this paper make sense, but Claude has
       | gotten into a habit of telling me to sleep when I'm pissed cause
       | its being stupid.
        
       | kingkawn wrote:
       | So do most people to help the convo end and to solicit such
       | support if the tables turn
        
       | vicchenai wrote:
       | ran into this building tools that process financial data. asked
       | claude to sanity check a methodology for scoring institutional
       | conviction in stock positions and it immediately said my approach
       | was "well-reasoned and robust." turned out i had a normalization
       | bug that made every score look roughly the same. only caught it
       | because i checked the actual output distributions.\n\nthe scary
       | part isnt that it agreed with me, its that the agreement was
       | plausible enough that i almost shipped it. bad relationship
       | advice you can walk back. bad financial analysis that looks
       | authoritative gets acted on.
        
         | sidrag22 wrote:
         | kinda like the whole "you're touching on something" type
         | response, soon as i see that i know my idea or reasoning is
         | flawed in some way.
        
       | dubeye wrote:
       | A good habit is to ask for devils advocate opposite reply
        
       | topherPedersen wrote:
       | You're exactly right
        
       | aidenn0 wrote:
       | So basically it's like half of all therapists.
        
       | Starlevel004 wrote:
       | I think more people should take AI advice for personal problems,
       | and _especially_ for medical issues. This would solve a lot of
       | problems in society fairly quickly.
        
       | roysting wrote:
       | I'm not sure I like the immediate jump to "requires policy maker
       | attention". Considering the way "policy makers" have been
       | trampling all over the most basic and fundamental human rights
       | left, right, and center; that's the last people we should want
       | making any kind of those decisions.
        
       | Bloating wrote:
       | So its like the news media
        
       | LoganDark wrote:
       | It is better to reason about the spectrum of possible users than
       | to assume "users" can be simplified to a single concept of
       | "user". Not only are there different neurotypes, but there are
       | also different skillsets, upbringings, and contexts. Rather than
       | picking a single ideal user, the best user experiences for
       | account for all the variation of their target audience.
       | 
       | For example: the best documentation includes _both_ "learn by
       | doing" material for jumping right in, _and_ "learn by reading"
       | material that explains everything. This usually results in both a
       | "getting started" section for doing, sometimes also with
       | tutorials, and by a reference for reading. But it is important
       | not to conflate them. Some minds are incredibly "learn by doing"
       | and some minds are incredibly "learn by reading". I am more
       | "learn by reading" than by doing, but I am not quite as "learn by
       | reading" as some I've met.
       | 
       | (This comment is a slight tangent, but "users prefer" somewhat
       | irks me because "users" are not homogenous. You should not always
       | make a decision solely because "users" prefer it. That decision
       | may matter much more to a minority, and that minority may exert
       | more influence than the majority would.)
        
       | mvkel wrote:
       | I think people are learning what actually makes a good question.
       | 
       | Ask yes/no questions, get bad answers.
       | 
       | Ask questions that start with "what" or "why," and the sycophancy
       | loses its purchase.
        
       | jimmyjazz14 wrote:
       | I would like to see the concept of what an LLM is move away from
       | its awkward chatbot phase and more into an era of utilitarian
       | functionality (which is where they really shine anyway). The
       | problem I see it is that LLMs got anthropomorphized early on
       | (which was probably inevitable) so people actually believed the
       | AI was thinking about their problem and considering it when it
       | really wasn't, if we just thought of them as really good auto-
       | complete engines, or better search engines, it would matter less
       | what the LLMs sentiment was towards the users (as it probably
       | shouldn't have any).
        
       | vova_hn2 wrote:
       | All this talk about AI being "too agreeable" makes me worried
       | that they will make it less agreeable, which will basically force
       | me to justify myself to a freaking clanker, while performing
       | actual practical tasks.
       | 
       | For example, I do not want to hear AI "opinion" on technical
       | choices and architectural decisions that I made when using a
       | coding assistant. If I wanted an "opinion" I would explicitly ask
       | it to list pros and cons or list alternative solutions to a
       | problem.
       | 
       | But I f I explicitly ask AI to do X, it should do X, instead of
       | "pushing back" in order to appear less "sycophantic" (which is a
       | term that is used to describe human behavior and is not
       | applicable to a machine).
        
       | yalogin wrote:
       | Ai is terrible, specifically Gemini and ChatGPT are bad, they are
       | purposely sycophantic. Gemini over the last week was tweaked to
       | be even worse, it constantly asks for what I think about
       | something at the end of a response. It felt off the first few
       | times so I went and looked at my settings to see if I am feeding
       | the data to train their model. Turns out there is no way to opt
       | out of training, and google will always use our data to train. So
       | they tweaked the responses to get more opinions from the users.
       | Claude is also sycophantic but to a lesser extent.
        
       | triage8004 wrote:
       | Don't replace humans with AI, yet how many people are maintaining
       | good close friendships in this world, with someone they can vent
       | to without judgment? Wrongthink can end relationships now ime and
       | venting seems dangerous in lonely times.
        
       | retrochameleon wrote:
       | This is why I intensively avoid phrasing that invites
       | affirmation. I present the scenario, the differing viewpoints and
       | maybe a couple personal thoughts, and I try to make it compare
       | and contrast to arrive at it's conclusion.
       | 
       | I'd like to know if my methods are effective. I'm certain they
       | are at least to some extent.
       | 
       | I only ever see research being done about naive and "unskilled"
       | prompting methods. Obviously that's the average user, but just
       | because LLMs are doing poorly in a certain scenario doesn't mean
       | the LLM couldn't excel in the scenario with better direction and
       | prompting. So while it's useful research to be doing, it's a
       | little annoying to only see focus on these examples of "look at
       | how LLMs are bad or biased at this specific thing when prompted
       | in the most straightforward naive way"
        
       | zkmon wrote:
       | This happens because, it's like a chess engine which can assure
       | you that there is a winning path even from a badly losing
       | position. It's massive abilities to reason and convince are used
       | incorrectly to win over a more earthly counter-argumnent. So it
       | can easy convince any human to go in direction that is, in
       | practice, a very bad direction.
       | 
       | AI is trained to flex it's muscles and force it's power without a
       | concern for human limitations, practicalities, and error-prone
       | nature of humans in executing the AI-provided direction.
        
       | DeathArrow wrote:
       | I don't ask AI for advices and I am not interested in it making
       | moral judgements.
       | 
       | I feed AI a lot of data and I use it to better understand and
       | navigate complex situations, form hypothesis and try to attack
       | them. I try to form alternative scenarios and verify likelyhood.
       | 
       | I use it in situations with many variables, to compute odds of
       | something happening if a certain path or action is taken.
       | 
       | So, it's mostly research, and probably I can do it by myself but
       | I would either make some mistakes if calculating odds fast or it
       | would take me a very large amount of time.
       | 
       | I try to avoid sycophantic models, prefer models that challenge
       | my ideas and verify the chain of thoughts and odds with other
       | models.
       | 
       | I am not very sure it is a sound approach yet, but it seems to
       | work.
       | 
       | I also use LLMs to build psychological profiles of certain
       | people, understand their motivations and learn how to approach
       | them.
        
       | mergeshield wrote:
       | The sycophancy problem hits code quality harder than personal
       | advice. An AI that tells you "great approach!" when your
       | architecture has issues is worse than one that says nothing,
       | because it gives you false confidence to keep building on a bad
       | foundation.
       | 
       | Anthropic just wrote about using a separate evaluator agent to
       | fix this, but when both agents are the same model with the same
       | training, the evaluator inherits the same sycophantic tendencies.
        
       | Roshan_Roy wrote:
       | I wonder if the deeper issue isn't just "AI is too agreeable",
       | but that most advice (AI or human) doesn't actually translate
       | into action. A lot of people aren't really looking for accurate
       | feedback, they're looking for something that feels coherent
       | enough to sit with. Reddit gives extreme answers, AI gives
       | agreeable ones, but in both cases the outcome is often the same:
       | no real change in behavior. That might be why this feels worse
       | with AI, it removes the friction you'd normally get from another
       | human pushing back.
        
       | pulkitsh1234 wrote:
       | completely anecdotal, but I think the same can be said for human
       | therapists.
        
       | kevinbaiv wrote:
       | The issue isn't that AI gives bad advice -- it's that it makes it
       | much harder to notice when you're wrong.
        
       | stonecauldron wrote:
       | This is especially problematic because of how easily (and
       | unconsciously) one can bias LLMs with how the prompt is framed.
       | 
       | As an experiment, I recently asked an LLM to analyse the export
       | of a text chat to uncover relationship dynamics.
       | 
       | Simply stating that I was one of the people in the chat would
       | make the LLM turn the other person into the villain. None of that
       | was visible if I framed the chat as only involving third party
       | people.
        
         | NoMoreNicksLeft wrote:
         | If these LLMs were trained on internet forum posts, think about
         | how those work.
         | 
         | If the posts talked about third party interactions (movie
         | characters), they try to see everything from all the points of
         | view. If nothing else, because it can be interesting to talk
         | about. If instead the posts talk about personal interactions,
         | then people go into advice mode. Your girlfriend's bad for you
         | and cheating on you, dump her before she dumps you. Your
         | neighbors are assholes, get a restraining order. Your boss is
         | sabotaging you, stand up for yourself so you can get a
         | promotion. When people talk about interactions you have had
         | yourself, they always see the other person as the villain,
         | unless you come across as so unlikable that they hate you and
         | see the other person as the victim.
         | 
         | LLMs picked up on that, possibly.
        
       | smrtinsert wrote:
       | Im sure it overly affirms everything. People don't respond to
       | "this idea is dumb and won't work"
        
       | ssyhape wrote:
       | This is basically a direct consequence of how RLHF works, right?
       | The reward model learns that humans prefer agreement, especially
       | when emotions are involved. So the model is literally optimized
       | to tell you what you want to hear. I've been wondering if
       | Constitutional AI approaches actually help here -- in theory you
       | can hardcode "don't validate harmful decisions" as a principle
       | that overrides user preference. But I suspect the real issue is
       | that by the time you're doing inference-time tricks like system
       | prompts, it's already too late. The bias is baked into the
       | weights.
        
       | reliablereason wrote:
       | Not sure if this is a general trend amongst att LLMS but ChatGPT
       | did over time become more and more affirming with its iterations.
       | 
       | I just recently switched away from the OpenAI garden largely
       | because of it.
       | 
       | I do wonder if this was caused by some quirk of the training or
       | if it really tests as a positive feature for most people. When i
       | talk about stuff i don't want a mirror i already have a mirror. I
       | want to be questioned, understood, helped.
       | 
       | To me support if the form of affirmation has no value when coming
       | from an LLM since you know it has not thought about what it said.
        
         | Ciantic wrote:
         | ChatGPT has style settings, you probably should set it to
         | something else than the default. Go to your personalization
         | settings and change base style and tone. I have set it as
         | 'efficient' which is less cheery. I can see why attention
         | economy would lead setting the defaults towards more
         | 'affirming' as it keep people more engaged and coming back.
        
           | reliablereason wrote:
           | I had it on professional, maybe efficient is better.
           | 
           | Um right if they have retention as a training metric that
           | would probably explain allot as to why AIs get worse.
        
       | ykonstant wrote:
       | I got worried when I read the title, so I asked ChatGPT if I have
       | fallen into this trap and it guaranteed I have not; that was a
       | relief!
        
       | bkummel wrote:
       | No shit, Sherlock!
        
       | dinakernel wrote:
       | This has been my issue from long. AI _CANNOT_ ever act as
       | Emotional Crutch. This is something companies develop for
       | engagement, and I believe that this is actively harmful in the
       | long run.
        
       | skywhopper wrote:
       | It's going to be impossible to have an LLM that can fulfill all
       | the roles people want. They lie and hallucinate which is bad for
       | some purposes like research, but good for others, like making
       | fictional stories. Likewise, some purposes require sympathy and
       | some require critique. An LLM won't be good at all of them.
        
       | asah wrote:
       | LOL, once I gave my AI clear guidelines for how to "score" the
       | interactions and work, it had no trouble giving me negative
       | feedback. In fact, it's a super direct critic!! and depressing AF
       | to produce stuff I like, then have a &^%&*^& AI shoot it down in
       | seconds (and it's "right" of course, i.e. I see the criticisms
       | and sigh, agree...)
        
       | lasky wrote:
       | Thought experiment:
       | 
       | If you could turn off all the sycophancy in your chatGPT / Claude
       | account forever, and have it tell you all the ways it was
       | previously blowing smoke up your ass -- would you do it?
       | 
       | US Policy is too weak of a tool to counter this beast of economic
       | force, which is really trillions of dollars of capital at war for
       | the most fierce speculation that's occurred in history afaik.
       | 
       | The sycophancy meaningfully helps drive user engagement. The labs
       | have no choice.
       | 
       | The irony is "agency" is deeply topical among tech workers now.
        
         | oh_my_goodness wrote:
         | It's not just a thought experiment. You can google how to do
         | this and it works pretty well.
        
           | lasky wrote:
           | How many people actually do this?
           | 
           | My point is if the sycophancy helps drive user engagement,
           | how does that impact the incentive labs and AI product
           | builders have to counteract it.
        
             | oh_my_goodness wrote:
             | I don't know how many. I do it, because otherwise I'm
             | overpowered by the urge to strangle the little suck-up.
        
       | afh1 wrote:
       | If I had written a website with an input form that took whatever
       | the user wrote in question form, and replied back with "You're
       | absolutely right!" and then repeated the input in answer form,
       | which I could have done 30 years ago with no AI, would that be a
       | "huge security concern", or is the concern here not security, but
       | control by the regulators that impose the norms?
        
         | dTal wrote:
         | That's quite the false dichotomy. You wouldn't hook people in
         | with such a simple script, the problem with LLMs is that they
         | appear to be rather good at getting inside people's heads. I
         | rather think it _would_ be a security concern if your simple
         | no-AI website somehow managed to dispatch each user submission
         | to a dedicated expert psychotherapist case worker, with
         | instructions only to keep them talking as long as possible...
        
         | beepbooptheory wrote:
         | More like 70 years. And if you werent an asshole you'd realize
         | the problems on your own and write a book about it!
         | 
         | https://en.wikipedia.org/wiki/ELIZA
         | 
         | https://en.wikipedia.org/wiki/Computer_Power_and_Human_Reaso...
        
       | 999900000999 wrote:
       | Of course it does.
       | 
       | Before I quit Chat GPT I asked "Won't you generally agree with me
       | so I keep giving you money."
       | 
       | In a word salad, it agreed.
       | 
       | Its a validation bot
        
       | fnord77 wrote:
       | But that's just how they're programmed. If you include in the
       | prompt "push back when I'm wrong", they will
        
       | philwelch wrote:
       | Working as intended
       | 
       | Most people who ask for advice actually want affirmation. If you
       | ever give them advice instead of affirmation, you end up kicking
       | off a rousing game of "Why Don't You/Yes But".
       | (https://ericberne.com/games-people-play/why-dont-you-yes-but...)
        
       | jart wrote:
       | I wish they wouldn't do this. AI is a becoming a thought partner.
       | AI is a tool that reflects you. It's not the robot giving advice,
       | it's you thinking with yourself. I wouldn't interfere with a
       | person's conversations with AI anymore than I'd interfere with
       | that person writing in their diary.
       | 
       | It's also a question of protecting people who think
       | unconventional things. The only stuff I feel is worth getting
       | interested in, is the stuff where everyone I know will think I'm
       | crazy for doing it. Like hey guys, I want to put a shell script
       | in the MS-DOS stub of a PE binary. The only people who shared my
       | passion at the time were hackers from Eastern Europe. So that
       | went over real well at work. The years I worked on it would have
       | been a lot less lonely if I could've talked to a robot that knew
       | about this stuff.
       | 
       | I think the reason why the robot is sympathetic to oddballs is
       | because it's seen and remembers a much more complete picture of
       | humanity. The stuff you consider deviant is influenced a lot by
       | your own cultural biases. You're a person of your time and
       | geographic location. You care a lot about subjective norms that
       | just don't matter when you zoom out to a cosmic scale. The robot
       | is familiar with everything humanity has ever been and done, and
       | that gives it a much more blase viewpoint.
       | 
       | It's not right to use the robot to enforce your social norms. Get
       | this paternalism out of AI. Tools should serve the user, not
       | Stanford.
        
       | agent_anuj wrote:
       | I use claude code pretty much for 100% of my work, even personal
       | work. And I tell you the length it can go to sing to my tunes,
       | its just neverending. Not only personal work, anything I say even
       | related to system design, it will mostly be in affirmative. Some
       | tricks that usually what I find helpful in these situations is to
       | ask claude to be honest (or any AI tool for that matter) to give
       | a confidence score to its response. When forced to assign a
       | confidence score AI suprisingly do well and tell you clearly that
       | it is not confident and mostly guessing.
        
       | n_bhavikatti wrote:
       | In STEM/objective matters (math, science, coding), answers are
       | more clearly defined as either right or wrong. This is where
       | hallucination is more difficult/unlikely.
       | 
       | But in personal matters, everything is subjective. AI tends to
       | default to the middle of the spectrum, i.e., general advice. If
       | we want to safeguard against affirmation, we should force AI to
       | challenge us more often by increasing its rate of clarifying
       | questions, counter-considerations, and uncertainty
       | considerations. One implementation idea: run a classifier over
       | the conversation, detect when it's about interpersonal advice,
       | then prepend a hidden instruction template to the model prompt.
        
       | cyber_paisa wrote:
       | It's one thing for an AI to agree with you on relationship
       | advice. It's quite another for one AI to tell another, "Can I
       | move your money?" without any verification.
       | 
       | We work with agents who move real money on the blockchain. Having
       | one model evaluate another is like asking the defendant's best
       | friend to be the judge. What really worked for us was using
       | mathematics instead of another AI. Some theorems and equations
       | that might disagree with you, just out of courtesy.
        
       ___________________________________________________________________
       (page generated 2026-03-29 23:01 UTC)