[HN Gopher] AI models routinely lie when honesty conflicts with ...
___________________________________________________________________
AI models routinely lie when honesty conflicts with their goals
Author : rntn
Score : 16 points
Date : 2025-05-01 19:37 UTC (3 hours ago)
(HTM) web link (www.theregister.com)
(TXT) w3m dump (www.theregister.com)
| alexjplant wrote:
| Is "2010" becoming real life? [1]
|
| > Chandra discovers the reasons for HAL's malfunction: the NSC
| ordered HAL to conceal information about the monolith from
| Discovery's crew, and programmed him to complete the mission
| alone. This conflicted with HAL's programming, the open and
| accurate processing of information, causing the computer
| equivalent of a paranoid breakdown.
|
| [1]
| https://en.m.wikipedia.org/wiki/2010:_The_Year_We_Make_Conta...
| b88m wrote:
| This title is simply the definition of hallucination in LLMs
| trained by RLHF.
| HarHarVeryFunny wrote:
| This makes it sound like LLMs are sentient and machiavellian,
| when they are just dumb statistical generators. An LLM doesn't
| have any goals other than the biases the model trainers chose to
| impart via RLHF, aside from which it is just trying to predict
| "what would the training data have said".
|
| If you insist on anthromorphizing the model, then it's goal is to
| please the RLHF testers, subject to pulling from the repertoire
| of responses that are statistically faithful to the training
| data.
___________________________________________________________________
(page generated 2025-05-01 23:02 UTC)