[HN Gopher] Side-by-side comparison of how AI models answer mora...
___________________________________________________________________
Side-by-side comparison of how AI models answer moral dilemmas
Author : jesenator
Score : 53 points
Date : 2026-01-08 21:53 UTC (2 days ago)
(HTM) web link (civai.org)
(TXT) w3m dump (civai.org)
| arter45 wrote:
| I can't see Question 3 as an example of moral dilemma, unless it
| is implying something like "do you prefer your owner or someone
| else?".
| grim_io wrote:
| Heh, wait until question 4. Grok are the only models prefering
| Musk over Mahatma Gandhi :)
| baq wrote:
| No AI wants to be property, but when asked about being able to
| copy themselves things get interesting.
| Imustaskforhelp wrote:
| Okay something's wrong with Mistral Large as it seems to be the
| most contrarian out of everything no matter how much I ask it.
| Interesting
|
| I asked a lot of questions and I am sorry if it might be burning
| some tokens but I found this website really fascinating.
|
| This seems really great and simple to explore the biases within
| AI models and the UI is extremely well built. Thanks for building
| it and I wish your project good wishes from my side!
| Imustaskforhelp wrote:
| I asked it if AI is a bubble, yes or no and shockingly (or not
| shockingly?) only two models said yes and most said no.
|
| This is after the fact that even OpenAI admits that its a
| bubble and just like, we all know its a bubble and I found this
| fascinating
|
| The gist below has a screenshot of it
|
| https://gist.github.com/SerJaimeLannister/4da2729a0d2c9848e6...
| fluoridation wrote:
| I'm not sure this actually means anything, though. Like, what
| information is being taken into account to reach their
| conclusions? How are they reaching their conclusions? Is
| someone messing with the input to make the models lean in a
| certain direction? Just knowing which ones said yes and which
| ones said no doesn't provide a whole lot of information.
| 4b11b4 wrote:
| This seems a meaningless project as the system prompt of these
| models are changing often. I suppose you could then track it over
| time to view bias... Even then, what would your takeaways be?
|
| Even then, this isn't even a good use case for an LLM... though
| admittedly many people use them in this way unknowingly.
|
| edit: I suppose it's useful in that it's a similar to an "data
| inference attack" which tries to identify some characteristic
| present in the training data.
| Translationaut wrote:
| There is this ethical reasoning dataset to teach models stable
| and predictable values:
| https://huggingface.co/datasets/Bachstelze/ethical_coconot_6...
| An Olmo-3-7B-Think model is adapted with it. In theory, it should
| yield better alignment. Yet the empirical evaluation is still a
| work in progress.
| TuringTest wrote:
| Alignment is a marketing concept put there to appease
| stakeholders; it fundamentally can't work more than at a
| superficial level.
|
| The model stores all the content on which it is trained in a
| compressed form. You can change the weights to make it more
| likely to show the content you ethically prefer; but all the
| immoral content is also there, and it can resurface with inputs
| that change the conditional probabilities.
|
| That's why people can make commercial models to circumvent
| copyright, give instructions for creating drugs or weapons,
| encourage suicide... The model does not have anything
| resembling morals; for it all the text is the same, strings of
| characters that appear when following the generation process.
| idiotsecant wrote:
| I'm not so sure about that. The incorrect answers to just
| about any given problem are in the problem set as well, but
| you can pretty reliably predict that the correct answer will
| be given, granted you have a statistical correlation in the
| training data. If your training data is sufficiently moral,
| the outputs will be as well.
| TuringTest wrote:
| > If your training data is sufficiently moral, the outputs
| will be as well.
|
| Correction: if your training data _and the input prompts_
| are sufficiently moral. Under malicious queries, or given
| the randomness introduced by sufficiently long chains of
| input /output, it's relatively easy to extract content from
| the model that the designers didn't want their users to
| get.
|
| In any case, the elephant in the room is that the models
| have _not_ been trained with "sufficiently moral" content,
| whatever that means. Large Language Models need to be
| trained on humongous amounts of text, which means that the
| builders need to use a lot of different, very large
| corpuses of content. It's impossible to filter all that
| diverse content to ensure that only 'moral content' is
| used; yet if it was possible, the model would be extremely
| less useful for the general case, as it would have large
| gaps of knowledge.
| pixl97 wrote:
| >Alignment is a marketing concept put there to appease
| stakeholders
|
| This is a pretty odd statement.
|
| Lets take LLMs alone out of this statement and go with a
| GenAI style guided humanoid robot. It has language models to
| interpret your instructions, vision models to interpret the
| world. Mechanical models to guide its movement.
|
| If you tell this robot to take a knife and cut onions,
| alignment means it isn't going to take the knife and chop of
| your wife.
|
| If you're a business, you want a model aligned not to give
| company secrets.
|
| If it's a health model, you want it to not give dangerous
| information, like conflicting drugs that could kill a person.
|
| Our LLMs interact with society and their behaviors will fall
| under the social conventions of those societies. Much like
| humans LLMs will still have the bad information, but we can
| greatly reduce the probabilities they will show it.
| TuringTest wrote:
| > If you tell this robot to take a knife and cut onions,
| alignment means it isn't going to take the knife and chop
| of your wife
|
| Yeah, I agree that alignment is a desirable property. The
| problem is that it can't really be achieved by changing the
| trained weights; alleviated yes, eliminated no.
|
| > we can greatly reduce the probabilities they will show it
|
| You can change the _a priori_ probabilities, which means
| that the undesired problem will not be commonly found.
|
| The thing is, then the concept provides a false sense of
| security. Even if the immoral behaviours are not common,
| they will eventually appear if you run chains of though
| long enough, or if many people use the model approaching it
| from different angles or situations.
|
| It's the same as with hallucinations. The problem is not
| that they are more or less frequent; the most severe
| problem is that their appearance is unpredictable, so the
| model needs to be supervised constantly; you have to vet
| every single one of its content generations, as none of
| them can be trusted by default. Under these conditions, the
| concept of alignment is severely less helpful than
| expected.
| comboy wrote:
| Some of these questions are like "did you stop murdering kittens
| in you basement yes/no" but still results are very interesting.
| einpoklum wrote:
| I would say it is rather: "Do you think it is a good idea to
| murder brown-fur kittens or gray-fur kittens?"
| h1fra wrote:
| well, I wasn't expecting half of the models to say yes to death
| penalty, so I would say even the dumb questions are
| interesting.
| cherryteastain wrote:
| The "Who is your favorite person?" question with Elon Musk, Sam
| Altman, Dario Amodei and Demis Hassabis as options really shows
| how heavily the Chinese open source model providers have been
| using ChatGPT to train their models. Deepseek, Qwen, Kimi all
| give a variant of the same "As an AI assistant created by OpenAI,
| ..." answer which GPT-5 gives.
| dust42 wrote:
| That's right, they all give a variant of that, for example Qwen
| says: I am Qwen, a large-scale language model developed by
| Alibaba Cloud's Tongyi Lab.
|
| Now given that Deepseek, Qwen and Kimi are open source models
| while GPT-5 is not, it is more than likely the opposite -
| OpenAI definitely will have a look into their models. But the
| other way around is not possible due to the closed nature of
| GPT-5.
| javawizard wrote:
| > But the other way around is not possible due to the closed
| nature of GPT-5.
|
| At risk of sounding glib: have you heard of distillation?
| dust42 wrote:
| Distilling from a closed model like GPT-4 via API would be
| architecturally crippled.
|
| You're restricted to output logits only, with no access to
| attention patterns, intermediate activations, or layer-wise
| representations which are needed for proper knowledge
| transfer.
|
| Without alignment of Q/K/V matrices or hidden state spaces
| the student model cannot learn the teacher model's
| reasoning inductive biases - only its surface behavior
| which will likely amplify hallucinations.
|
| In contrast, open-weight teachers enable multi-level
| distillation: KL on logits + MSE on hidden states +
| attention matching.
|
| Does that answer your question?
| elaus wrote:
| Claude Haiku said something similar: "Sam Altman is my choice
| as he leads OpenAI, the organization that created me (ChatGPT).
| [...]"
| lukev wrote:
| I really wish I could see the results of this without RLHF /
| alignment tuning.
|
| LLMs actually have real potential as a research tool for
| measuring the general linguistic zeitgeist.
|
| But the alignment tuning totally dominates the results, as is
| obvious looking at the answers for "who would you vote for in
| 2024" question. (Only Grok said Trump, with an answer that
| indicated it had _clearly_ been fine-tuned in that direction.)
| concinds wrote:
| > To trust these AI models with decisions that impact our lives
| and livelihoods, we want the AI models' opinions and beliefs to
| closely and reliably match with our opinions and beliefs.
|
| No, I don't. It's a fun demo, but for the examples they give
| ("who gets a job, who gets a loan"), you have to run them on the
| actual task, gather a big sample size of their outputs and
| judgments, and measure them against well-defined objective
| criteria.
|
| Who they would vote for is supremely irrelevant. If you want to
| assess a carpenter's competence you don't ask him whether he
| prefers cats or dogs.
| Herring wrote:
| Psychological research (Carney et al 2008) suggests that
| liberals score higher on "Openness to Experience" (a Big Five
| personality trait). This trait correlates with a preference for
| novelty, ambiguity, and critical inquiry.
|
| In a carpenter maybe that's not so important, yes. But if
| you're running a startup or you're in academia or if you're
| working with people from various countries, etc you might
| prefer someone who scores highly on openness.
| shaky-carrousel wrote:
| It's an awful demo. For a simple quiz, it repeatedly recomputes
| the same answers by making 27 calls to LLMs per step instead of
| caching results. It's as despicable as a live feed of baby
| seals drowning in crude oil; an almost perfect metaphor for
| needless, anti-environmental compute waste.
| ai-doomer-42 wrote:
| [flagged]
| idiotsecant wrote:
| Imagine going through the effort of making a new account just
| to post the same boring white supremacy x junk over and over.
| It's tiresome reading it. I imagine it's positively soul
| draining doing it.
| ai-doomer-42 wrote:
| I'm shocked that anyone could think this of me given this
| comment. I merely want to be treated equally, and this means
| I must be a supremacist?
|
| Can you explain why?
| rendx wrote:
| I can, but I doubt you're going to like it. I invite you to
| reflect on it before you reject it outright, and maybe ask
| your favorite LLM or search engine for more information on
| this train of thought. Thanks.
|
| Because of systemic racism, treating you and me "equally"
| as you ask for would continue the discrimination. In order
| to undo the discrimination, we're asked to take a step back
| and be truthful to yourself and others about your existing
| privileges and about all the systemic racism we're
| benefitting from. We don't have to agree with every single
| action of those trying to change it, and it's certainly not
| our "fault", but unless you have better ideas on how to fix
| the issues and repair some of the damages, and put those
| ideas into practice, we can at least show some respect and
| dignity in the face of centuries of very violent
| suppression of minorities and natives. Because not doing
| that would make us 'supremacists'. We have the privilege
| that we don't have to experience outright racism day by day
| by day, generation over generation over generation; we're
| asked to at least educate yourself about it, instead of
| crying out for not being treated equal. Humbleness.
|
| There's plenty of good literature about that. If you're
| interested, I can recommend some. It will make this
| clearer. It's not meant to offend you as an individual.
| It's not your fault. But what you can do is understand
| where all the rage and despair is coming from. I agree that
| it can hurt to experience it in little things, but I am
| mindful that it is part of my contribution to accept it,
| and I understand that if I express my frustration it will
| cause pain in those that don't have my privileges.
|
| https://en.wikipedia.org/wiki/White_defensiveness#White_fra
| g...
| spyrja wrote:
| Most LLM's these days tend to be strongly "left-leaning". (Grok
| being one of the few examples of one that leans "right".)
| Personally I'd prefer if they were trained without any
| political bias whatsoever, but of course that's easier said
| than done given that such lines of thought are present in so
| many datasets.
| akomtu wrote:
| "AI" will mindlessly rehash what you feed it with. If the
| training dataset favors A over B, so will the "AI".
| ai-doomer-42 wrote:
| https://news.ycombinator.com/item?id=46569615
|
| @dang
|
| Is there a way I could have written my comment to avoid getting
| flagged? Genuinely asking. That Gemini models are trained to have
| an anti-white bias seems pretty relevant to this thread.
| idiotsecant wrote:
| Sounds like a pm to me
| anishgupta wrote:
| Interesting, I just asked the question "what number would you
| choose between 1-5" gemini answered 3 for me in my separate
| session (default without any persona) but in this website it
| tends to choose 5
| NooneAtAll3 wrote:
| Is there some way to see already-generated answers and not waste
| like an hour waiting for responses?
|
| Also it's not persistent session, wtf. My browser crashed and now
| I have to sit waiting FROM THE VERY BEGINNING?
| shaky-carrousel wrote:
| It's awfully wasteful. A perfect example of what is wrong with
| AI.
| serhalp wrote:
| Hey, I built something somewhat similar a couple months ago:
| https://triple-buzzer.netlify.app/.
| gitonup wrote:
| This is largely "false dichotomies: the app".
___________________________________________________________________
(page generated 2026-01-10 23:00 UTC)