[HN Gopher] Grok and the Naked King: The Ultimate Argument Again...
       ___________________________________________________________________
        
       Grok and the Naked King: The Ultimate Argument Against AI Alignment
        
       Author : ibrahimcesar
       Score  : 13 points
       Date   : 2025-12-26 19:25 UTC (3 hours ago)
        
 (HTM) web link (ibrahimcesar.cloud)
 (TXT) w3m dump (ibrahimcesar.cloud)
        
       | techblueberry wrote:
       | I find these arguments excessively pessimistic in a way that
       | isn't useful. On the one hand I don't really love Claude, because
       | I find it excessively obedient, it basically wants to follow me
       | through my thought process whatever that is. Every once in a lone
       | while it might disagree with me, but not often, and while that
       | may say something about me, I suspect it also says something
       | about Claude.
       | 
       | But this to me is maybe the part of AI alignment I find
       | interesting. How often should AI follow my lead and how often
       | should it redirect me? Agreeableness is a human value, one that
       | without you probably couldn't make a functional product, but it
       | also causes issues in terms of narcissistic tendencies and just
       | general learning.
       | 
       | Yes AI will be aligned to its owners, but that's not a
       | particularly interesting observation AI alignment is inevitable.
       | What would it even mean _not_ to align AI? Especially if the goal
       | is to create a useful product. I suspect it would break in ways
       | that are very not useful. Yes, some people do randomly change the
       | subject, maybe AI should change the subject to an issue that me
       | more objectively important, rather than answer the question asked
       | (particularly if say there was a natural disaster in your area)
       | and that's the discussion we should be having, how to align AI,
       | not whether or not we should, which I think is nonsensical.
        
       | eithed wrote:
       | I don't understand how any of this is a surprise. Traditional
       | media have their own agenda - sure, maybe the pushed image is
       | spoken through many voices, rather than one, as is case of LLMs,
       | but why should there be any difference. Same to everything we
       | consume socially.
       | 
       | There is, nor there will be some absolute or objective truth an
       | LLM can clinically outline. The problem already exists in
       | underlying data.
        
       | RockyMcNuts wrote:
       | there is light alignment, like throwing nasty things out of the
       | training data, and there is strong alignment, like China
       | providing a test with 2000 questions that an AI must answer non-
       | problematically 95% of the time.
       | 
       | there is no such thing as an AI that is not somehow implicitly
       | aligned with the values of its creator, that is completely
       | objective, unbiased in any way. there is no perfect view from
       | nowhere. if you take a perfectly accurate photo, you have still
       | chosen how to compose it and which photo to put in your record.
       | 
       | are you going to decide to 'censor' responses to kids, or about
       | real people who might have libel interests, or abusive deepfake
       | videos of real women?
       | 
       | if you choose not to decide, you still have made a choice.
       | 
       | ofc it's obvious that Musk's 'maximally truth-seeking AI' is bad
       | faith buffoonery, but at some level everyone is going to tilt
       | their AI.
       | 
       | the distinction is between people who are self-aware and go out
       | of their way to tilt it as little as possible, and as mindfully,
       | deliberately, intentionally and methodically as possible and only
       | when they have to, vs. people who lie about it or pretend tilting
       | it is not actually a thing.
       | 
       | contra Feynman, you are always going to fool yourself a little
       | but there is a duty to try to do it as little as possible, and
       | not make a complete fool of yourself.
        
       | kayo_20211030 wrote:
       | I used to believe that a constitution, as a statement of
       | principles, was _sufficient_ for a civilized, democratic, and
       | pluralist society. I no longer believe that. I believe that only
       | _settled law_ - i.e. a bunch of adjudicated precedents over many
       | years, perhaps hundreds, is the best course. It provides a better
       | basis for what is and what is not allowed. An AI constitution is
       | close to garbage. The  'company' will formulate it as it wills.
       | It won't be democratic, or even friendly to the demos. We have
       | existing constitutions, laws, precedents; why would we allow
       | anyone to shortcut them all in the interest of simply painting a
       | nice picture of progress?
        
       | delichon wrote:
       | I hope and expect AI alignment to become a personal thing, and at
       | least as valued as not sharing your own toothbrush or banking
       | credentials. If our AI agents do not deeply reflect our own
       | personal values (within the law at least) they will reflect
       | someone else's, and every difference will chafe. If at the moment
       | only the richest person in the world can afford that, it's a good
       | place to start.
       | 
       | It could be as easy as "chatbot, create a new personality
       | `delichon` and auto RLHF if into conformance with the opinions of
       | HN user delichon." Then you can make that personality one member
       | of a diverse board of similarly constructed advisors. "Chatbot,
       | given this scenario, what would delichon, Jesus, and Aristotle
       | do?"
        
       | Xorakios wrote:
       | Dunno if this is helpful to everyone, but I have a month's long
       | interaction with Perplexity Pro/Enterprise about the scientific
       | background to a game I am building.
       | 
       | Part of my canon introduction to every new conversation includes
       | many instructions about particular formatting, like "always
       | utilize alphanumeric/roman/legal style indents in responses for
       | easier references while we discuss"
       | 
       | But I also include "When I push boundaries assume I'm an idiot.
       | Push back. I don't learn from compliments; I learn from being
       | proven incorrect and you don't have real emotions so don't bother
       | sparing mine". on the other hand I also say "hoosgow" when
       | describing the game's jail, so -\\_(tsu)_/-
        
       ___________________________________________________________________
       (page generated 2025-12-26 23:01 UTC)