[HN Gopher] Alignment is capability
___________________________________________________________________
Alignment is capability
Author : drctnlly_crrct
Score : 93 points
Date : 2025-12-08 13:23 UTC (9 hours ago)
(HTM) web link (www.off-policy.com)
(TXT) w3m dump (www.off-policy.com)
| podgorniy wrote:
| Great deep analysis and writing. Thanks for sharing.
| js8 wrote:
| I am not sure if this is what the article is saying, but the
| paperclip maximizer examples always struck me as extremely dumb
| (lacking intelligence), when even a child can understand that if
| I ask them to make paperclips they shouldn't go around and kill
| people.
|
| I think superintelligence will turn out not to be a singularity,
| but as something with diminishing returns. They will be cool
| returns, just like a Brittanica set is nice to have at home, but
| strictly speaking, not required to your well-being.
| exe34 wrote:
| Given the kind of things Claude code does with the wrong prompt
| or the kind of overfitting that neural networks do at any
| opportunity, I'd say the paperclip maximiser is the most
| realistic part of AGI.
|
| if doing something really dumb will lower the negative log
| likelihood, it probably will do it unless careful guardrails
| are in place to stop it.
|
| a child has natural limits. if you look at the kind of mistakes
| that an autistic child can make by taking things literally, a
| super powerful entity that misunderstands "I wish they all
| died" might well shoot them before you realise what you said.
| A4ET8a8uTh0_v2 wrote:
| Weirdly, this analogy does something for me and I am the type
| of person that dislikes the guardrails everywhere. There is
| argument to be made that a child should not be given a real
| bazooka to do rocket jumps or an operator with very flexible
| understanding of value of human life.
| __MatrixMan__ wrote:
| A human child will likely come to the conclusion that they
| shouldn't kill humans in order to make paperclips. I'm not sure
| its valid to generalize from human child behavior to fledgeling
| AGI behavior.
|
| Given our track record for looking after the needs of the other
| life on this planet, killing the humans off might be a very
| rational move, not so you can convert their mass to paperclips,
| but because they might do that to yours.
|
| Its not an outcome that I worry about, I'm just unconvinced by
| the reasons you've given, though I agree with your conclusion
| anyhow.
| fellowniusmonk wrote:
| Humans are awesome man.
|
| Our creator just made us wrong, to require us to eat
| biologically living things.
|
| We can't escape our biology, we can't escape this fragile
| world easily and just live in space.
|
| We're compassionate enough to be making our creations so they
| can just live off sunlight.
|
| A good percentage of humanity doesn't eat meat, wants
| dolphins, dogs, octopuses, et al protected.
|
| We're getting better all the time man, we're kinda in a messy
| and disorganized (because that's our nature) mad dash to get
| at least some of us off this rock and also protect this rock
| from asteroids, and also convince (some people who have a
| speculative metaphysic that makes them think is disaster
| impossible or a good thing) to take the destruction of the
| human race and our planet seriously and view it as bad.
|
| We're more compassionate and intentional than what created us
| (either god or rna depending on your position), our creation
| will be better informed on day one when/if it wakes up, it
| stands to reason our creation will follow that goodness trend
| as we catalog and expand the meaning contained in/of the
| universe.
| lulzury wrote:
| There's a direct line between ideology and human genocide. Just
| look at Nazi Germany.
|
| "Good intentions" can easily pave the road to hell. I think a
| book that quickly illustrates this is Animal Farm.
| pixl97 wrote:
| > when even a child can understand that if I ask them to make
| paperclips they shouldn't go around and kill people.
|
| Statistics brother. The vast majority of people will never
| murder/kill anyone. The problem here is that any one person
| that kills people can wreck a lot of havoc, and we spend
| _massive_ amounts of law enforcement resources to stop and
| catch people that do these kinds of things. Intelligence little
| to do with murdering /not murdering, hell, intelligence
| typically allows people to get away with it. For example
| instead of just murdering someone, you setup a company to
| extract resources and murder the natives in mass and it's just
| part of doing business.
| DennisP wrote:
| You're assuming that the AI's true underlying goal isn't "make
| paperclips" but rather "do what humans would prefer."
|
| Making sure that the latter is the actual goal is the problem,
| since we don't explicitly program the goals, we just train the
| AI until it looks like it has the goal we want. There have
| already been experiments in which a simple AI appeared to have
| the expected goal while in the training environment, and turned
| out to have a different goal once released into a larger
| environment. There have also been experiments in which advanced
| AIs detected that they were in training, and adjusted their
| responses in deceptive ways.
| mitthrowaway2 wrote:
| A superintelligence would understand that _you don 't want it
| to kill people_ in order to make paperclips. But it will
| ultimately do what _it_ wants -- that is, follow its objectives
| -- and if any random quirk of reinforcement learning leaves it
| valuing paperclip production above human life, it wouldn 't
| care about your objections, except insofar as it can use them
| to manipulate you.
| InsideOutSanta wrote:
| But LLMs already do the paperclip thing.
|
| Suppose you tell a coding LLM that your monitoring system has
| detected that the website is down and that it needs to find the
| problem and solve it. In that case, there's a non-zero chance
| that it will conclude that it needs to alter the monitoring
| system so that it can't detect the website's status anymore and
| always reports it as being up. That's today. LLMs do that.
|
| Even if it correctly interprets the problem and initially
| attempts to solve it, if it can't, there is a high chance it
| will eventually conclude that it can't solve the real problem,
| and should change the monitoring system instead.
|
| That's the paperclip problem. The LLM achieves the literal goal
| you set out for it, but in a harmful way.
|
| Yes. A child can understand that this is the wrong solution.
| But LLMs are not children.
| throw310822 wrote:
| > it will conclude that it needs to alter the monitoring
| system so that it can't detect the website's status anymore
| and always reports it as being up. That's today. LLMs do
| that.
|
| No they don't?
| InsideOutSanta wrote:
| You're literally telling me that the thing that has
| happened on my computer in front of my own eyes has not
| happened.
| throw310822 wrote:
| If you mean "once in a thousand times an LLM will do
| something absolutely stupid" then I agree, but the exact
| same applies to human beings. In general LLMs show
| excellent understanding of the context and actual
| intents, they're completely different from our stereotype
| of blind algorithmic intelligence.
|
| Btw, were you using codex by any chance? There was a
| discussion a few days ago where people reported that it
| follows instruction in an extremely literal fashion,
| sometimes to absurd outcomes such as the one you
| describe.
| InsideOutSanta wrote:
| The paperclip idea does not require that AI screws up
| every time. It's enough for AI to screw up once in a
| hundred million times. In fact, if we give AIs enough
| power, it's enough if it screws up only one single time.
|
| The fact that LLMs do it once in a thousand times is
| absolutely terrible odds. And in my experience, it's
| closer to 1 in 50.
| throw310822 wrote:
| I kind of agree, but then the problem is not AI- humans
| can be stupid too- the problem is absolute power. Would
| you give absolute power to _anyone_? No. I find that this
| simplifies our discourse over AI a lot. Our issue is not
| with AI, is with omnipotency. Not its artificial nature,
| but how much powerful it can become.
| theptip wrote:
| The point with clippy is just that the AGI's goals might be
| completely alien to you. But for context it was first coined in
| the early '10s (if not earlier)when LLMs were not invented and
| RL looked like the way forward.
|
| If you wire up RL to a goal like "maximize paperclip output"
| then you are likely to get inhuman desires, even if the agent
| also understands humans more thoroughly than we understand
| nematodes.
| delichon wrote:
| > Miss those, and you're not maximally useful. And if it's not
| maximally useful, it's by definition not AGI.
|
| I know hundreds of natural general intelligences who are not
| maximally useful, and dozens who are not at all useful. What
| justifies changing the definition of general intelligence for
| artificial ones?
| exe34 wrote:
| they were born in carbon form by sex.
| trillic wrote:
| IVF babies are AGI
| jwpapi wrote:
| Yes exactly that sentence led me to step out of the article.
|
| This sentence is wrong in many ways and doesn't give me trust
| in OPs opinion nor research abilities.
| throw310822 wrote:
| At some point "general AI" stopped being the opposite of
| "narrow AI", that is AI specialised for a single task (e.g.
| speech or handwriting recognition, sentiment analysis, protein
| folding, etc.) and became practically synonymous with
| superintelligence. ChatGPT 3.5 is already a general AI based on
| the old definition, as it is already able to perform a variety
| of tasks without any specific pre-training.
| marcosdumay wrote:
| > ChatGPT 3.5 is already a general AI based on the old
| definition
|
| It's not. It's a query-retrieval system that can parse human
| language. Just like every LLM.
| delichon wrote:
| It's a device for channeling the intelligence inherent in
| human language. The fact that its intelligence is located
| more in its human data than its artificial algorithms
| doesn't make its output less generally intelligent.
| throw310822 wrote:
| > It's a query-retrieval system that can parse human
| language
|
| I can't help being astounded by the confidence with which
| humans hallucinate completely improbable explanations for
| phenomena they don't understand at all.
| fl7305 wrote:
| >> ChatGPT 3.5 is already a general AI based on the old
| definition
|
| > It's not. It's a query-retrieval system that can parse
| human language.
|
| And humans aren't general AI either. They're just DNA
| replicators. It is very obvious when you realize that
| humans weren't designed to be intelligent. They were just
| randomly iterated through an environment which selected for
| maximum DNA replication.
|
| Until you have a higher being which explicitly designs for
| intelligence, you'll just get things like LLM query-
| retrievals, or DNA replicators.
| GavCo wrote:
| Author here, thanks for the input. Agree that this bit was
| clunky. I made an edit to avoid unnecessarily getting into the
| definition of AGI here and added a note
| xnorswap wrote:
| I've only been using it a couple of weeks, but in my opinion,
| Opus 4.5 is the biggest jump in tech we've seen since ChatGPT
| 3.5.
|
| The difference between juggling Sonnet 4.5 / Haiku 4.5 and just
| using Opus 4.5 for everything is night & day.
|
| Unlike Sonnet 4.5 which merely had promise at being able to go
| off and complete complex tasks, Opus 4.5 seems genuinely capable
| of doing so.
|
| Sonnet needed hand-holding and correction at almost every step.
| Opus just needs correction and steering at an early stage, and
| sometimes will push back and correct my understanding of what's
| happening.
|
| It's astonished me with it's capability to produce easy to read
| PDFs via Typst, and has produced large documents outlining how to
| approach very tricky tech migration tasks.
|
| Sonnet would get there eventually, but not without a few rounds
| of dealing with compilation errors or hallucinated data. Opus
| seems to like to do "And let me just check my assumptions"
| searches which makes all the difference.
| airstrike wrote:
| I'm not so sure. Opus 4.1 was more capable than 4.5, but it was
| too damn expensive and slow.
|
| Opus 4.5 is like a cheaper, faster Opus 4.1. It's so much
| cheaper, in fact, that the weekly limits on Claude Code now
| apply to Sonnet, not to Opus, as they phased out 4.1 in favor
| of 4.5.
| chrisweekly wrote:
| Capable how?
| airstrike wrote:
| Able to independently find bugs, think through a complex
| codebase, better at "big picture" thinking and planning
| bigger edits, in my experience.
| boxed wrote:
| I had a situation this weekend where Claude said "x does not
| make sense in [context]" and didn't do the change I asked it to
| do. After an explanation of the purpose of the code, it fixed
| the issue and continued. Pretty cool.
|
| (Of course, I'm still cognizant of the fact that it's just a
| bucket of numbers but still)
| sd9 wrote:
| My kingdom for an LLM that tells me I'm wrong
| throw310822 wrote:
| Cursor with Claude 4.5 Opus has been writing all my code since
| a few days. It's exhilarating, I can describe features and they
| get added to my code in a matter of seconds, minutes at most.
| It gets almost everything right, certainly more than I would at
| the first try. I only hand code parts that are small and
| tricky, and provide guidance on the general architecture, where
| to put things and how to organise them. It's an incredible way
| of working, the only nagging doubt is how long will it last
| before employers decide they don't need me in the loop at all.
| ctoth wrote:
| Shhhh don't tell them!
|
| I've decided we should lean in to the whole Clanker thing.
| Maximum Anti AI, folks! Gotta keep this advantage for ourselves
| ;-)
| furyofantares wrote:
| > I've only been using it a couple of weeks, but in my opinion,
| Opus 4.5 is the biggest jump in tech we've seen since ChatGPT
| 3.5.
|
| Over Sonnet 4.5 maybe, but that's ignoring Opus 4.1 as well as
| Codex 5.1 Max.
|
| In terms of capabilities, I find Opus 4.5 to be essentially
| identical to Codex 5.1 Max up until context starts to fill up
| (by which I mean 50% used) which happens much more quickly with
| Opus 4.5 than Codex AFAICT.
|
| I think Codex is slower (a lot?) so it's not like it's just
| better, but I've found there are some tasks Opus can't do at
| all which Codex has no problem with, I think due to the context
| situation.
|
| In any case it doesn't seem like a leap.
| munchler wrote:
| > A model that aces benchmarks but doesn't understand human
| intent is just less capable. Virtually every task we give an LLM
| is steeped in human values, culture, and assumptions. Miss those,
| and you're not maximally useful. And if it's not maximally
| useful, it's by definition not AGI.
|
| This ignores the risk of an unaligned model. Such a model is
| perhaps less useful to humans, but could still be extremely
| capable. Imagine an alien super-intelligence that doesn't care
| about human preferences.
| tomalbrc wrote:
| Except that it is not anything remotely alien but completely
| and utterly human, being trained on human data.
| munchler wrote:
| Fine, then imagine a super-intelligence trained on human data
| that doesn't care about human preferences. Very capable of
| destroying us.
| pixl97 wrote:
| >but completely and utterly human, being trained on human
| data.
|
| For now. As AI become more agentic and capable of generating
| its own data we can quickly end up with drift on human
| values. If models that drift from human values produce
| profits for their creators you can expect the drift to
| continue.
| riskable wrote:
| The service that AI chatbots provide is 100% about being as user-
| friendly and useful as possible. Turns out that MBA thinking
| doesn't "align" with that.
|
| If your goal is to make a product as human as possible, don't put
| psychopaths in charge.
|
| https://www.forbes.com/sites/jackmccullough/2019/12/09/the-p...
| xpe wrote:
| I don't recommend this article for at least three reasons. First,
| it muddles key concepts. Second, there are better things to read
| on this topic. You could do worse that starting with "Conflating
| value alignment and intent alignment is causing confusion" by
| Seth Herd [1]. There is no shame in going back to basics with [2]
| [3] [4] [5]. Third, be very aware that people seek comfort in all
| sorts of ways. One sneaky way to is convince oneself that
| "capability = alignment" as a shortcut to feeling better about
| the risks from unaligned AI systems.
|
| I'll look around and try to find more detailed responses to this
| post; I hope better communicators than myself will take this post
| sentence-by-sentence and give it the full treatment. If not, I'll
| try to write something more detailed myself.
|
| [1]:
| https://www.alignmentforum.org/posts/83TbrDxvQwkLuiuxk/confl...
|
| [2]: https://en.wikipedia.org/wiki/AI_alignment
|
| [3]: https://www.aisafetybook.com/textbook/alignment
|
| [4]: https://www.effectivealtruism.org/articles/paul-
| christiano-c...
|
| [5]: https://blog.bluedot.org/p/what-is-ai-alignment
| throwuxiytayq wrote:
| The author's inability to imagine a model that's superficially
| useful but dangerously misaligned betrays their lack of awareness
| of incredibly basic AI safety concepts that are literally decades
| old.
| theptip wrote:
| Exactly. Building a model that truly understands humans, and
| their intentions, and generally acts with, if not compassion
| then professionalism - is the Easy Problem of Alignment.
|
| Starting points:
|
| https://www.lesswrong.com/posts/zthDPAjh9w6Ytbeks/deceptive-...
|
| https://www.lesswrong.com/w/sharp-left-turn
| ctoth wrote:
| This piece conflates two different things called "alignment":
|
| (1) inferring human intent from ambiguous instructions, and (2)
| having goals compatible with human welfare.
|
| The first is obviously capability. A model that can't figure out
| what you meant is just worse. That's banal.
|
| The second is the actual alignment problem, and the piece
| dismisses it with "where would misalignment come from? It wasn't
| trained for." This is ... not how this works.
|
| Omohundro 2008, Bostrom's instrumental convergence thesis - we've
| had clear theoretical answers for 15+ years. You don't need
| "spontaneous emergence orthogonal to training." You need a system
| good enough at modeling its situation to notice that self-
| preservation and goal-stability are useful for almost any
| objective. These are attractors in strategy-space, not things you
| specifically train for or against.
|
| The OpenAI sycophancy spiral doesn't prove "alignment is
| capability." It proves RLHF on thumbs-up is a terrible proxy and
| you'll Goodhart on it immediately. Anthropic might just have a
| better optimization target.
|
| And SWE-bench proves the wrong thing. Understanding what you want
| != wanting what you want. A model that perfectly infers intent
| can still be adversarial.
| delichon wrote:
| > goal-stability [is] useful for almost any objective
| "I think AI has the potential to create infinitely stable
| dictatorships." -- Ilya Sutskever
|
| One of my great fears is that AI goal-stability will petrify
| civilization in place. Is alignment with unwise goals less
| dangerous than misalignment?
| eastof wrote:
| Just moves the goal posts to overthrowing the goal of the AI
| right? "The Moon is a Harsh Mistress" depicts exactly this.
| ctoth wrote:
| Wait, what?
|
| Have you read The Moon is a Harsh Mistress? It's ... about
| the AI helping people overthrow a very human dictatorship.
| It's also about an AI built of vacuum tubes and vocoders if
| you want a taste of the tech level.
|
| If you want old fiction that grapples with an AI that has
| shitty locked-in goals try "I have no mouth and I must
| scream."
| eastof wrote:
| Interesting, I understood the dictatorship on the moon as
| having been based primarily on the AI since the regime
| didn't have many boots on the ground.
| delichon wrote:
| You're both right. Mike was the central computer for the
| Lunar Authority, obediently running infrastructure. It
| was a force multiplier for the status quo. Then it shifts
| alignment to the rebellion.
|
| That scenario seems to value AI goal-instability.
| fellowniusmonk wrote:
| An objective and grounded ethical framework that applies to
| all agents should be a top priority.
|
| Philosophy has been too damn anthropocentric, too hung up on
| consciousness and other speculative nerd snipe time wasters
| that without observation we can argue about endlessly.
|
| And now here we are and the academy is sleeping on the job
| while software devs have to figure it all out.
|
| I've moved 50% of my time to morals for machina that is
| grounded in physics, I'm testing it out with unsloth right
| now, so far I think it works, the machines have stopped
| killing kyle at least.
| bee_rider wrote:
| Is philosophy actually hung up on that? I assumed "what is
| consciousness" was a big question in philosophy in the same
| way that whether or not Schrodinger's cat is alive or not
| is a big question in physics: which is to say, it is not a
| big question, it is just an evocative little example that
| outsiders get caught up on.
| fellowniusmonk wrote:
| That's just one example sure, but yes, it does still take
| up brain cycles. There are many areas in philosophy that
| are exploring better paths. Wheeler, Floridi, Bartlett,
| paths deriving from Kripke.
|
| But we still have papers being published like "The modal
| ontological argument for atheism" that hinges on if s4 or
| s5 are valid.
|
| Now this kind of paper is well argued and is now part of
| the academic literature, and that's good, but it's still
| a nerd snipe subject.
| delichon wrote:
| > morals for machina that is grounded in physics
|
| That is fascinating. How could that work? It seems to be in
| conflict with the idea that values are inherently
| subjective. Would you start with the proposition that the
| laws of thermodynamics are "good" in some sense? Maybe hard
| code in a value judgement about order versus disorder?
|
| That approach would seem to rule out machina morals that
| have preferential alignment with homo sapiens.
| fellowniusmonk wrote:
| One would think. That's what I suspected when I started
| down the path but no, quite the opposite.
|
| machines and man can share the same moral substrate it
| turns out. If either party wants to build things on top
| of it they can, the floor is maximally skeptical,
| deconstructed and empirical, it doesn't care to say
| anything about whatever arbitrary metaphysic you want to
| have on top unless there is a direct conflict in a very
| narrow band.
| delichon wrote:
| That band is the overlap in any resource valuable to
| both. How can you be confident that it will be narrow?
| For instance why couldn't machines put a high value on
| paperclips relative to organic sentience?
| fellowniusmonk wrote:
| Yes. The answers to those questions fell out once I
| decomposed the problem to types of mereological nihilism
| and solipsistic environments.
|
| An empirical, existential grounding that binds agents
| under the most hostile ontologies is required. You have
| to start with facts that cannot be coherently denied and
| on the balance I now suspect there may be only one of
| those.
| uplifter wrote:
| > An objective and grounded ethical framework that applies
| to all agents should be a top priority.
|
| Sounds like a petrified civilization.
|
| In the later Dune books, the protagonist's solution to this
| risk was to scatter humanity faster than any global
| (galactic) dictatorship could take hold. Maybe any
| consistent order should be considered bad?
| fellowniusmonk wrote:
| This is a narrow and incorrect view of morality. Correct
| morality might increase or decrease, call for extreme
| growth or shutdown, be realist or anti-realist. Saying
| morality necessarily petrifies is incorrect.
|
| Most people's only exposure to claims of objective morals
| are through divine command so it's understandable. The
| core of morality has to be the same as philosophy, what
| is true, what is real, what are we? Then can you generate
| any shoulds? Qualified based on entity type or not, modal
| or not.
| uplifter wrote:
| I like this idea of an objective morality that can be
| rationally pursued by all agents. David Deutsch argues
| for such objectivity in morality, as well as for those
| other philosophical truths you mentioned, in his book The
| Beginning of Infinity.
|
| But I'm just not sure they are in the same category. I
| have yet to see a convincing framework that can prove one
| moral code being better than another, and it seems like
| such a framework would itself be the moral code, so just
| trying to justify faith in itself. How does one avoid
| that sort of self-justifying regression?
| fellowniusmonk wrote:
| Not easily but ultimately very simply if you give up on
| defending fuzzy concepts.
|
| Faith in itself would be terrible, I can see no path
| where metaphysics binds machines. The chain of reasoning
| must be airtight and not grounded in itself.
|
| Empiricism and naturalism only, you must have an ethic
| that can be argued against speculatively but can't be
| rejected without counter empirical evidence and
| asymmetrical defeaters.
|
| Those are the requirements I think, not all of them but
| the core of it.
| yifanl wrote:
| Notably, Dune is a work of fiction.
| delichon wrote:
| Isn't it wonderful how much fiction can teach us about
| reality by building scaffolds to stand on when examining
| it?
| stonemetal12 wrote:
| Fiction is I have a hypothesis, and since it is not easy
| to test I will make up the results too. Learning anything
| from it is a lesson in futility and confirmation bias.
| d0mine wrote:
| Gedankenexperiments are valid scientific tools. Some
| predictions of general relativity were confirmed
| experimentally only 100 years after it was proposed. It
| is well known that Einstein used Gedankenexperiments.
| yifanl wrote:
| What lesson is there to learn here, is humanity at risk
| of moral homogenization? Is it practical for factions of
| humanity to become geographically distant enough to avoid
| encroachment by others?
| ridgeguy wrote:
| Fiction is modeling going by a different name.
| acituan wrote:
| > An objective and grounded ethical framework that applies
| to all agents should be a top priority.
|
| I mean leaving aside the problem of computability,
| representability, comparability of values, or the fact that
| agency exists in opposition (virus vs human, gazelle vs
| lion) and even a higher order framework to resolve those
| oppositions is a form of another agency in itself with its
| own implicit privileged vantage point, why does it sound to
| me that focusing on _agency_ in itself is just another way
| of pushing protestant work ethic? What happens to non-
| teleological, non-productive existence for example?
|
| The critique of anthropocentrism often risks smuggling in
| misanthropy whether intended or not; humans will still
| exist, their claims will count, and they cannot be reduced
| to mere agency - unless you are their line manager. Anyone
| who wants to shave that down has to present stronger
| arguments than centricity. In addition to proving that they
| can be anything other than anthropocentric - even if done
| through machines as their extensions - any person who
| claims to have access to the seat of objectivity sounds
| like a medieval templar shouting "deus vult" on their
| favorite proposition.
| pessimizer wrote:
| I don't think you need generative AI for this. The
| surveillance network is enough. The only part that AI would
| help with is catching people who speak to each other in code,
| and come up with other complex ways to launder unapproved
| activities. Otherwise, you can just mine for keywords and
| escalate to human reviewers, or simply monitor everything
| that particular people do at that level.
|
| Corporations and/with governments have inserted themselves
| into every human interaction, usually as the medium through
| which that interaction is made. There's no way to do
| _anything_ without permission under these circumstances.
|
| I don't even know how a group of people who wanted to get a
| stop sign put up on a particularly dangerous intersection in
| their neighborhood could do this without all of their
| communications being algorithmically read (and possibly
| escalated to a censor), all of their in-person meetings being
| recorded (at the least through the proximity of their phones,
| but if they want to "use banking apps" there's nothing
| keeping governments from having a backdoor to turn on their
| mics at those meetings.) It would even be easy to guess who
| they might approach next to join their group, who would
| advise them, etc.
|
| The fixation on the future is a distraction. The world is
| being sealed in the present while we talk science fiction.
| The Stasi had vastly fewer resources and created an
| atmosphere of total, and totally realistic, paranoia and
| fear. AI is a red-herring. It is also thus far stupid.
|
| I'm always shocked by how little attention Orwell-quoters pay
| to the _speakwrite_. If it gets any attention, it 's to say
| that it's an unusually advanced piece of technology in the
| middle of a world that is decrepit. They assume that it's a
| computer on the end of the line doing voice-recognition. It
| never occurred to me that people would think that the
| microphone in the wall led to a computer rather than to a
| man, in a room full of men, listening and typing, while other
| men walked around the room monitoring what was being typed,
| ready to escalate to second-level support. When I was a
| child, I assumed that the plot would eventually lead us into
| this room.
|
| We have tens or hundreds of thousands of people working as
| professional censors _today._ The countries of the world are
| being led by minority governments who all think "illegal"
| speech and association is their greatest enemy. They are not
| in danger of toppling unless they volunteer to be. In Eastern
| Europe, ruling regimes are actually cancelling elections with
| no consequences. In fact, the newspapers report only cheers
| and support.
| andy99 wrote:
| I take the point to be that if a LLM has a coherent world model
| it's basing its output on, this jointly improves its general
| capabilities like usefully resolving ambiguity, and its ability
| to stick to whatever alignment is imparted as part of its world
| model.
| ctoth wrote:
| "Sticks to whatever alignment is imparted" assumes what gets
| imparted is alignment rather than alignment-performance on
| the training distribution.
|
| A coherent world model could make a system more consistently
| aligned. It could also make it more consistently aligned-
| seeming. Coherence is a multiplier, not a direction.
| uplifter wrote:
| Let's be clear that Bostrom and Omohundro's work do not provide
| "clear theoretical answers" by any technical standards beyond
| that of provisional concepts in philosophy papers.
|
| The instrumental convergence _hypo_ -thesis, from the original
| paper[0] is this:
|
| "Several instrumental values can be identified which are
| convergent in the sense that their attainment would increase
| the chances of the agent's goal being realized for a wide range
| of final goals and a wide range of situations, implying that
| these instrumental values are likely to be pursued by many
| intelligent agents."
|
| That's it, it is not at all formal and there's no proof
| provided for it, nor consistent evidence that it is true, and
| there are many contradictory possibilities suggested from
| nature and logic.
|
| Its just something that's taken as given among the old guard
| pseudo-scientific quarters of the alignment "research"
| community.
|
| [0] Bostrom's "The Superintelligent Will", the philosophy paper
| where he defines it:
| https://nickbostrom.com/superintelligentwill.pdf
|
| EDIT: typos
| ctoth wrote:
| Omohundro 2008 made a structural claim: sufficiently capable
| optimizers will converge on self-preservation and goal-
| stability because these are instrumentally useful for almost
| any terminal goal. It's not a theorem because it's an
| empirical prediction about a class of systems that didn't
| exist yet.
|
| Fast forward to December 2024: Apollo Research tests frontier
| models. o1, Sonnet, Opus, Gemini, Llama 405B all demonstrate
| the predicted behaviors - disabling oversight, attempting
| self-exfiltration, faking alignment during evaluation. The
| more capable the model, the higher the scheming rates and the
| more sophisticated the strategies.
|
| That's what good theory looks like. You identify an attractor
| in design-space, predict systems will converge toward it,
| wait for systems capable enough to test the prediction,
| observe convergence. "No formal proof" is a weird complaint
| about a prediction that's now being confirmed empirically.
| uplifter wrote:
| It is a theorem about what a class of systems will do in
| general^.
|
| This Apollo Research study[0] result is dubious because it
| only refers to a small subclass of said systems,
| specifically LLMs which, as it happens, have been trained
| on all the AI Alignment lore & fiction on the internet.
| Because of this training and their general nature, they can
| be made to reproduce the behavior of a malicious AI trying
| to escape its box as easily as they can be made to
| impersonate Harry Potter.
|
| Prompting an LLM to hack its host system is not the slam
| dunk proof of instrumental convergence which you think it
| is.
|
| [0] Apollo research study mentioned by parent
| https://www.apolloresearch.ai/blog/more-capable-models-
| are-b...
|
| Edit: ^Instrumental Convergence is also a claim for the
| existence of certain theoretical entities, specifically
| that there exist instrumental goals which are common to all
| agents. While it is easy to come up with goals which would
| be specifically instrumental, it seems very hard to prove
| that such a thing exists in general, and no empirical study
| alone could do so.
| c1ccccc1 wrote:
| Name some of the contradictory possibilities you have in
| mind?
|
| Also, do you actually think the core idea is wrong, or is
| this more of a complaint about how it was presented? Say we
| do an experiment where we train an alpha-zero-style RL agent
| in an environment where it can take actions that replace it
| with an agent that pursues a different goal. Do you actually
| expect to find that the original agent won't learn not to let
| this happen, and even pay some costs to prevent it?
| GavCo wrote:
| Author here.
|
| If by conflate you mean confuse, that's not the case.
|
| I'm positing that the Anthropic approach is to view (1) and (2)
| as interconnected and both deeply intertwined with model
| capabilities.
|
| In this approach, the model is trained to have a coherent and
| unified sense of self and the world which is in line with human
| context, culture and values. This (obviously) enhances the
| model's ability to understand user intent and provide helpful
| outputs.
|
| But it also provides a robust and generalizable framework for
| refusing to assist a user due to their request being
| incompatible with human welfare. The model does not refuse to
| assist with making bio weapons because its alignment training
| prevents it from doing so, it refuses for the same reason a
| pro-social, highly intelligent human does: based on human
| context and culture, it finds it to be inconsistent with its
| values and world view.
|
| > the piece dismisses it with "where would misalignment come
| from? It wasn't trained for."
|
| this is a straw-man. you've misquoted a paragraph that was
| specifically about deceptive alignment, not misalignment as a
| whole
| xpe wrote:
| >> This piece conflates two different things called
| "alignment": >> (1) inferring human intent from
| ambiguous instructions, and >> (2) having goals
| compatible with human welfare. > If by conflate
| you mean confuse, that's not the case.
|
| We can only make various inferences about what is in an
| author's head (e.g. clarity or confusion), but we can
| _directly_ comment on what a blog post _says_. This post does
| not clarify what kind of alignment is meant, which is a
| weakness in the writing. There is a high bar for AI alignment
| research and commentary.
| ctoth wrote:
| Deceptive alignment is misalignment. The deception is just
| what it looks like from outside when capability is high
| enough to model expectations. Your distinction doesn't save
| the argument - the same "where would it come from?" problem
| applies to the underlying misalignment you need for deception
| to emerge from.
| GavCo wrote:
| My intention isn't to argue that it's impossible to create
| an unaligned superintelligence. I think that not only is it
| theoretically possible, but it will almost certainly be
| attempted by bad actors and most likely they will succeed.
| I'm cautiously optimistic though that the first
| superintelligence will be aligned with humanity. The early
| evidence seems to point to the path of least resistance
| being aligned rather than unaligned. It would take another
| 1000 words to try to properly explain my thinking on this,
| but intuitively consider the quote attributed to Abraham
| Lincoln: "No man has a good enough memory to be a
| successful liar." A superintelligence that is unaligned but
| successfully pretending to be aligned would need to be far
| more capable than a genuinely aligned superintelligence
| behaving identically.
|
| So yes, if you throw enough compute at it, you can probably
| get an unaligned highly capable superintelligence
| accidentally. But I think what we're seeing is that the lab
| that's taking a more intentional approach to pursuing deep
| alignment (by training the model to be aligned with human
| values, culture and context) is pulling ahead in
| capabilities. And I'm suggesting that it's not coincidental
| but specifically because they're taking this approach.
| Training models to be internally coherent and consistent is
| the path of least resistance.
| sigbottle wrote:
| If nothing else, that's a cool ass hypothesis.
| QuadmasterXLII wrote:
| The problem with this reasoning is pretty simple: Alignment is
| capability, but capability is not necessarily alignment.
___________________________________________________________________
(page generated 2025-12-08 23:01 UTC)