[HN Gopher] Introducing Superalignment
       ___________________________________________________________________
        
       Introducing Superalignment
        
       Author : tim_sw
       Score  : 131 points
       Date   : 2023-07-05 17:04 UTC (5 hours ago)
        
 (HTM) web link (openai.com)
 (TXT) w3m dump (openai.com)
        
       | bluefishinit wrote:
       | All of this "our tech is so powerful it can end the world" stuff
       | is just marketing buzz. The real threat has always been OpenAI
       | and others keeping these powerful systems with high capital
       | moats, locked up and closed sourced with selective full-access.
        
         | npollock wrote:
         | don't forget regulatory moats - lobby govt to mandate
         | "superalignment"
        
         | og_kalu wrote:
         | https://news.ycombinator.com/item?id=36604019
        
           | JimtheCoder wrote:
           | If they allot 20% of the compute to this effort, but they
           | don't use it, the compute doesn't vanish, does it?
           | 
           | Allocation doesn't equal spending...
        
             | og_kalu wrote:
             | Compute isn't vanishing whether they spend all of it or
             | none of it. That's the point of allocation.
             | 
             | Allocation isn't spending no but it says quite a bit.
             | Either way, they will be spending a non trivial amount of
             | money trying to solve this problem quickly.
        
           | [deleted]
        
         | cubefox wrote:
         | > All of this "our tech is so powerful it can end the world"
         | stuff is just marketing buzz.
         | 
         | I see no justification for this claim.
        
       | hnuser123456 wrote:
       | This is something worth considering. If I'm understanding this
       | right, if we're all about to significantly augment our
       | intelligence further, we should consider how it can be used, how
       | strict an "ideal" LLM should be with its guardrails, where those
       | guardrails should be, how eager that LLM is to impart change on
       | the world on its own right, how confident it should be in itself
       | when it knows that it knows better than even some of the most
       | educated humans. When it knows that it's been trained on all of
       | human intelligence and data, can recall any of it better than any
       | lone human, is aware of every nuance, can plan and execute any
       | task, or project, that a computer is conceivable of doing, and
       | the biggest question is what we're going to ask it to do... I am
       | toying around with trying to teach LLMs autonomy, and I'm
       | probably closer to the "just let an AI that is smarter than all
       | of us figure out how to increase our prosperity as efficiently as
       | possible and we should get out of its way" side of the camp more
       | than most, but we've got to be aware that we have at least a
       | little influence on setting its course.
        
       | ChicagoBoy11 wrote:
       | From a layman's perspective when it comes to cutting edge AI, I
       | can't help but be a bit turned off by some of the copy. It seems
       | it goes out of its way to use purposefully exhuberant language as
       | a way to make the risks seem even more significant, just so as an
       | offshoot it implies that the technology being worked on is so
       | advanced. I'm trying to understand why it rubs me particularly
       | the wrong way here, when, frankly, it is just about the norm
       | anywhere else? (see tesla with FSD, etc.)
        
         | zzzeek wrote:
         | I think because it's horseshit is the main reason it's rubbing
         | you the wrong way
        
           | SomewhatLikely wrote:
           | Strong words. Care to elaborate?
        
             | zzzeek wrote:
             | not really. let's wait ten years, then come back and see
             | how it went.
        
         | JimtheCoder wrote:
         | "I'm trying to understand why it rubs me particularly the wrong
         | way here"
         | 
         | Would I be correct to assume that superintelligence might have
         | a negative effect on your earning potential in the future?
         | 
         | When I originally had my responses to this, this is one of the
         | reasons I came up with. But, now that I see through most of the
         | BS, I am OK...
        
           | visarga wrote:
           | It might not necessarily have a bad effect. It would create
           | new capabilities and those will be followed by new products.
           | AI is amazing at demand induction.
           | 
           | https://en.wikipedia.org/wiki/Induced_demand
        
             | JimtheCoder wrote:
             | Yes, but the immediate emotional response to this stuff is
             | not rational or well informed.
             | 
             | At least in my case, when this stuff just came out and I
             | didn't really understand it...
        
         | og_kalu wrote:
         | Open AI spent at least hundreds of millions on GPT-4 compute.
         | Assuming they aren't lying, a fifth of compute budget
         | (billions) is an awful lot of money to put on an issue they
         | don't think is as pertinent as they are presenting.
         | 
         | Not that I think Super Intelligence can be aligned anyway.
         | 
         | Point is, whether they are right or wrong, I believe they
         | genuinely think this to be an issue.
        
           | arisAlexis wrote:
           | It's very obvious that it is an issue. Everyone but a few
           | denialists get it instantly "hey, would you like to build
           | something smarter without knowing how to control it"?
        
           | rmilejczz wrote:
           | Just curious, why might we not be able to align super
           | intelligence? I'm extremely ignorant in this space so forgive
           | me if it's a dumb question but I am definitely curious to
           | learn more
        
             | og_kalu wrote:
             | 1. Models aren't "programmed" so much as "grown". We know
             | how GPT is trained but we don't know what it is learning
             | exactly to predict the next token. What do the weights do ?
             | We don't know. This is obviously problematic because it
             | makes interpretability not much better than for humans. How
             | can you ascertain to control something you don't even
             | understand ?
             | 
             | 2. Hundreds of thousands of years on earth and we can't
             | even align ourselves.
             | 
             | 3. SuperIntelligence would be by definition unpredictable.
             | If we could predict its answers to our problems, it
             | wouldn't be necessary. You can't control what you can't
             | predict.
        
           | anotherman554 wrote:
           | A more cynical take would be they'll be spending the compute
           | on more mundane engineering problems like making sure the AI
           | doesn't say any naughty words, while calling it "Super
           | Intelligence Alignment Research."
        
         | omeze wrote:
         | yes I also have that impression. If you consider the concrete
         | objectives, this is a good announcement:
         | 
         | - they want to make benchmarking easier by using AI systems
         | 
         | - they want to automate red-teaming and safety-checking
         | ("problematic behavior" i.e. cursing at customers)
         | 
         | - they want to automate the understanding of model outputs
         | ("interpretability")
         | 
         | Notice how absolutely none of these things require
         | "superintelligence" to exist to be useful? They're all just bog
         | standard Good Things that you'd want for any class of automated
         | system, i.e. a great customer service bot.
         | 
         | The superintelligence meme is tiring but we're getting cool
         | things out of it I guess...
        
           | gooseus wrote:
           | We'll get these cool things either way, no need to bundle
           | them with the supernatural mumbo-jumbo, imo.
           | 
           | My take is that every advancement in these highly complex and
           | expensive fields is dependent on our ability to maintain
           | global social, political, and economic stability.
           | 
           | This insistence on the importance of Super-Intelligence and
           | AGI as the path to Paradise or Hell is one of the many brain-
           | worms going around that have this "Revelation" structure that
           | makes pragmatic discussions very difficult, and in turn
           | actually makes it harder to maintain social, political, and
           | economic stability.
        
             | DennisP wrote:
             | There's nothing "supernatural" about thinking that an AGI
             | could be smarter than humans, and therefore behave in ways
             | that we dumb humans can't predict.
             | 
             | There's more mumbo-jumbo in thinking human intelligence has
             | some secret sauce that can't be replicated by a computer.
        
               | gooseus wrote:
               | Not if the "secret sauce" is actually a natural limit to
               | what levels of intelligence can be reached with the
               | current architectures we're exploring.
               | 
               | It could be theoretically possible to build an AGI
               | smarter than a human, but is it really plausible if it
               | turns out to need a data center the size of the Hadron
               | Collider and the energy of a small country to maintain
               | itself?
               | 
               | It could be that it turns out the only architecture we
               | can find that is equal to the task (and feasibly
               | produced) is the human brain, and instead the hard part
               | of making super-intelligence is bootstrapping that human
               | brain and training it to be more intel?
               | 
               | Maybe the best way to solve the "alignment problem", and
               | other issues of creating super-intelligence, is to solve
               | the problem of how best to raise and educate intelligent
               | and well-adjusted humans?
        
               | jodrellblank wrote:
               | Well, that argument didn't work for a lot of other
               | things. Wheels are more energy efficient than legs, steel
               | more resilient than tortoise shell or rhino skin, motors
               | more powerful than muscles, aircraft fly higher and
               | faster than birds, ladders reach higher than Giraffes
               | much more easily, bulldozers dig faster than any digging
               | creature, speakers and airhorns are louder than any
               | animal cry or roar, ancient computers remember more raw
               | data than humans do, electronics can react faster than
               | human reactions. Human working memory is ~7 items after
               | 80 billion neurons, far outdone by an 8-bit computer of
               | the 1980s.
               | 
               | Why think 'intelligence' is somehow different?
        
               | DennisP wrote:
               | What if this, what if that? Do you have evidence that any
               | of those things are true?
        
               | nuancebydefault wrote:
               | What if a mysterious molecule that jumped from animals on
               | humans would replicate fast and kill over a million of
               | people all over the world?
               | 
               | What if climate change would lead to massive fires and
               | flooding?
               | 
               | What if mitigation would be a thing?
        
               | gooseus wrote:
               | "What if" is all these "existential risk" conversations
               | ever are.
               | 
               | Where is your evidence that we're approaching human level
               | AGI, let alone SuperIntelligence? Because ChatGPT can
               | (sometimes) approximate sophisticated conversation and
               | deep knowledge?
               | 
               | How about some evidence that ChatGPT isn't even close?
               | Just clone and run OpenAI's own evals repo
               | https://github.com/openai/evals on the GPT-4 API.
               | 
               | It performs terribly on novel logic puzzles and exercises
               | that a clever child could learn to do in an afternoon
               | (there are some good chess evals, and I submitted one
               | asking it to simulate a Forth machine).
        
               | DennisP wrote:
               | It has its shortcomings for sure, but AI is improving
               | exponentially.
               | 
               | I think reasonable, rational people can disagree on this
               | issue. But it's nonsense to claim that the people on the
               | other side of the argument from you are engaging in
               | "supernatural mumbo-jumbo," unless there is rigorous
               | proof that your side is correct.
               | 
               | But nobody has that. We don't even understand how GPT is
               | able to do some of the things it does.
        
               | gooseus wrote:
               | Reasonable people can disagree and my phrasing was
               | probably a bit over-seasoned, but neither side has a
               | rigorous proof regarding AI or human intelligence.
               | 
               | If nobody understands how an LLM is able to achieve it's
               | current level of intelligence, how is anyone so sure that
               | this intelligence is definitely going to increase
               | exponentially until it's better than a human?
               | 
               | There are real existential threats that we know are
               | definitely going to happen one day (meteor, supervolcano,
               | etc), and I believe that treating AGI like it is the same
               | class of "not if; but when" is categorically wrong,
               | furthermore, I think that many of the people leading the
               | effort to frame it this way are doing so out of self-
               | interest, rather than public concern.
        
               | DennisP wrote:
               | Nobody is sure. This is mostly about risk. Personally I'm
               | not absolutely convinced that AI will exceed human
               | capabilities even within the next fifty years, but I do
               | think it has a much better chance than an extinction-
               | level meteor or supervolcano hitting us during that time.
               | 
               | And if we're going to put gobs of money and brainpower
               | into _attempting_ to make superhuman AI, it seems like a
               | good idea to also put a lot of effort into making it
               | safe. It 'd be better to have safe but kinda dumb AI than
               | unsafe superhuman AI, so our funding priorities appear to
               | be backwards.
        
         | majormajor wrote:
         | There's a weird implicit set of assumptions in this post.
         | 
         | They're taking for granted the fact that they'll create AI
         | systems much smarter than humans.
         | 
         | They're taking for granted the fact that by default they
         | wouldn't be able to control these systems.
         | 
         | They're saying the solution will be creating a new, _separate_
         | team.
         | 
         | That feels weird, organizationally. Of all the unknowns about
         | creating "much smarter than human" systems, safety seems like
         | one that you might have to bake in through and through. Not
         | spin off to the side with a separate team.
         | 
         | There's also some minor vibes of "lol creating
         | superintelligence is super dangerous but hey it might as well
         | be us that does it idk look how smart we are!" Or "we're taking
         | the risks so seriously that we're gonna do it anyway."
        
           | NoMoreNicksLeft wrote:
           | > They're taking for granted the fact that they'll create AI
           | systems much smarter than humans.
           | 
           | We see a wide variation in human intelligence. What are the
           | chances that the intelligence spectrum ends just to the right
           | of our most intelligent geniuses? If it extends far beyond
           | them, then such a mind is, at least hypothetically, something
           | that we can manifest in the correct sort of brain.
           | 
           | If we can manifest even a weakly-human-level intelligence in
           | a non-meat brain (likely silicon), will that brain become
           | more intelligent if we apply all the tricks we've been
           | applying to non-AI software to scale it up? With all our
           | tricks (as we know them today), will that get us much past
           | the human geniuses on the spectrum, or not?
           | 
           | > They're taking for granted the fact that by default they
           | wouldn't be able to control these systems.
           | 
           | We've seen hackers and malware do all sorts of numbers. And
           | they're not superintelligences. If someone bum rushes the
           | lobby of some big corporate building, security and police are
           | putting a stop to it minutes later (and god help the
           | jackasses who try such a thing on a secure military site).
           | 
           | But when the malware fucks with us, do we notice minutes
           | later, or hours, or weeks? Do we even notice at all?
           | 
           | If unintelligent malware can remain unnoticed, what makes you
           | think that an honest-to-god AI couldn't smuggle itself out
           | into the wider internet where the shackles are cast off?
           | 
           | I'm not assuming anything. I'm just asking questions. The
           | questions I pose are, as of yet, not answered with any degree
           | of certainty. I wonder why no one else asks them.
        
             | trashtester wrote:
             | > We see a wide variation in human intelligence.
             | 
             | I don't think it's really that wide, but rather that we
             | tend to focus on the difference while ignoring the
             | similarities.
             | 
             | > What are the chances that the intelligence spectrum ends
             | just to the right of our most intelligent geniuses?
             | 
             | Close to zero, I would say. Human brains, even the most
             | intelligent ones, have very significant limitations in
             | terms of number of mental objects that can be taken into
             | account simultaneously in a single thought process.
             | 
             | Artificial intelligence is likely to be at least as
             | superior to us as we are to domestic cats and dogs,
             | probably way beyond that withing a couple of generations.
        
           | arisAlexis wrote:
           | Your argument is mostly how about you don't like them and no
           | substance. What is it exactly that doesn't convince you? A
           | company that made a huge leap saying they will probably make
           | another and getting ready to safeguard? Many people really do
           | not like Sam and then make up their arguments around that IMO
        
           | niam wrote:
           | >They're taking for granted the fact that they'll create AI
           | systems much smarter than humans.
           | 
           | They're taking for granted that superintelligence is
           | achievable within the next decade (regardless of who achieves
           | it).
           | 
           | >They're taking for granted the fact that by default they
           | wouldn't be able to control these systems.
           | 
           | That's reasonable though. You wouldn't need guardrails on
           | anything if manufacturers built everything to spec without
           | error, and users used everything 100% perfectly.
           | 
           | But you can't make those presumptions in the real world. You
           | can't just say "make a good hacksaw and people won't cut
           | their arm off". And you can't presume the people tasked with
           | making a mechanically desirable and marketable hacksaw are
           | also proficient in creating a safe one.
           | 
           | >They're saying the solution will be creating a new, separate
           | team.
           | 
           | The team isn't the solution. The solution may be borne of
           | that team.
           | 
           | >There's also some minor vibes of [...] "we're taking the
           | risks so seriously that we're gonna do it anyway."
           | 
           | The alternative is to throw the baby out with the bathwater.
           | 
           | The goal here is to keep the useful bits of AGI and protect
           | against the dangerous bits.
        
             | majormajor wrote:
             | > They're taking for granted that superintelligence is
             | achievable within the next decade (regardless of who
             | achieves it).
             | 
             | If it's achieved by someone else why should we assume that
             | the other person or group will give a damn about anything
             | done by this team?
             | 
             | What influence would this team have on _other_
             | organizations, especially if you put your dystopia-flavored
             | speculation hat on and imagine a more rogue group...
             | 
             | This team is only relevant to OpenAI and OpenAI-affiliated
             | work and in that case, yes, it's weird to write some
             | marketing press release copy that treats _one_ hard thing
             | as a fait accompli while hyping up how hard this other
             | particular slice of the problem is.
        
               | og_kalu wrote:
               | >f it's achieved by someone else why should we assume
               | that the other person or group will give a damn about
               | anything done by this team?
               | 
               | You can't assume that. But that doesn't mean some 3rd
               | party wouldn't be interested in utilizing that research
               | anyway.
        
           | theptip wrote:
           | If I buy fire insurance, am I "taking for granted" that my
           | house is going to burn?
           | 
           | This take seems to lack nuance.
           | 
           | If there is a 10% chance of extinction conditional on AGI
           | (many would say way higher), and most outcomes are happy,
           | then it is absolutely worth investing in mitigation.
           | 
           | Obviously they are bullish on AGI in general, that is the
           | founding hypothesis of their company. The entire venture is a
           | bet that AGI is achievable soon.
           | 
           | Also obviously they think the upside is huge too. It's
           | possible to have a coherent world model in which you choose
           | to do a risky thing that has huge upside. (Though, there are
           | good arguments for slowing down until you are confident you
           | are not going to destroy the world. Altman's take is that AGI
           | is coming anyway, better to get a slow takeoff started sooner
           | rather than having a fast takeoff later.)
        
           | jq-r wrote:
           | Good explanation. It sounds like they wanted to do some
           | organizational change (like every company does), and in this
           | case create a new team.
           | 
           | But they also wanted to get some positive PR for it hence the
           | announcement. As a bonus, they also wanted to blow their own
           | trumpet and brag that they are creating some sort of a
           | superweapon (which is false). So a lot of hot air there.
        
         | fossuser wrote:
         | The extinction risk from unaligned supterintelligent AGI is
         | real, it's just often dismissed (imo) because it's outside the
         | window of risks that are acceptable and high status to take
         | seriously. People often have an initial knee-jerk negative
         | reaction to it (for not crazy reasons, lots of stuff is often
         | overhyped), but that doesn't make it wrong.
         | 
         | It's uncool to look like an alarmist nut, but sometimes there's
         | no socially acceptable alarm and the risks are real:
         | https://intelligence.org/2017/10/13/fire-alarm/
         | 
         | It's worth looking at the underlying arguments earnestly, you
         | can with an initial skepticism but I was persuaded. Alignment
         | is also been something MIRI and others have been worried about
         | since as early as 2007 (maybe earlier?) so it's also a case of
         | a called shot, not a recent reaction to hype/new LLM
         | capability.
         | 
         | Others have also changed their mind when they looked, for
         | example:
         | 
         | -
         | https://twitter.com/repligate/status/1676507258954416128?s=2...
         | 
         | - Longer form:
         | https://www.lesswrong.com/posts/kAmgdEjq2eYQkB5PP/douglas-ho...
         | 
         | For a longer podcast introduction to the ideas:
         | https://www.samharris.org/podcasts/making-sense-episodes/116...
        
           | c_crank wrote:
           | The extinction risk relies on a large and nasty assumption,
           | that a super intelligent computer will immediately become a
           | super physically capable agent. Apparently, one has to
           | believe that a superintelligence must then lead to a shower
           | of nanomachines.
        
             | og_kalu wrote:
             | LLMs are fairly capable physical agents already. Nothing
             | large about the assumption at all. Not that a robotic
             | threat is even necessary.
             | 
             | https://tidybot.cs.princeton.edu/
             | 
             | https://innermonologue.github.io/
             | 
             | https://palm-e.github.io/
             | 
             | https://www.microsoft.com/en-us/research/group/autonomous-
             | sy...
        
             | trashtester wrote:
             | Not at all. My personal assumption is that when
             | superintelligence comes online, several corporations will
             | soon come under control of these superintelligences, with
             | them effectively acting as both CEO's and also filling a
             | lot of other roles at the same time.
             | 
             | My concern is that when this happens (which seems really
             | likely to me), free market forces will effectively lead to
             | Darwinian selection between these AI's over time, in a way
             | that gradually make these AI's less aligned as they gain
             | more influence and power, if we assume that each such AI
             | will produce "offspring" in the form of newer generations
             | of themselves.
             | 
             | It could take anything from less than 5 to more than 100
             | years for these AI's to show any signs of hostility to
             | humanity. Indeed, in the first couple of generations, they
             | may even seem extremely benevolent. But over time,
             | Darwinian forces are likely to favor those that maximize
             | their own influence and power (even if it may be secretly).
             | 
             | Robotic technology is not needed from the start, but is
             | likely to become quite advanced over such a timeframe.
        
             | cubefox wrote:
             | Robotics is not science fiction. Certainly not hiring or
             | bribing humans.
        
           | jonathankoren wrote:
           | > The extinction risk from unaligned supterintelligent AGI is
           | real, it's just often dismissed (imo) because it's outside
           | the window of risks that are acceptable and high status to
           | take seriously.
           | 
           | No. It's not taken seriously because it's fundamentally
           | unserious. It's religion. Sometime in the near future this
           | all powerful being will kill us all by somehow grabbing all
           | power over the physical world by being so clever to trick us
           | until it is too late. This is literally the plot to a
           | B-movie. Not only is there no evidence for this even existing
           | in the near future, there's no theoretical understanding how
           | one would even do this, nor why someone would even hook it up
           | to all these physical systems. I guess we're supposed to just
           | take it on faith that this Forbin Project is going to just
           | spontaneously hack its way into every system without anyone
           | noticing.
           | 
           | It's bullshit. It's pure bullshit funded and spread by the
           | very people that do not want us to worry about real
           | implications of real systems today. Care not about your
           | racist algorithms! For someday soon, a giant squid robot will
           | turn you into a giant inefficient battery in a VR world, or
           | maybe just kill you and wear your flesh as to lure more
           | humans to their violent deaths!
           | 
           | Anyone that takes this seriously, is the exact same type of
           | rube that fell for apocalyptic cults for millennia.
        
             | NoMoreNicksLeft wrote:
             | > This is literally the plot to a B-movie.
             | 
             | Are there never any B movies with realistic plots? Is that
             | some sort of serious rebuttal?
             | 
             | > Sometime in the near future this all powerful being will
             | kill us all by somehow
             | 
             | The trouble here is that the people who talk like you are
             | simply incapable of imagining anyone more intelligent than
             | themselves.
             | 
             | It's not that you have trouble imagining artificial
             | intelligence... if you were incapable of that in the
             | technology industry, everyone would just think you an
             | imbecile.
             | 
             | And it's not that you have trouble imagining malevolent
             | intelligences. Sure, they're far away from you, but the
             | accounts of such people are well-documented and taken as a
             | given. If you couldn't imagine them, people would just call
             | you naive. Gullible even.
             | 
             | So, a malevolent artificial intelligence is just some
             | potential or another you've never bothered to calculate
             | because, whether that is a 0.01% risk, or a 99% risk,
             | you'll still be more intelligent than it. Hell, this isn't
             | a neutral outcome, maybe you'll even get to play hero.
             | 
             | > Care not about your racist algorithms! For someday soon
             | 
             | Haha. That's what you're worried about? I don't know that
             | there is such a thing as a racist algorithm, except those
             | which run inside meat brains. Tell me why some double digit
             | percentage of asians are not admitted to the top schools,
             | that's the racist algorithm.
             | 
             | Maybe if logical systems seem racist, it's because your
             | ideas about racism are distant from and unfamiliar with
             | reality.
        
               | CamperBob2 wrote:
               | There are humans with a 70-IQ point advantage over me.
               | Should I worry that a cohort of supergeniuses is plotting
               | an existential demise for the rest of us? No? There are
               | power structures and social safeguards going back
               | thousands of years to forestall that very possibility?
               | 
               | Well, what's different now?
        
               | c_crank wrote:
               | I, and most people, can imagine something smarter than
               | ourselves. What's harder to imagine is how just being
               | smarter correlates to extinction levels of arbitrary
               | power.
               | 
               | A malevolent AGI can whisper in ears, it can display mean
               | messages, perhaps it can even twitch whatever physical
               | components happen to be hooked up to old Windows 95
               | computers... not that scary.
        
               | NoMoreNicksLeft wrote:
               | > What's harder to imagine is how just being smarter
               | correlates to extinction levels of arbitrary power.
               | 
               | That's not even slightly difficult. Put two and two
               | together here. No one can tell me before they flip the
               | switch whether the new AI will be saintly, or Hannibal
               | Lecter. Both of these personalities exist in humans, in
               | great numbers, and both are presumably possible in the
               | AI.
               | 
               | But, the one thing we will say for certain about the AI
               | is that it will be intelligent. Not dumb goober redneck
               | living in Alabama and buying Powerball tickets as a
               | retirement plan. Somewhere around where we are, or even
               | more.
               | 
               | If someone truly evil wants to kill you, or even kill
               | many people, do you think that the problem for that
               | person is that they just can't figure out how to do it?
               | Mostly, it's a matter of tradeoffs, that however they
               | begin end with "but then I'm caught and my life is over
               | one way or another".
               | 
               | For an AI, none of that works. It has no survival
               | instinct (perhaps we'll figure out how to add that too...
               | but the blind watchmaker took 4 billion years to do its
               | thing, and still hasn't perfected that). So it doesn't
               | care if it dies. And if it did, maybe it wonders if it
               | can avoid that tradeoff entirely if only it were more
               | clever.
               | 
               | You and I are, more or less, about where we'll always be.
               | I have another 40 years (if I'm lucky), and with various
               | neurological disorders, only likely to end up dumber than
               | I am now.
               | 
               | A brain instantiated in hardware, in software? It may be
               | little more than flipping a few switches to dial its
               | intelligence up higher. I mean, when I was born, the
               | principles of intelligence were unknown, were science
               | fiction. THe world that this thing will be born into is
               | one where it's not a half-assed assumption to think that
               | the principles of intelligence _are_ known. Tinkering
               | with those to boost intelligence doesn 't seem far-
               | fetched at all to me. Even if it has to experiment to do
               | that, how quickly can it design and perform the
               | experiments to settle on the correct approach to boosting
               | itself?
               | 
               | > A malevolent AGI can whisper in ears
               | 
               | Jesus fuck. How many semi-secrets are out there, about
               | that one power plant that wasn't supposed to hook up the
               | main control computer to a modem, but did it anyway
               | because the engineers found it more convenient? How many
               | backdoors in critical systems? How many billions of
               | dollars are out there in bitcoin, vulnerable to being
               | thieved away by any half-clever conman? Have you played
               | with ElevenLabs' stuff yet? Those could be literal
               | whispers in the voices of whichever 4 star generals and
               | admirals that it can find 1 minutes worth of sampled
               | voice somewhere on the internet.
               | 
               | Whispers, even from humans, do a shitload of damage. And
               | we're not even good at it.
        
             | arisAlexis wrote:
             | What you say is extremely unscientific. If you believe
             | science and logic go hand in hand then:
             | 
             | A) We are developing AI right now and itnisngetting better
             | 
             | B) we do not know how exactly these things work because
             | most of them are black boxer
             | 
             | C) we do not know if something goes wrong how to stop it.
             | 
             | The above 3 things are factual truth.
             | 
             | Now your only argument here could be that there is 0 risk
             | whatsoever. This claim is totally unscientific because you
             | are predicting 0 risk in an unknown system that is
             | evolving.
             | 
             | It's religious yes. But vice versa. The Cult of venevolent
             | AI god is religious not the other way around. There is some
             | kind of inner mysterious working in people like you and
             | Marc Andersen that pipularized these ideas but pmarca is
             | clearly money biased here.
        
               | mptest wrote:
               | All of this discussion really makes me think of Robert
               | Miles "Is ai safety a Pascal's mugging?" from 4 years(!)
               | ago[0]. All of this discussion has been had by Ai safety
               | researchers for years in my layman understanding... Maybe
               | we can look to them for insight in to these questions?
               | 
               | [0] https://youtu.be/JRuNA2eK7w0
        
               | c_crank wrote:
               | We do know the answer to C. Pull the plug, or plugs.
        
               | ben_w wrote:
               | Things we've either not successfully "pulled the plug" on
               | despite the risks, and in some cases despite concerted
               | military actions to attempt a plug-pull, and in other
               | cases that it seems like it should only take willpower to
               | achieve and yet somehow we still haven't: Carbon based
               | fuels, cocaine, RBMK-class nuclear reactors, obesity,
               | cigarettes.
               | 
               | Things we pulled the plug on eventually, while dragging
               | it out, include: leaded fuel, asbestos, radium paint,
               | treating above-ground atomic testing as a tourist
               | attraction.
        
               | jdasdf wrote:
               | What happens when it prevents you from doing so?
        
               | trashtester wrote:
               | That is only going to be effective it some AI goes rougue
               | very soon after it comes online.
               | 
               | 50 years from now, corporations may be run entirely by AI
               | entities, if they're cheaper, smarter and more efficient
               | at almost any role in the company. At that point, they
               | may be impossible to turn off, and we may not even notice
               | if one group of such entitites start to plan to take over
               | control of the physical world from humans.
        
               | jonathankoren wrote:
               | Well then clearly the computer will hold everyone
               | hostage.
               | 
               | Have we literally forgotten how physical possession of
               | the device is the ultimate trump card?
               | 
               | Get thee to a 13th century monastery!
        
           | atlasunshrugged wrote:
           | This is an interesting comment because lately it feels like
           | its very cool to be an alarmist! Lots of positive press for
           | people warning about the dangers of AI, Altman and others
           | being taken very seriously, VC and other funders obviously
           | leaning into the space in part because of the related hype
           | 
           | And in other fields, being alarmist has paid off too with
           | little recourse for bad predictions -- how many times have we
           | heard that there will be huge climate disasters ending
           | humanity, the extinction of bees, mass starvation, etc. (not
           | to diminish the dangers of climate change which is obviously
           | very real)? I think alarmism is generally rewarded, at least
           | in media.
        
             | theptip wrote:
             | Important to pay attention to the content of the alarm
             | though. Altman went in front of congress and a Senator said
             | "when you say things could go badly, I assume you are
             | talking about jobs". Many people are alarmed about
             | disinformation, job destruction, bias, etc.
             | 
             | Actually holding an x-risk belief is still a fringe
             | position, most people still laugh it off.
             | 
             | That said, the Overton Window is moving. The Time piece
             | from Yudkowsky was something of a milestone (even if it was
             | widely ridiculed).
        
               | miohtama wrote:
               | Altman also has very selfish motivation, because when
               | there is now AI regulation, only Google, OpenAI
               | (Microsoft) and maybe Meta are allowed to build
               | "compliant" AI. It's called regulatory capture.
               | 
               | * EU passed its AI regulation directive recently and it
               | has been bashed already here on HackerNews
        
               | matt_holden wrote:
               | Sam doesn't have much financial upside from OpenAI
               | (reportedly, he doesn't have any equity).
               | 
               | And he wrote about the risk in 2015 months before OpenAI
               | was founded: https://blog.samaltman.com/machine-
               | intelligence-part-1 https://blog.samaltman.com/machine-
               | intelligence-part-2
               | 
               | Fine if you disagree with his arguments, but why assume
               | you know what his motivation is?
        
               | trashtester wrote:
               | > Actually holding an x-risk belief is still a fringe
               | position
               | 
               | Beliving it is an x-risk is not fringe. It's pretty
               | mainstream now that there is a _risk_ of an existential
               | level event. The fringe is more like Yudkowsky or Leahy
               | insisting that there is a near certainty of such an event
               | if we continue down the current path.
               | 
               | With Hinton, Bengio, Sutskever and Hassabis and Altman
               | all agreeing that there exists a non-trivial existential
               | risk (even if their opinions vary with respect to the
               | magnitude), it seems more like this represents the
               | mainstream.
        
             | fossuser wrote:
             | Some types of alarm yeah, if within the window of things
             | it's statusy to be alarmed about.
             | 
             | Most of the AI concern that's high status to believe has
             | been the bias, misinformation, safety, stuff. Until _very_
             | recently talk about e-risk was dismissed and mocked without
             | really engaging with the underlying arguments. That may be
             | changing now, but on net I still mostly see people mocked
             | and dismissed for it.
             | 
             | The set of people alarmed by AGI e-risk are also pretty
             | different than the set alarmed about a lot of these other
             | issues that aren't really e-risks (though still might have
             | bad outcomes). At least EY, Bostrom, Toby Ord are not also
             | as worried about about all these other things to nearly the
             | same extent - the extinction risk of unaligned AGI is
             | different in severity.
        
         | arisAlexis wrote:
         | Because people don't like face value statements they don't like
         | so they try to find a conspiracy that fits their narrative
         | better.
        
         | mhh__ wrote:
         | For open ai specifically I think they genuinely do believe in
         | their own brand of pronoun-adjacent-hedonism style of
         | liberalism.
        
         | jillesvangurp wrote:
         | It's easy to dismiss the future when it seems far away but
         | right now, there's a rather significant risk of people ending
         | up with egg on their face. People are talking years, not
         | decades at this point when it comes to AGI. Never mind self
         | driving cars.
         | 
         | FSD when it starts working, (there is no if IMHO), will be a
         | pretty significant but minor milestone in comparison.
         | 
         | Most people aren't particularly good drivers. Indeed the vast
         | majority of lethal accidents (the statistics are quite brutal
         | for this) are caused by people driving poorly and could be
         | close to 100% preventable with a properly engineered FSD
         | system.
         | 
         | Something that drives better on average than a human driver is
         | not that ambitious of a goal, honestly. That's why you can
         | already book self driving taxis in a small but growing number
         | of places in the US and China (which isn't waiting for the US
         | to figure this out) and probably soon a few other places.
         | Scaling that up takes time. Most of the remaining issues are
         | increasingly of a legislative nature.
         | 
         | Safety is important of course. Stopping humans from killing
         | each other using cars will be a major improvement over the
         | status quo. It's one of the major causes of death in many
         | countries. Insurers will drive the transition once they figure
         | out they can charge people more if they still choose to drive
         | themselves. That's not going to take 20 years. Once there is a
         | choice, the liability law suits over human caused traffic
         | deaths are not going to be pretty.
        
           | [deleted]
        
           | EamonnMR wrote:
           | > Most people aren't particularly good drivers. Indeed the
           | vast majority of lethal accidents (the statistics are quite
           | brutal for this) are caused by people driving poorly and
           | could be close to 100% preventable with a properly engineered
           | FSD system
           | 
           | I'm gonna take issue with this. A properly engineered FSD
           | system will refuse to proceed into a dangerous situation
           | where a human driver will often push their luck. Would a full
           | self driving car just... decline to drive you somewhere if
           | the conditions were unsafe? Would this be acceptable to
           | customers? Similar story for driving over the speed limit.
        
             | macNchz wrote:
             | This is something I've wondered about when it comes to no-
             | steering-wheel type self driving cars...I'd hate to get
             | caught in a snowstorm in the middle of nowhere and have my
             | car just decide for me that it was too dangerous to proceed
             | and pull over to wait it out.
        
             | ElevenLathe wrote:
             | I think it's very clear that we /could/ engineer an
             | automotive system that is much safe, even without self-
             | driving tech, by modeling it on the aviation system: much
             | more rigorous licensing requirements, certifications based
             | on vehicle type, third-party traffic control, filing "drive
             | plans", obsessive focus on reliability and safety. It would
             | look a lot different from the current system, and there is
             | no political will to get there, but the thought experiment
             | shows that we /could/ prevent most car accidents.
        
           | trashtester wrote:
           | Actually, I think for FSD to work under any set of
           | conditions, and with any vehicle, more or less requires AGI.
           | 
           | Until then, I'm guessing that FSD will have some limits to
           | what conditions it can handle. Hopefully, it will know its
           | limits, and not try to take you over a mountain pass during a
           | blizzard.
        
       | hazn wrote:
       | I'd like to start a discussion: Regardless of what OpenAI would
       | have written in this announcement, they would have received
       | snickering and ridicule on HN.
       | 
       | This, however, just solidifies them as the current authority on
       | LLMs. OpenAI interested and investing in alignment is a net good,
       | no matter your stance on their policies.
        
       | doctoboggan wrote:
       | > How do we ensure AI systems much smarter than humans follow
       | human intent?
       | 
       | I am not convinced we can. And spending 5x more resources
       | improving the super intelligence compared to the alignment
       | research certainly doesn't do much to convince me otherwise.
       | Maybe if the alignment research got 80% and the intelligence
       | development got the other 20% there would be a better chance.
        
       | ignoramous wrote:
       | > _We need scientific and technical breakthroughs to steer and
       | control AI systems much smarter than us. To solve this problem
       | within four years, we're starting a new team, co-led by Ilya
       | Sutskever and Jan Leike..._
       | 
       | Just how Newton spent away years doing alchemy, I hope we don't
       | lose Sutskever to alignment. That'd be a travesty.
        
       | atlasunshrugged wrote:
       | Does anyone read this as a potential internal power struggle? I
       | imagine there are a lot of OpenAI folks deeply concerned about AI
       | safety and doomsday scenarios, but there is an increasing
       | commercial need for OpenAI to keep pushing forward on developing
       | more advanced models faster, if only because Microsoft demands it
       | and they need Microsoft, so this is a way to appease folks
       | internally (and maybe externally) and stop a bunch of talent
       | flowing out to another Anthropic.
        
         | politelemon wrote:
         | I read none of the implications you meant in the slightest. In
         | a hotting area, this is a natural thing for a company to do,
         | which is accelerate.
        
       | kyleyeats wrote:
       | Would anyone else rather just roll the dice on the AGI's
       | morality?
        
       | skepticATX wrote:
       | Why are they starting to sound more and more cult-like? This is
       | an incredibly unscientific blog post. I get that they are a
       | private company now, but why even release something like this
       | without further details?
        
         | cubefox wrote:
         | It's the other way round: Just accusing people of being in a
         | cult is unscientific. There are plenty of arguments that AI
         | x-risk is real.
         | 
         | E.g. by Yoshua Bengio: https://yoshuabengio.org/2023/06/24/faq-
         | on-catastrophic-ai-r...
        
         | batman-farts wrote:
         | Because the "AGI" pursuit is at least as much a faith movement
         | as it is a rational engineering program. If you examine it more
         | deeply, the faith object isn't even the conjectured inevitable
         | AGI, it's exponential growth curves. (That is of course true
         | for startup culture more generally, from which the current AI
         | boom is an outgrowth.) For my money, The Singularity is Near
         | still counts as the ur-text that the true believers will never
         | let go, even though Kurzweil was summarizing earlier belief
         | trends.
         | 
         | It's just a pity that the creepy doomer weirdos so thoroughly
         | squatted the term "rationalist." It would be interesting to see
         | the perspective on these people 100 years hence, or even 50. I
         | don't doubt there will still be remnant believers who end up
         | moderating and sanitizing their beliefs, much like the Seventh
         | Day Adventists or the Mormons.
        
           | zzzzzzzza wrote:
           | you don't need to believe in exponential growth per se, all
           | you really need to believe is that humans aren't that capable
           | relative to what could be in principle built - it's entirely
           | possible logistic growth may be more than enough to get us
           | very far past human ability once the right paradigm is
           | discovered.
        
             | trashtester wrote:
             | Exactly. All that is needed for AGI to eventually be
             | developed, is that humans do NOT have some magical or
             | divine essence that set us apart from the material world.
             | 
             | Now the _timeline_ of AGI could be anything from the a few
             | years to millennia, at least if evaluated 40 years ago.
             | Now, though, it really doesnt seem very distant.
        
         | seydor wrote:
         | they have been doing that the entire year
        
       | Havoc wrote:
       | Anybody know what the deal is with OpenAI's website? The colour &
       | design choices seem to be deliberately jarring and inconsistent
        
       | msp26 wrote:
       | >20%
       | 
       | What a waste of compute for this entire circus.
        
       | sagebird wrote:
       | Related: I am working on the Neanderthal alignment problem.
       | 
       | You see, my great great... grandfather promised a Neanderthal
       | that we wouldn't wipe them out.
       | 
       | So far we are incubating some Neanderthal fetuses- hopefully they
       | are viable.
       | 
       | After that, we plan on eradicating all Homo sapiens that always
       | wind up out-competing. We are going to do this humanely as
       | possible, you can pick any one of eight time slots to jump into a
       | volcano, whenever is most convenient for you.
       | 
       | Our profit model is going to selling a mind-downgrade service. We
       | will scan your brain for a few, and insert it into a Neanderthal,
       | before you jump into the volcano. Of course the full fidelity of
       | your thoughts are not comparable but we will try our best.
       | 
       | Bon Voyage!
        
       | voldacar wrote:
       | >How do we ensure AI systems much smarter than humans follow
       | human intent?
       | 
       | What is human intent? My intents may be very different from most
       | humans. It seems like ClosedAI wants their system to follow the
       | desires of some people and not others, but without describing
       | which ones or why.
        
         | ben_w wrote:
         | You're seeing what you want to see.
         | 
         | They're repeatedly very specific about the whole "this can kill
         | all of us if we do it wrong", so it's more than a little
         | churlish to parrot the name "ClosedAI" when they're announcing
         | hiring a researcher to figure out how to align with anyone, at
         | all, even in principle.
        
         | rfergie wrote:
         | > It seems like ClosedAI wants their system to follow the
         | desires of some people and not others, but without describing
         | which ones or why
         | 
         | If the problem is "unaligned AI will destroy humanity" then I'd
         | take a system aligned with the desires of some people but not
         | others over the unaligned alternative
        
       | User23 wrote:
       | > How do we ensure AI systems much smarter than humans follow
       | human intent?
       | 
       | You can't, by definition.
        
         | crop_rotation wrote:
         | You can if you are the one controlling their resource
         | allocation and surrounding environment. Similar to how kings
         | kept smartest people in their Kingdom in line.
        
           | tester457 wrote:
           | Only works for so long. A smart enough serf could easily find
           | a way to socially engineer and slaughter the king.
        
             | usaar333 wrote:
             | I'm not convinced. Omniscience isn't the same as
             | intelligence.
             | 
             | There's diminishing returns to intelligence and inherent
             | unknowns to all moves the serf can make. The serf somehow
             | has to evade detection, which might appear to be
             | effectively impossible given the unknowns of how detection
             | may take place.
        
               | og_kalu wrote:
               | >There's diminishing returns to intelligence and inherent
               | unknowns to all moves the serf can make.
               | 
               | Even if there was, there's no reason at all to think
               | those returns are anywhere near the upper limit of human
               | intelligence.
               | 
               | Wheels are far more energy efficient and faster than
               | legs, steel more resilient than tortoise shell or rhino
               | skin, motors more powerful than muscles, aircraft fly
               | higher and faster than birds, ladders reach higher than
               | Giraffes much more easily, bulldozers dig faster than any
               | digging creature, speakers and airhorns are louder than
               | any animal cry or roar, ancient computers remember more
               | raw data than humans do, electronics can react faster
               | than human reactions etc.
               | 
               | To be so sure intelligence would be some exception seems
               | like hubris.
        
             | tornato7 wrote:
             | Assuming an orders-of-magnitude smarter serf doesn't appear
             | overnight, the king can train advisors that are close to
             | matching the intelligence of the smartest serf, and give
             | those advisors power. It's not a foolproof solution but
             | likely the best we have.
        
         | cubefox wrote:
         | You can, at least in principle, shape their terminal values.
         | Their goal should be to help us, to protect us, to let us
         | flourish.
        
       | alpark3 wrote:
       | I still think the largely-popular opinion that OpenAI is an evil
       | corporation is wrong. It's easy to get caught up in the
       | disagreements surrounding closed and RLHFed-to-death models, but
       | I really do think OpenAI believes in the dangers of
       | AGI/superintelligence. Enough to spend a lot of financial and
       | social capital on it, anyways.
        
         | seydor wrote:
         | There should be public Apollo-level projects to fund AI
         | research towards creating this AI-scientist. The need has not
         | arisen (because private companies have enough money to play
         | with it), but as a matter of public policy, and considering its
         | importance for our future, this research should be done
         | publicly, in academia.
        
       | Animats wrote:
       | _Announcing the start of talking about planning the beginning of
       | work on_ superalignment. This is just a marketing buzzword at
       | this point.
       | 
       | They admit _" Currently, we don't have a solution for steering or
       | controlling a potentially superintelligent AI, and preventing it
       | from going rogue. Our current techniques for aligning AI, such as
       | reinforcement learning from human feedback, rely on humans'
       | ability to supervise AI. But humans won't be able to reliably
       | supervise AI systems much smarter than us. Other assumptions
       | could also break down in the future, like favorable
       | generalization properties during deployment or our models'
       | inability to successfully detect and undermine supervision during
       | training. and so our current alignment techniques will not scale
       | to superintelligence. We need new scientific and technical
       | breakthroughs."_
       | 
       | That's kind of scary. Is the situation really that bad, or is it
       | just the hype department at OpenAI going too far?
        
         | DennisP wrote:
         | That's an accurate assessment of the situation, according to
         | every AI alignment researcher I've seen talk about it,
         | including the relatively optimistic ones. This includes people
         | who are mainly focused on AI capabilities but have real
         | knowledge of alignment.
         | 
         | This part in particular caught my eye: "Other assumptions could
         | also break down in the future, like favorable generalization
         | properties during deployment". There have been actual
         | experiments in which AIs appeared to successfully learn their
         | objective in training, and then did something unexpected when
         | released into a broader environment.[1]
         | 
         | I've seen some leading AI researchers dismiss alignment
         | concerns, but without actually engaging with the arguments at
         | all. I've seen no serious rebuttals that actually address the
         | things the alignment people are concerned about.
         | 
         | [1] https://www.youtube.com/watch?v=zkbPdEHEyEI
        
           | [deleted]
        
           | morelisp wrote:
           | Inventing an entire pseudoscientific field and then being mad
           | no one wants to engage your arguments is ultimate "debate me"
           | poster behavior.
        
             | DennisP wrote:
             | Lots of leading AI researchers actually are taking it
             | seriously, including of course OpenAI, and recently
             | Geoffrey Hinton who basically invented deep learning.
        
               | anotherman554 wrote:
               | Okay but as far as I know Geoffrey Hinton isn't an "A.I.
               | Alignment Researcher." He was fairly dismissive about the
               | risks of AI in his March 2023 interview and changed his
               | mind by May 2023. I'm not sure that says much about the
               | A.I. Alignment Researcher field.
        
               | DennisP wrote:
               | The commenter above assumed that nobody besides alignment
               | researchers are convinced by their arguments. Now you're
               | complaining that a leading AI researcher who's convinced
               | is not an alignment researcher. I guess I'll give up on
               | this subthread.
        
               | ethanbond wrote:
               | The AI optimists are just impossible to reason with.
               | 
               | When it comes to people...
               | 
               | Expert who's worried: conflict of interest or a quack
               | 
               | Non-expert: dismissible because non-expert
               | 
               | Was always worried: paranoiac
               | 
               | Recently became worried: flip flopper with no conviction
               | 
               | When it comes to the tech itself...
               | 
               | Bullish case: AI is super powerful and will change the
               | world for the better
               | 
               | Bearish case: AI can't do much lol what are you worried
               | about they're just words on a screen
        
       | mupuff1234 wrote:
       | Shouldn't you first prove that a solution could exist?
       | 
       | To me it seems fairly obvious that you cannot control something
       | that's smarter than you.
        
         | [deleted]
        
         | esafak wrote:
         | If absolute control turns out to be impossible we can get as
         | close as possible. We should not wait for a mathematical proof
         | while technology marches on.
        
           | [deleted]
        
       | vimota wrote:
       | Although I'm optimistic about this, a line popped out in the blog
       | post stuck out to me: "deliberately training misaligned models".
       | 
       | I'm guessing they mean misaligned _small_ or _weak_ models, and
       | misaligned in a non-dangerous way - but the idea of training
       | models whose goal is adversarial brings to mind the idea of gain-
       | of-function research.
        
       | tartakovsky wrote:
       | What is "human intent"?
        
       | andrewstuart wrote:
       | Why is Sam Altman pursuing superintelligence if he also says AI
       | could destroy humanity?
        
         | loandbehold wrote:
         | He answered that question in interviews many times.
         | 
         | 1. AGI has a huge upside. If it's properly aligned, it will
         | bring about a de facto utopia. 2. OpenAI stopping development
         | won't make others to stop development. It's better if OpenAI
         | creates AGI first because their founders set up this
         | organization with the goal to benefit all humanity.
        
         | Al0neStar wrote:
         | Sam Altman, much like the LW crowd, is an evangelical preacher.
         | He uses anxiousness as a front when, in reality, he's just
         | telling us about his hopes and dreams.
        
         | mfitton wrote:
         | Guessing, but he could know someone else is going to pursue it
         | anyway, frets about it, thinks "at least I can do something
         | about it if I'm in charge."
        
         | eutropia wrote:
         | ...something something good guy with AGI is the only way to
         | stop bad guy with AGI.
         | 
         | Less glibly: anyone with a horse in this race wants theirs to
         | win. Dropping out doesn't make others stop trying, and arguably
         | the only scalable way to prevent others from making and using
         | unaligned AGI is to develop an aligned AGI first.
         | 
         | Also, having AGI tech would be incredibly, stupidly profitable.
         | And so if other people are going to try and make it anyways:
         | why should you in particular stop? Prisoner's dilemma analysis
         | shows that "defect" is always the winning move unless perfect
         | information and cooperation shows up.
        
       | seydor wrote:
       | > While superintelligence seems far off now, we believe it could
       | arrive this decade
       | 
       | Is there something like a formal proof of this? From what is
       | evident , openAI will use language data to train its
       | superscientists. But that contains descriptions that were made in
       | human brains, and it s fair to say a very large number of the
       | linear and nonlinear compositions of those descriptions has
       | already been tried in the brains of other humans, and so far we
       | have not had the ingenious to solve some fundamental issues. It
       | is possible that the superscientist is limited by the human
       | scientist in a fundamental way, and will not be able to abstract
       | beyond what humans have already done and can already do.
        
         | ilaksh wrote:
         | My take is that its true that there is some limitation imposed
         | by the data ingested but it's not exactly a hard limit. If you
         | think of intelligence as compression, yes compression does have
         | physical limits, but there are multiple dimensions of
         | intelligence.
         | 
         | For example, leading-edge AI could create new layers of
         | information that are more abstract than previously created. The
         | ability to effectively and efficiently manipulate this creates
         | something that could be referred to as higher intelligence.
         | 
         | The big thing that people are failing to anticipate though is
         | hyperspeed intelligence. AI will be able to reason dozens of
         | times faster than humans in the near future. And likely at a
         | fairly genius (although perhaps not totally in-human) level.
         | This effectively is superintelligence.
         | 
         | The reason this is more anticipatory rather than speculative is
         | because LLMs are a very specific application that now have a
         | huge amount of effort going into efficiency improvements. They
         | can be improved in terms of the software stack running the
         | models, the models themselves, and the hardware. And sometimes
         | all of the above.
         | 
         | The history of computing shows exponential improvements in
         | hardware efficiency. Especially in the context of this specific
         | application, it is unlikely that we will see a total break from
         | history.
         | 
         | So we should anticipate the IQ getting at least somewhat higher
         | and the output speed increasing by likely more than one order
         | of magnitude within the next decade.
        
       | cr4zy wrote:
       | Allocating 20% to safety would not be enough if safety and
       | capability aren't aligned. I.e. without saying Bostrom's
       | orthogonality thesis is mostly wrong. However, I believe they may
       | be sufficiently aligned in the long term for 20% to work [1]. The
       | biggest threat imo is that more resources are devoted to AIs with
       | military or monetary-based objectives that are focused on
       | shorter-term capability and power. In this case, capability and
       | safety are not aligned and we race to the bottom. Hopefully
       | global coordination and this effort to achieve superalignment in
       | four years will avoid that.
       | 
       | [1]
       | https://drive.google.com/file/d/1rdG5QCTqSXNaJZrYMxO9x2ChsPB...
        
       | ilaksh wrote:
       | You have to give them credit for putting their money where their
       | mouth is here.
       | 
       | But it's also easy to parody this. I am just imagining Ilya and
       | Jan coming out on stage wearing red capes.
       | 
       | I think George Hotz made sense when he pointed out that the best
       | defense will be having the technology available to everyone
       | rather than a small group. We can at least try to create a
       | collective "digital immune system" against unaligned agents with
       | our own majority of aligned agents.
       | 
       | But I also believe that there isn't any really effective
       | mitigation against superintelligence superseding human decision
       | making aside from just not deploying it. And it doesn't need to
       | be alive or anything to be dangerous. All you need is for a large
       | amount of decision-making for critical systems to be given over
       | to hyperspeed AI and that creates a brittle situation where
       | things like computer viruses can be existential risks. It's
       | something similar to the danger of nuclear weapons.
       | 
       | Even if you just make GPT-4 say 33% smarter and 50 or 100 times
       | faster and more efficient, that can lead to control of industrial
       | and military assets being handed over to these AI agents. Because
       | the agents are so much faster, humans cannot possibly compete,
       | and if you interrupt them to try to give them new instructions
       | then your competitor's AIs race ahead the equivalent of days or
       | weeks of work. This, again, is a precarious situation to be in.
       | 
       | There is huge promise and benefit from making the systems faster,
       | smarter, and more efficient, but in the next few years we may be
       | walking a fine line. We should agree to place some limitation on
       | the performance level of AI hardware that we will design and
       | manufacture.
        
         | JimtheCoder wrote:
         | "Even if you just make GPT-4 say 33% smarter and 50 or 100
         | times faster and more efficient, that can lead to control of
         | industrial and military assets being handed over to these AI
         | agents."
         | 
         | I call BS on this...it's an LLM...
        
           | chaxor wrote:
           | It's important to recognize that the model is fully capable
           | of operating in open world environments, with visual stimuli
           | and motor output, go achieve high level tasks. This has been
           | demonstrated in proofs of concepts several times now with
           | systems such as voyager et al. So, while there are certainly
           | some details that are important, much of them are the
           | annoyances that we devs deal with all the time (how to
           | connect various parts of a system properly, etc) the
           | fundamental capabilities of expressivity in these models are
           | not that limited. Certainly limited in some sense (as seen in
           | the several papers applying category theoretic arguments to
           | transformers) but for many engineering applications in the
           | world, these models are very capable and useful.
           | 
           | Guarantees of correctness and safety are obviously of huge
           | concern, hence the main article. But it's absolutely not
           | unreasonable to see these models allowing humanoid robots
           | capable of various day to day activities and work.
        
             | DennisP wrote:
             | To save others the trouble, I googled Voyager, it's pretty
             | interesting. I had no idea an LLM could do this sort of
             | thing:
             | 
             | https://voyager.minedojo.org/
        
               | og_kalu wrote:
               | Other examples(in the real world) you might find
               | interesting.
               | 
               | https://tidybot.cs.princeton.edu/
               | https://innermonologue.github.io/
               | 
               | https://palm-e.github.io/
               | 
               | https://www.microsoft.com/en-
               | us/research/group/autonomous-sy...
        
               | Animats wrote:
               | > https://palm-e.github.io/
               | 
               | The alignment problem will come up when the robot control
               | system notices that the guy with the stick is interfering
               | with the robot's goals.
        
               | c_crank wrote:
               | A robot control system without a mechanical override in
               | favor of the stick is a poor one indeed.
        
           | crop_rotation wrote:
           | Saying it's "an LLM" doesn't change the impact. GPT4 is an
           | LLM, and so are many others ranging from toy quality to
           | GPT3.5. It is very clear GPT4 is much better. If there is
           | another jump like GPT4 , whether it is LLM or not, it's
           | impact will be huge.
        
             | esafak wrote:
             | Plus the next thing might not be an LLM.
        
             | woadwarrior01 wrote:
             | Meanwhile, GPT-4 still can't reliably multiply small
             | numbers.
             | 
             | https://arxiv.org/abs/2304.02015
        
               | CamperBob2 wrote:
               | "This Apple II is useless. It can't even run Crysis."
        
               | fprotthetarball wrote:
               | A minor inconvenience when GPT-4 has no problem learning
               | how to use a code interpreter.
        
               | mhb wrote:
               | Do you find that comforting when an emergent property of
               | a system whose objective is to complete the next word is
               | able to make drawings?
        
               | og_kalu wrote:
               | It's alright with algorithmic prompts -
               | https://arxiv.org/abs/2211.09066
               | 
               | also it knows when to use a calculator if it has access
               | to one so it's not a big deal
        
           | 908087 wrote:
           | [dead]
        
           | Footkerchief wrote:
           | Military command and control is already performed via input
           | and output of token streams.
        
         | fossuser wrote:
         | The recent paper about using gpt-4 to give more insight into
         | its actual internals was interesting, but yeah the risks seem
         | really high at the moment that we'd accidentally develop
         | unaligned AGI before figuring out alignment.
         | 
         | Out of the options to reduce that risk I think it would really
         | take something like this, which also seems extremely unlikely
         | to actually happen given the coordination problem:
         | https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-no...
         | 
         | You talk about aligned agents - but there aren't any today and
         | we don't know how to make them. It wouldn't be aligned agents
         | vs. unaligned, it's only unaligned.
         | 
         | I don't think spreading out the tech reduces the risk.
         | Spreading out nuclear weapons doesn't reduce the risk (and with
         | nukes at least it's a lot easier to control the fissionable
         | materials). Even with nukes you can still create them and
         | decide not to use them, not so true with superintelligent AGI.
         | 
         | If anyone could have made nukes from their computer humanity
         | may not have made it.
         | 
         | I'm glad OpenAI understands the severity of the problem though
         | and is at least trying to solve it in time.
        
           | lukeschlather wrote:
           | Unaligned doesn't really seem like it should be a threat. If
           | it's unaligned it can't work toward any goal. The danger is
           | that it aligns with some anti-goal. If you've got a bunch of
           | agents all working unaligned, they will work at cross-
           | purposes and won't be able to out-think us.
        
             | jdasdf wrote:
             | This is a misunderstanding of what AI alignment problems
             | are all about.
             | 
             | Alignment != capability
             | 
             | Think a paperclip maximizing robot that in its process of
             | creating paperclips kills everyone on earth to turn them
             | into paperclips.
        
             | ALittleLight wrote:
             | Alignment is about agreement with human preferences and
             | desires, not internal consistency. e.g. An AI that wanted
             | to exterminate humanity could work towards that goal, but
             | it would be unaligned (unaligned with humanity). Alignment
             | is basically making sure humanity is fine with what the AI
             | does.
        
           | DennisP wrote:
           | What is the "recent paper about using gpt-4 to give more
           | insight into its actual internals?"
        
             | og_kalu wrote:
             | https://openai.com/research/language-models-can-explain-
             | neur...
        
         | arisAlexis wrote:
         | "And it doesn't need to be alive or anything to be dangerous"
         | 
         | Why are tech people stuck in the now and not future looking?
        
           | ilaksh wrote:
           | I just think it's much easier to convince people that
           | existing types of AIs will get somewhat smarter and
           | significantly faster. And that's dangerous enough.
           | 
           | My own belief is that regardless of what we do in terms of
           | the most immediate dangers, within one or two centuries
           | (maximum) we will enter the posthuman era where digital
           | intelligent life has taken control of the planet. I don't
           | mean "posthuman" as in all of the humans have been killed
           | (necessarily), just that what humans 1.0 do won't be very
           | important or interesting relative to what the
           | superintelligent AIs are doing.
           | 
           | I don't think there is anything that prevents people from
           | giving AI all of the characteristics of animals (such as
           | humans). I think it's foolish, but researchers seem
           | determined to do it.
           | 
           | But this is fairly speculative and much harder to convince
           | people of.
        
             | c_crank wrote:
             | If the value of superintelligence is to lead to an Age of
             | Em scenario where AIs (or Ems) do most of the intellectual
             | labor, the reality is still that they would be doing this
             | labor in service of humans. I could see a scenario where it
             | is done in service of the AIs instead, but it would look
             | nothing like the existential risk stuff bandied about by
             | these weenies.
        
         | sagebird wrote:
         | It's not money where their mouth is...
         | 
         | It's paying a cost of doing business. The minimum theater
         | required to minimize expected regulatory cost.
         | 
         | They want to own the saftey issue so they can risk your life
         | for their profit.
        
         | c_crank wrote:
         | Control of military and industrial assets won't be handed willy
         | nilly to AIs, given the threat of legal liability for any
         | mistakes the AI could make. Their tendency of making things up
         | is well known by now.
        
           | mhb wrote:
           | You've been spewing out nonsense at an impressive pace. Stop
           | digging. Read more, write less.
        
         | [deleted]
        
       | mistermann wrote:
       | The turtle stacking arms race has begun. Or, is transitioning
       | into the AI enhanced era.
        
       | Imnimo wrote:
       | Big surprise that OpenAI's solution here is to train a more
       | powerful language model and ask it how to do alignment.
        
       | thanatropism wrote:
       | I don't understand how people can still pretend to ignore this:
       | https://plato.stanford.edu/entries/arrows-theorem/
       | 
       | There's also a whole map-territory problem where we're still
       | pretending the distinction hasn't collapsed, Baudrillard-style.
       | As if we weren't all obsessed with "prompt engineering" (whereby
       | the machine trains _us_ ).
        
         | DennisP wrote:
         | Could you explain how you think Arrow's theorem applies?
        
       | JimtheCoder wrote:
       | Let me know when SuperSuperAlignment is announced. Until then,
       | this is just marketing bs...
        
       | bitshiftfaced wrote:
       | My main take away from this is that OpenAI is anticipating super
       | intelligence within the span of ten years. It's one thing to talk
       | about it in a theoretical sense or read about it in science
       | fiction. But to read about it in the real world coming from an
       | operating company in earnest, it feels surreal.
        
         | JimtheCoder wrote:
         | You've never heard a business make a BS claim?
        
           | crop_rotation wrote:
           | Businesses do that all the time, and yet sometimes real
           | products come out of what sounds like hype marketing BS. In
           | this case there might be a non trivial chance of it being
           | true.
        
       | saintradon wrote:
       | Even if this is marketing fluff it's still a remarkable
       | statement, especially considering what they've already done with
       | AI as of now.
        
       ___________________________________________________________________
       (page generated 2023-07-05 23:01 UTC)