[HN Gopher] The changing goalposts of AGI and timelines
___________________________________________________________________
The changing goalposts of AGI and timelines
Author : skandium
Score : 315 points
Date : 2026-03-08 17:16 UTC (5 hours ago)
(HTM) web link (mlumiste.com)
(TXT) w3m dump (mlumiste.com)
| bilekas wrote:
| Hah can you imagine a world where OpenAi says to all the people
| who have dumped billions in : "well we lost guys, sorry about
| that, were just gonna help Google now".
|
| I'll eat my hat after I sell you a bridge.
| skerit wrote:
| Gemini might be great at benchmarks, it is terrible at actual
| agentic coding. So Anthropic seems like a more logical choice.
| bilekas wrote:
| The particulars don't matter. OpenAI will never do this.
| kirubakaran wrote:
| """ I resigned from OpenAI. I care deeply about the Robotics team
| and the work we built together. This wasn't an easy call. AI has
| an important role in national security. But surveillance of
| Americans without judicial oversight and lethal autonomy without
| human authorization are lines that deserved more deliberation
| than they got. This was about principle, not people. I have deep
| respect for Sam and the team, and I'm proud of what we built
| together. """
|
| - Caitlin Kalinowski, previously head of robotics at OpenAI
|
| https://www.linkedin.com/posts/ckalinowski_i-resigned-from-o...
| sheepscreek wrote:
| I 100% understand and agree the AI community argument around
| lethal autonomy.
|
| But I am trying to understand this from the perspective of
| defence & govt. Why is it so business as usual for them? Do
| they consider this at par with missiles with infra-red/heat
| sensors for tracking/locking? Where does the definition of
| lethal autonomy begin and end?
|
| Just putting this out there as a point to ponder on. By itself,
| this may rightly be too broad and should be debated.
| bigyabai wrote:
| For surveillance at least, multimodal AI is old hat: https://
| en.wikipedia.org/wiki/Sentient_(intelligence_analysi...
|
| If you're one of the contractors working in NRO or aware of
| Sentient, OpenAI and Anthropic probably _do_ look like supply
| chain risks. They want to subsume the work you 're already
| doing with more extreme limitations (ones that might already
| be violated). So now you're pitching backup service
| providers, analyzing the cost of on-prem, and pricing out
| your own model training; it would be _really_ convenient if
| OpenAI just agreed to terms. As a contractor, you can make
| them an offer so good that it would be career suicide to
| refuse it.
|
| Autonomous weapons are a horse of a different color, but it's
| safe to assume the same discussions are happening inside
| Anduril et. al.
| lich_king wrote:
| My most straightforward read is that the military simply
| doesn't want their contractors to have a say in the war
| doctrine. Raytheon doesn't get to say "you can only bomb the
| countries we like, and no hitting hospitals or schools". It
| doesn't necessarily mean the Pentagon wants to bomb
| hospitals, but they also don't want to lose autonomy.
|
| A less charitable interpretation is that the current doctrine
| is "China / Russia will build autonomous killbots, so we
| can't allow a killbot gap".
|
| I'm frankly less concerned about "proper" military uses than
| I am about the tech bleeding into the sphere of domestic law
| enforcement, as it inevitably will.
| remarkEon wrote:
| >A less charitable interpretation is that the current
| doctrine is "China / Russia will build autonomous killbots,
| so we can't allow a killbot gap".
|
| What's the reason this is less charitable, exactly? Do we
| think this isn't true, or that we think it's immoral to
| build the Terminator even if China/Russia already have
| them?
| lich_king wrote:
| I don't know what you're trying to argue about here. I
| meant "charitable" as in "not necessarily implying the
| thing critics worry about". The less charitable
| interpretation is that the implied thing is true but is
| seen as a necessity.
|
| We'll leave the morality of war for another time.
| BoxFour wrote:
| I don't think it's about lethal autonomy specifically as much
| as it's just about government autonomy period. They don't
| think private companies should have any veto power over how
| the government uses some technology they're provided.
|
| On its face that's not a crazy stance: Governments are meant
| to represent the public, while private companies obviously
| aren't. I think it's somewhat understandable why the
| government might reject that kind of "we know better than
| you" type of clause.
|
| Of course, the reaction is wildly out of proportion. A normal
| response would just be to stop doing business with the
| company and move on. Labeling them a supply chain risk is an
| extreme response.
| remarkEon wrote:
| Agree, and I think the labeling of them (Anthropic) a
| supply chain risk was handled poorly and will likely be
| reverted over time. That being said, I would be nervous if
| I was in the Pentagon and depended on Anthropic tooling for
| something, even if that something was unrelated to kinetic
| operations. How do they audit that Anthropic can't alter
| model outputs for contexts they (the ethics board or
| whatever it's called, can't remember) don't like? If you
| sell a weapon to the department that is in charge of
| killing people and breaking things, you don't get a say in
| who gets killed or how. It's never worked like that.
|
| Maybe the argument is that they should, but I don't agree
| with that. If Anthropic or any of these other vendors have
| reservations about the logical conclusion of how these
| tools will be/are used then they should not sell to the
| government. Simple as. However ... if the claims Anthropic
| et al make about how these systems will develop and the
| capabilities they will have are at all true, then _the
| government will come knocking anyway_.
| BoxFour wrote:
| > the government will come knocking anyway.
|
| Dario has even said something along these lines at one
| point: As the technology matures, it's very possible the
| government either nationalizes or semi-nationalizes
| companies like Anthropic.
|
| That doesn't seem out of the realm of possibility if they
| can't land on a relationship similar to existing defense
| contractors like Raytheon, where these kinds of
| discussions obviously don't seem to happen.
| nradov wrote:
| If the government wants a frontier LLM for military
| purposes then they can just put out a tender. Defense
| contractors like Anduril will bid on it. The end product
| might be slightly worse than what Anthropic sells but, as
| my dad used to say, "close enough for government work".
| BoxFour wrote:
| They don't even need to do that. Elon is almost certainly
| pushing Grok as hard as possible right now to them, and
| it's not like this administration is especially concerned
| with running a fair procurement process.
|
| So it's probably some mix of two things:
|
| 1) A punitive "bend the knee us or we'll destroy you,"
| which fits their track record.
|
| 2) Skepticism that Grok is actually as strong as the
| benchmarks suggest, which is also a pretty reasonable
| possibility.
| spacemanspiff01 wrote:
| > How do they audit that Anthropic can't alter model
| outputs for contexts they (the ethics board or whatever
| it's called, can't remember) don't like?
|
| I was thinking that Anthropic would just be providing the
| models/setup support to run their models in aws gov
| cloud. They do not have any real insight into what is
| being asked. Maybe a few engineers have the specific
| clearances to access and debug the running systems, but
| that would one or two people who are embedded to debug
| inference issues - not something that would be analyzed
| by others in the company.
|
| The whole 'do not use our models for mass surveillance'
| is at the end of the day an honor system. Companies have
| no real way of enforcing that clause, or determining that
| it has been violated. That being said, at least
| historically, one has been able to trust the government
| to abide by commercial agreements. The people who work in
| cleared positions are generally selected for honesty, and
| ability, willingness to follow rules.
| btown wrote:
| A counter-argument here: if a private company _knows_
| that its technology may be used for human-not-in-loop
| targeting /surveillance, and _knows_ that its technology
| is not yet ready to fulfill that use case without
| meaningful unintended casualties... does that company
| have an ethical obligation to contractually delineate its
| inability to offer that service?
|
| In a version of a trolley problem where you're on a track
| that will kill innocent people, and you have the
| opportunity to set up a contract that effectively moves a
| switch to a track without anyone on it, is it not
| imperative to flip that switch?
|
| (One might argue that increased reaction times might save
| service members' lives - but the whole point is that if
| the autonomous targeting is incorrect, it may just as
| well lead to increased violence and service member
| casualties in the aggregate.)
|
| And we're not talking about the ethics board manipulating
| individual token outputs subtly, which would indeed be a
| supply chain risk - we're talking about a contractual
| relationship in which, if a supplier detects use outside
| of the scope of an agreed contract, it has the
| contractual right to not provide the service for that
| novel use, while maintaining support for prior use cases.
|
| The fact that the government would use the threat of
| supply chain risk to enforce a better contract is
| unprecedented, and it deteriorates the government's
| standing as a reliable counterparty in general.
| mlinhares wrote:
| It's not obvious that the government should have to power
| to overwrite this, the US constitution was written as a
| collection of negative rights exactly to rein in government
| dictatorial impulses.
|
| And now that we see the government blatantly disrespecting
| the constitution and the rule of law the civil community
| must react.
| BoxFour wrote:
| > It's not obvious that the government should have to
| power to overwrite this
|
| The government shouldn't be able to set the terms of its
| contracts with private companies and walk away if those
| terms aren't acceptable? That seems like a stretch.
|
| The constitution is a wildly different premise from
| government contracting with private companies.
| mlinhares wrote:
| There was no contract, the government wanted to have a
| contract where they'd be able to use the tool to violate
| privacy rights of its citizens and issue kill orders
| without a human present and the company said no.
|
| The government shouldn't be able to coerce a business to
| do whatever it wants.
| dmschulman wrote:
| Additionally that kind of public trust only works if you
| have a government operating under the constraints of a
| legal framework, and to a lesser extent, an ethical
| framework. When a government serves the whims of an
| individual and instead of the function of their office,
| shirking agreed upon laws, etc, then you no longer have a
| government serving the people.
| BoxFour wrote:
| Sure, that's why I said "on its face." This
| administration is obviously very different than most.
|
| I don't think Anthropic is wrong to include that clause
| with this particular administration, and I doubt the
| administration is internally framing the issue the way I
| did rather than defaulting to simple authoritarian
| instincts.
|
| But a more reasonable administration could raise the same
| concern, and I think I would agree with them.
| whatshisface wrote:
| I don't think it's reasonable to take something the
| government is supposed to be protecting (right to
| contract) and turn them into its biggest threat. That's
| not security, it's letting the night guard raid the
| museum.
| BoxFour wrote:
| Sure, I said as such:
|
| > Of course, the reaction is wildly out of proportion. A
| normal response would just be to stop doing business with
| the company and move on. Labeling them a supply chain
| risk is an extreme response.
| crote wrote:
| > They don't think private companies should have any veto
| power over how the government uses some technology they're
| provided.
|
| On the other hand, why should the government have infinite
| power to override how a business operates? If you're not
| able to refuse to sell to the government, isn't that
| basically forced speech and/or forced labor?
| nradov wrote:
| If you want to engage on this topic then you should start
| by reading up on the Defense Production Act of 1950.
|
| https://www.congress.gov/crs-product/R43767
| marcosdumay wrote:
| > But I am trying to understand this from the perspective of
| defence & govt.
|
| Hum...
|
| The one thing domestic surveillance enables is defining
| targets inside the country, and the one thing lethal autonomy
| enables is executing targets that a soldier would refuse to.
|
| Those things don't have other uses.
| nradov wrote:
| The military has deployed lethal autonomous weapons since at
| least 1979. LLMs might be useful for certain missions but
| from a military perspective they're nothing fundamentally
| new.
|
| https://www.vp4association.com/aircraft-
| information-2/32-2/m...
| hirvi74 wrote:
| > _This was about principle, not people._
|
| Why do I not believe this at all? Were things truly sunshine
| and roses at OpenAI up until this Pentagon debacle? Perhaps I
| am mistaken, but it seemed like the writing was on the wall
| years ago.
|
| > _I have deep respect for Sam and the team_
|
| I have even more questions now.
| lucianbr wrote:
| How can you respect someone who betrays a principle you care
| about for money?
|
| Not to mention that the principles are not being betrayed now
| for the first time.
| hirvi74 wrote:
| > _respect someone who betrays a principle you care about
| for money?_
|
| That is one of my many questions too. I am not certain I
| believe her either. People predicted AI would be used in
| such nefarious manners way before AI even existed.
|
| Something about the whole resignation and immediate social
| media post seems more like an attention grab than anything
| else to me. Whatever her prerogative is, I still believe
| she is still partially culpable for anything that becomes
| of this technology -- good or bad.
| Lerc wrote:
| You could believe it is not about money.
|
| Most importantly, this seems to rest more on if you believe
| the principle was being followed or not.
|
| It is possible to believe one thing, have another person
| believe another thing, and respect that their decision is
| sincerely held but subject to a different perspective, as
| is our own beliefs.
|
| You can stand up for what you believe and still
| respectfully disagree with someone with a different stance.
|
| The problem is when you decide that reality always conforms
| to you opinion. If you assume the other person is aware of
| that reality and decides differently than it becomes a
| betrayal of principles. Assuming to know the internal state
| of another's mind to declare that it is for money becomes
| disrespectfully presumptuous.
|
| Your problem is not in understanding how X can occur if Y.
| It is assuming that everyone agrees with you on Y.
|
| You might be right about Y, you might be wrong. Even if you
| are right, it is still possible that Y is a belief a
| rational person can hold if their perspective has been
| different,
| daheza wrote:
| Sounds like a statement to ensure they aren't blacklisted or
| seen as anti executive.
| hirvi74 wrote:
| > _they aren't blacklisted or seen as anti executive._
|
| Which further solidifies my belief that this person is
| being disingenuous.
| ronnier wrote:
| Do Chinese do this in China? Walk away from companies that will
| be used for war? I doesn't seem to be prevalent and instead
| they try to take every advantage they can to push their
| country, China, to become the most dominate in the world. They
| must be elated to watch the world's premier tech companies
| protest the American government and refusal to work with them.
| If I wanted China to be weaker I'd hope that Chinese companies
| protested and refused to work with the Chinese government.
| tkz1312 wrote:
| You have used chatgpt presumably. Based on your interactions
| with it, do you seriously think it should be allowed to shoot
| a gun without any human oversight?
| ronnier wrote:
| That simplistic question is not how things will work. I
| guess we'll just get shot by Chinese AI, they will not
| stop.
| esafak wrote:
| You'd rather get shot by domestic bots first?
| ronnier wrote:
| We have nukes, missiles, bombs, all capable of mass
| widespread death. Should we give those up too and just
| let adversaries be the only ones in possession of these
| types of weapons?
| esafak wrote:
| Autonomous robots are one of the adversaries. They're
| their own side.
| user3939382 wrote:
| Yes we should dispense with ethics so we can win at all
| costs. Like your point isn't invalid but what's the point of
| restating something akin to the trolley problem but this
| time, as if the answer is obvious.
| qwerpy wrote:
| We can debate philosophy while our adversaries use any
| means at their disposal. Or we can invest in different
| ideas, see what works, and choose the best option.
| _DeadFred_ wrote:
| What are we if we throw away the Constitution and allow
| the Government to punish people/companies that exercise
| their rights?
|
| China's constitution includes freedom of speech and
| elections.
|
| Funny thing when you put rights on hold today for
| 'reasons' they tend to just go away. Look at the US today
| versus pre 9/11. It's a completely different country with
| completely different attitudes about freedom and privacy
| and government over reach and power.
| fancy_pantser wrote:
| It's explicitly illegal in China.
|
| A 2017 national intelligence law compels Chinese companies
| and individuals to cooperate with state intelligence when
| asked and without and public notice.
|
| China has no equivalent of the whistleblower protection that
| enables resignations with public letters explaining why,
| protests, open letters with many signatures, etc. Whenever
| you see "Chinese whistleblower" in the news, you're looking
| at someone who quietly fled the country first and then blew
| the whistle. Example:
| https://www.cnn.com/2026/02/27/us/china-nyc-whistleblower-
| uf...
| crote wrote:
| Isn't that basically the same as a National Security Letter
| and its attached gag order in the USA?
| fancy_pantser wrote:
| It's along the same lines, but an NSL can be challenged
| in court (the FISC is a secret and lopsided court, alas).
| Companies like Apple and Google have fought specific
| orders publicly (and possibly some secretly), and some
| have won.
|
| NSLs are also narrow in scope: they compel data
| disclosure, not active technical assistance in building
| surveillance systems like the Chinese law.
|
| The Chinese laws can compel any citizen anywhere in the
| world to perform work on supporting state military and
| intelligence capabilities with no recourse. There have
| been no cases of companies or individuals fighting those
| orders.
| nradov wrote:
| Not at all. If you're an employee at a company that
| receives a National Security Letter then you can just
| quit if you want to. Unlike in China, the US government
| can't force you to keep working there to suit their
| purposes.
| pfortuny wrote:
| One of the things about slave coups in ancient times was that
| they really believed there are things more important than
| life.
| yorwba wrote:
| Yes, of course there are people in China who, when their job
| puts them in conflict with their ethics, will decide to do
| something ethical. I can't think of any war-related examples,
| since it's been a while since China was involved in any big
| wars, but I like the story of Liu Lipeng, who used to work as
| an internet censor:
| https://madeinchinajournal.com/2025/04/03/me-and-my-censor/
| rishabhaiover wrote:
| > The impotence of naive idealism in the face of economic
| incentives
|
| A great point. I saw blinding idealism during the early days of
| GPT era.
| coliveira wrote:
| This was never idealism, it was much more about gaslighting.
| Billionaire investors have a playbook where they say something
| to gaslight people into agreeing with them, and then go to the
| opposite direction. It is already a pattern. For example, most
| companies would swear they were decided on cutting carbon
| emissions, just to forget everything when it comes to build
| data centers. They say some technology is just for the growth
| of mankind when in fact they're looking for monopoly and
| destroying jobs. They say they're donating money for "charity"
| when they're in fact investing in technologies that they have
| vested financial interest. They say we need to contribute to a
| not-for-profit institution which will later be used to create
| another monopoly. It is so predictable that I wonder why anyone
| can still be fooled by this.
| CamperBob2 wrote:
| _For example, most companies would swear they were decided on
| cutting carbon emissions, just to forget everything when it
| comes to build data centers._
|
| This doesn't seem contradictory if you consider that success
| at AGI will solve the problem of carbon emissions, one way or
| another. If one data center ultimately replaces a whole
| medium-sized city of commuters...
| bluefirebrand wrote:
| > If one data center ultimately replaces a whole medium-
| sized city of commuters...
|
| Then we find out how long it takes for a medium sized city
| of commuters to start killing each other, elites and
| burning down data centers. Once they're hungry enough it'll
| happen for sure
| Ekaros wrote:
| Data centres should have plenty of good loot. Raw
| materials like copper or at least aluminium. Maybe even
| steel, but value proposition there is less likely. I
| suppose someone will be interested in example fuel too if
| there is fuel based backup generation.
| coliveira wrote:
| Thinking about that, a world dominated by data centers
| will be relatively easy do disrupt, someone just needs to
| destroy a dozen or so datacenters to bring everything
| down.
| irishcoffee wrote:
| A given country is about 6 missed meals away from
| complete anarchy.
| croes wrote:
| > Therefore, if a value-aligned, safety-conscious project comes
| close to building AGI before we do
|
| > It can be debated whether arena.ai is a suitable metric for
| AGI, a strong case can probably be made for why it's not.
| However, that's irrelevant, as the spirit of the self-sacrifice
| clause is to avoid an arms race, and we are clearly in one.
|
| No, the spirit is clearly meant for near AGI and we aren't near
| AGI
| bigyabai wrote:
| AFAIK we're still working on a unified definition and testing
| theory for whatever "AGI" is.
| MichaelDickens wrote:
| Altman has personally claimed that we are close to AGI.
| Therefore, according to him, OpenAI should invoke the self-
| sacrifice clause.
| croes wrote:
| Of course he claims that, he seeks money from investors but
| the charter is likely be written by people who took it
| seriously
| rvz wrote:
| The "I" in AGI stands for IPO.
|
| The "S" stands for Safety.
| choult wrote:
| The writing was on the wall as soon as it went all-in on
| commercializing the tech.
|
| This will never happen, LLMs are already being used very
| unsafely, and if this HN headline stays where it is OpenAI will
| quietly remove their charter from their website.
| diabllicseagull wrote:
| > The impotence of naive idealism in the face of economic
| incentives.
|
| I don't think it was so much the naivety of idealism, but more an
| adoption of idealism and related language to help market what was
| actually being built: a profit-first organization that's taking
| its true form little by little.
| samrus wrote:
| I genuinely dont think it was profit first thenm i think it was
| all in good faith. Altman definitely got greedy and stabbed all
| that in the back
| runarberg wrote:
| This is an extraordinary claim and it requires extraordinary
| evidence. There is nothing in Altman's behavior current or
| past to suggest this was anything other then a money making
| grift. The easiest explanation for his betrayal is that he
| was simply lying.
| conradkay wrote:
| I'm not fan of Altman but the financial angle doesn't make
| much sense when he doesn't have equity in OpenAI
|
| There's some indirect exposure and potential of being
| granted significant equity, but his actions don't read as
| being for his own wealth
|
| As an example, he walks away with nothing in the very
| plausible timeline where he was fired but not then
| reinstated
| dataflow wrote:
| It's clever and funny, but nobody is legitimately near AGI, and
| their own AML Corp link proves Altman believes as much:
|
| > Achieving AGI, he conceded, will require "a lot of medium-sized
| breakthroughs. I don't think we need a big one."
|
| > At the Snowflake Summit in June 2025, Altman predicted that
| 2026 would mark a breakthrough when AI systems begin generating
| "novel insights" rather than simply recombining existing
| information. This represents a threshold he considers critical on
| the path to AGI.
|
| Though I'm sure they'll try to change the charter before we get
| to that point, but yeah.
| labrador wrote:
| The way Sam Altman bungled the Pentagon deal by swooping in a few
| hours after Anthropic was fired should be grounds for OpenAI
| finding another CEO.
| mirsadm wrote:
| Why? Give it a couple weeks and everybody will forget about
| this. They'll be earning more money than previously. Job well
| done.
| dgroshev wrote:
| Just like everyone forgot about this
| https://www.wired.com/story/openai-staff-walk-protest-sam-
| al...
| wongarsu wrote:
| Employees are the ones with the real power to make this
| hurt. The customers switching over are easily offset by the
| DoD contract. But losing talent over this, and having a
| harder time to attract future talent? That could hurt them
|
| Sam probably expects to solve this by just offering more
| money. It worked in the past
| integralid wrote:
| Yeah, they will have to raise salary by 10% to attract
| people. This will no doubt hurt their bottom line. Poor
| starving SV text workers will have no choice but to
| accept working for them, lest they starve.
|
| Maybe my sarcasm is not justified, but I don't think most
| people care that that work for a company that does
| unethical things. In fact I think all large companies are
| more or less immoral (or rather amoral) - that's just how
| the system is built.
| trollbridge wrote:
| Given the mass layoffs happening, I don't think hiring
| talent is as hard as it's made out to be.
| PunchyHamster wrote:
| Traitors to humanity
| bigyabai wrote:
| What, you think the DoD can only designate one supply chain
| risk at a time?
| micromacrofoot wrote:
| it started before that, the openai president donated 20mil to
| trump the month prior... ellision and kushners also pretty
| heavily involved with openai and altman is tight with peter
| thiel
|
| the whole public debacle was planned, the tos isn't stopping
| the pentagon from doing anything (as we seen with openai now)
| fiatpandas wrote:
| Nytimes reported that DoW prepped that deal in the background
| in parallel, as backup for issues with Anthopic. But I agree,
| optics are terrible.
| swingboy wrote:
| Purely anecdotal, but GPT 5.4 has been better than Opus 4.6 this
| past week or so since it came out. It's interesting to see it
| rank fairly low on that table. Opus "talks" better and produces
| nicer output (or, it renders better Markdown in OpenCode) than
| 5.4.
| sigmoid10 wrote:
| Chatbot Arena is notoriously unreliable for several reasons.
| First it's (at least in theory) based on normal human feedback.
| Given by normal people's current voting trends, they clearly
| are not very good at identifying experts or at least remotely
| correct statements. Second, the leaderboards are gamed hard by
| the big companies. Even ARC AGI entered the actively gamed
| stage by now. Sure the current gen models are certainly better
| than the last and if two are vastly different in leaderboards
| there may be something fundamental to it, but there is hardly
| any reason to use these kinds of comparison tables for anything
| useful among the latest models.
| EugeneOZ wrote:
| Not in my experience. Quoting my tweet:
|
| Gave the same prompt to GPT 5.4 (high) and Opus 4.6 (high).
|
| GPT 5.4 implemented the feature, refactored the code (was not
| asked to), removed comments that were not added in that
| session, made the code less readable, and introduced a bug.
| "Undo All".
|
| Opus 4.6 correctly recognized that the feature is already
| implemented in the current code (yeah, lol) and proposed
| implementing tests and updating the docs.
|
| Opus 4.6 is still the best coding agent.
|
| So yeah, GPT 5.4 (high) didn't even check if the feature was
| already implemented.
|
| Tried other tasks, tried "medium" reasoning - disappointment.
| hirvi74 wrote:
| I make ChatGPT and Claude code review each other's outputs.
| ChatGPT thinks its solutions are better than what Claude
| produces. What was more surprising to me is that Claude, more
| often than not, prefers ChatGPT's responses too.
|
| I am to sure one can really extrapolate much out of that, but
| I do find it interesting nonetheless.
|
| I think language is also an important factor. I have a hard
| time deciding which of the two LLMs is worse at Swift, for
| example. They both seem equally great and awful in different
| ways.
| frde wrote:
| Is this sample size of one task, or a consistent finding
| across many tasks?
| throwaw12 wrote:
| OpenAI:
|
| - we are building Open AI - only if you have more than $10B net
| worth
|
| - we are against using AI for military purposes - except when
| that case is allowed by government
|
| - we are on a mission to help humanity - again, we define
| humanity as set of people with more than $10B net worth
|
| - surrender? - sure, sure, we will, only to people with more than
| $10B net worth, they can do whatever they want to our models, we
| will surrender to them
| dkwmdkfkdk wrote:
| Anyone can make these silly negations. As if any other big corp
| is different. And as if your own imaginary big corp (if you
| took the time and effort to try to achieve, that is) would have
| behaved differently given enough pressure from investors and
| share holders.
|
| This is all just very naive
| tombert wrote:
| Sorry, no, you shouldn't just handwave away bad behavior just
| because it's common. That's ridiculous.
|
| If big corporations do things that are unethical, they should
| be called out, even if they're common. Saying "well
| everyone's doing it", isn't a good excuse to do things that
| are unethical.
|
| It's not "naive" to point out the lies that OpenAI told to
| get to the point that they are now. They were claiming to be
| a non-profit for awhile, they grew in popularity based in
| part on that early good-will, and now they are a for-profit
| company looking to IPO into one of the most valuable
| corporations on the planet. That's a weird thing. That's a
| thing that seems to be kind of antithetical to their initial
| purpose. People should point that out.
| PunchyHamster wrote:
| > Anyone can make these silly negations. As if any other big
| corp is different.
|
| The point you're desperately trying to miss is that most
| other companies don't put up those moral claims in the first
| place
| reppap wrote:
| Sam Altman is just another Elon Musk. Saying whatever he thinks
| sounds good in the moment.
| kakacik wrote:
| Its much, much broader. Almost all of them up there are
| hardcore sociopaths. When they look at you, they see a tool
| and source of revenue, not a human being. To be used, and
| then thrown away. Musk, Bezos, Gates, Schmidt, Altman, Thiel
| and so on and on. Most career politicians are similar.
|
| Is this really the garbage that should lead humanity to our
| future? Because inevitably that will be a dark future for
| 99.99% of the humans. And no, _you_ and _you_ won 't be part
| of that 0.01% or whatever tiny number of elite think they are
| better than rest of us.
| wrsh07 wrote:
| You've attached to the $10b number - is that illustrative or is
| there something specific you're referring to?
| bluegatty wrote:
| AI will be used wherever computers, silicon, RAM, software, GPUs
| and robots are today.
|
| And that's it.
|
| Everything beyond that is nuance.
|
| Nuance matters, but it's not the real story, it's the side show.
| p-o wrote:
| I think the brunt of the disruption regarding AI is already
| behind us for LLMs at least. It's possible we'll see improvements
| over the following months/years, but government will inevitably
| start to catchup to the level of disinformation and confusion
| that AI has brought to this world.
|
| Laws & regulations that needs to be created to reign in AI will
| undoubtedly increase the opportunity cost of training LLMs.
|
| For some, it might be similar to the early 2000s, but I think
| it's just a healthy rebalance of what AI is, and how the society
| needs to implement this new, hardly controllable, paradigm. With
| this perspective, OpenAI has a lot to lose as it hasn't been able
| to create a moat for itself compared to, let's say, Anthropic.
| eloisant wrote:
| I think that even if the models were to plateau today, there
| are still a lot of room for improvement in all the tooling
| around them, people finding ideas of applications, and users
| getting used of them. So we're not done with the disruption.
|
| Some of the apps made possible by smartphones only appeared a
| decade after they were made technically possible. A lot of the
| new use cases made possible by the Internet and broadband
| connections only became widely used because of Covid.
|
| I was already using Skype 20 years ago to make video calls, but
| I've only seen PTA meetings over Zoom since Covid.
| p-o wrote:
| Yes, I think you're right that it is not the end of the road
| for LLM and the application of LLM might be adopted over time
| across a variety of industries.
|
| I guess what I failed to convey in my original comment was
| that, like the Internet 20 years ago, the current advancement
| made by AI might _stall_ at a foundational level, while the
| landscape evolves.
|
| Essentially, I believe what you're saying is really close in
| spirit to what I'm saying.
| dmix wrote:
| This is taking Sam Altmans PR statements as proof of AGI?
|
| Even the quote they used questions the premise of the article
|
| > "We basically have built AGI" (later: "a spiritual statement,
| not a literal one")
| wongarsu wrote:
| AGI is so nebulous we will never be able to tell if we hit it.
| We have hit human-level abilities in some narrow tasks, and are
| still leagues away in others. And humans have so vastly
| different skill-levels that we can't even agree what human-
| level really means. As bad as the economic definition of AGI in
| OpenAI's Microsoft deal is, at least it's measurable.
|
| Imho that's a big part of why people are shifting to ASI. Not
| because we reached AGI, but because 'we reached ASI' is a well-
| defined verifiable statement, where 'we reached AGI' just isn't
| sebastiennight wrote:
| > 'we reached ASI' is a well-defined verifiable statement,
| where 'we reached AGI' just isn't
|
| So... we can't tell when the rocket has left Earth
| atmosphere, but we can tell when the rocket has entered
| space?
|
| I'm not getting how "superior in all tasks" is better-defined
| for you than "equal in all tasks".
| wongarsu wrote:
| Because ASI is 'superior to the best human' while AGI is
| 'matches or surpasses humans'. It's hard to agree on what
| 'matches humans' actually means. Would matching humans in
| coding involve the level of the average human (basically no
| coding ability), the level of a junior (the largest group
| of professionals), a senior ('someone competent') or of the
| best human we can find?
|
| Or as the scene from 'I, Robot' goes: Will Smith asks the
| android: 'Can a robot write a symphony? Can a robot turn a
| blank canvas into a masterpiece?' and the android simply
| answers 'Can you?' ASI sidesteps that completely
| hirvi74 wrote:
| > AGI is so nebulous we will never be able to tell if we hit
| it.
|
| I completely agree. We can't even measure each other well,
| let alone machines.
| Jensson wrote:
| It is very easy to tell if we still need humans in the
| loop. We still do so its not AGI.
| wongarsu wrote:
| For certain types of "human in the loop". If it can't
| write working code without a human in the loop then it's
| not AGI. But a human-level coder also has lots of humans
| in the loop: a more senior developer doing code review,
| several layers of management, a product owner that
| interfaces the project with outside reality, sales
| people, etc.
|
| Now I already hear you typing "but those roles should
| also be handles by AI if it's AGI" and I agree that an AI
| that can claim to be AGI should be able to handle those
| roles (as separate agents if necessary). But in a real
| setup it probably won't be the best choice to do those
| roles for cultural and legal reasons. Or it might simply
| not be cost effective. Not to mention that under most
| definitions of AGI there can still be humans more capable
| than the AI, as long as the AI hits the 50th percentile
| mark or something like that. So even if it's an AGI with
| the ability to do these roles we will still have humans
| in the loop for a long long time
| hirvi74 wrote:
| I am unaware of any definition of AGI that states AGI
| cannot have humans in the loop.
| eloisant wrote:
| I feel like AI is actually helping us getting a better
| understanding of what "human intelligence" really is.
|
| I remember when computers became better than humans at chess,
| many people were shocked and saw that as machines becoming
| more intelligent than humans. Because being good at chess
| what considered equivalent to "being smart".
| 0xbadcafebee wrote:
| AGI isn't going to happen within the next 30 years so this is
| moot. The actual researchers have said so many times. It's only
| the business people and laypeople whooping about AGI always being
| imminent.
|
| You cannot get real, actual AGI (the same ability to perform
| tasks as a human) without a continuous cycle of learning and deep
| memory, which LLMs cannot do. The best LLM "memory" is a search
| engine and document summarizer stuffed into a context window
| (which is like having someone take an entire physics course,
| writing down everything they learn on post-it notes, then you ask
| a _different person_ a physics question, and that different
| person has to skim all the post-it notes, and then write a new
| post-it note to answer you). To learn it would need RL (which
| requires specific novel inputs) and retraining (so that it can
| retain and compute answers with the learned input). This would
| all take too much time and careful input /engineering along with
| novel techniques. So AGI is too expensive, time consuming, and
| difficult for us to achieve without radically different designs
| and a whole lot more effort.
|
| Not only are LLMs not AGI, they're still not even that great at
| being LLMs. Sure, they can do a lot of cool things, like write
| working code and tests. But tell one "don't delete files in X/",
| and after a while, it _will_ delete all the files in "X/",
| whereas a human would likely remember it's not supposed to delete
| some files, and go check first. It also does fun stuff like
| follow arbitrary instructions from an attacker found in random
| documents, which most humans also wouldn't do. If they had a real
| memory and RL in real-time, they wouldn't have these problems.
| But we're a long way away from that.
|
| LLMs are fine. They aren't AGI.
| nerdsniper wrote:
| > _which is like having someone take an entire physics course,
| writing down everything they learn on post-it notes, then you
| ask a different person a physics question, and that different
| person has to skim all the post-it notes, and then write a new
| post-it note to answer you_
|
| This is the best summary of an LLM that I've ever seen (for
| laypeople to "get it") and is the first that accurately
| describes my experience. I will say, _usually_ the notes passed
| to the second person are very impressive quality for the topic.
| But the "2nd person" still rarely has a deep understanding of
| it.
| trollbridge wrote:
| A layperson analogy I use is that an LLM is like Dora with a
| really high IQ - it effectively needs everything reexplained
| to it, and you can't give it more than a few seconds of
| context before it just forgets.
| slavik81 wrote:
| Do you mean Dory, the fish from Finding Nemo?
| herodoturtle wrote:
| Agree with your overall point. Curious what you're basing the
| "not in the next 30 years" claim on, if you'd care to expand.
| mattlondon wrote:
| I disagree. There is some argument to be had that they're
| already generally intelligent. They're already certainly
| _better_ than me in basically anything I can ask them to do.
|
| So that leads to the question of what qualifies as intelligent?
| And do we need sentience for intelligence? What about self-
| agency/-actuation? Is that _needed_ for "generally
| intelligent"?
|
| I don't know.
|
| But I _feel_ like we 're not there yet, even for non-sentient
| intelligence. I personally think we need an "unlimited" context
| (as good as human memory context windows anyway, which some
| argue we've already surpassed) and genuine self-learning before
| we get close. I don't think we need it to be an infallible
| genius (i.e ASI) to qualify as generally intelligent ... or to
| put it another way "about as smart and reliable as the average
| human adult" which frankly is quite a low bar!
|
| One thing for sure though, I think this will creep up on us and
| one day it will suddenly become apparent that it's already
| there and we just didn't appreciate/notice/comprehend. There
| won't be a big fireworks display the moment it happens, more of
| a creeping realisation I think.
|
| I give it 5 years +/-2.
| kovek wrote:
| Models need pre-training and fine tuning. Humans can do
| online learning.
| mattlondon wrote:
| How much of our brain's innate wiring is "pre-training"?
| We're born pre-wired to breath, to swallow, to blink,
| sleep, cry etc. Fine tuned over many many many epochs and
| baked into our model weights/DNA.
|
| Is a newborn baby without learnt-knowledge not an
| intelligent being to you?
|
| Or is an empty vessel such as a newborn baby intelligent
| merely because it has the _ability_ to learn?
|
| It gets pretty philosophical pretty quick. This is why I
| don't think there'll be a "moment" when AGI happens - there
| is so many ways to interpret what constitutes intelligence.
|
| But yeah I agree that until models can learn in real time
| then I think we're probably not there yet. As I said - 5
| years give or take I reckon.
| irishcoffee wrote:
| > Is a newborn baby without learnt-knowledge not an
| intelligent being to you?
|
| A newborn baby without learnt knowledge is a phenomenal
| comparison to an LLM: reactionary, incapable of
| communicating a novel thought, extremely inconsistent
| reactions to similar stimuli, costs a lot of money, and
| the best part is they almost never turn out to be an
| income-bearing investment.
|
| Despite all this, people vehemently defend their ugly,
| obnoxious, screaming, drooling babies with the fierceness
| of a lion, because they're so blinded by emotion they're
| incapable of logical thought.
|
| I have three kids, I've earned the right to say this.
|
| What a great comparison!
| fwipsy wrote:
| Given how many "fundamental" limitations of AI have been
| resolved within the past few years, I'm skeptical. Even if
| you're right, I am not sure that the limitations you identified
| matter all that much in practice. I think very few human
| engineers are working on problems which are so novel and unique
| that AIs cannot grasp them without additional reinforcement
| learning.
|
| > it will delete all the files in "X/"
|
| How many "I deleted the prod database" stories have you seen?
| Humans do this too.
|
| > follow arbitrary instructions from an attacker found in
| random documents
|
| This is just the AI equivalent of phishing - inability to
| distinguish authorized from unauthorized requests.
|
| Whenever people start criticizing AI, they always seem to
| conveniently leave out all the stupid crap humans do and
| compare AI against an idealized human instead.
| surgical_fire wrote:
| > Given how many "fundamental" limitations of AI have been
| resolved within the past few years
|
| Eh? Which limitations were solved?
| monsieurbanana wrote:
| I can answer that one: none.
|
| The only thing I can think of is massively increased
| context windows (around 4k for gpt3), but a million context
| token with degraded performance when full is not what I'd
| qualify as resolved.
| fwip wrote:
| > How many "I deleted the prod database" stories have you
| seen? Humans do this too.
|
| Humans generally do it on accident. They don't preface it
| with "Let me delete the production database," which LLMs do.
| sulam wrote:
| Sorry, but you're mistaking outputs with process. If you
| actually know what models are doing under the hood to product
| output that (admittedly) looks very convincing, you'll
| quickly realize that they are simply exceptionally good at
| statistically predicting the next token in a stream of
| tokens. The reason you are having to become an expert at
| context engineering, and the reason the labs still hire
| engineers, is because turning next token prediction into
| something that can simulate general intelligence isn't easy.
|
| The boundaries of these systems is very easy to find, though.
| Try to play any kind of game with them that isn't a
| prediction game, or perhaps even some that are (try to play
| chess with an LLM, it's amusing).
| MadxX79 wrote:
| I enjoyed playing mastermind with LLMs where they pick the
| code and I have to guess it.
|
| It's not aware that it doesn't know what the code is (it
| isn't in the context because it's supposed to be secret),
| but it just keeps giving clues. Initially it works, because
| most clues are possible in the beginning, but very quickly
| it starts to give inconsistent clues and eventually has to
| give up.
|
| At no point does it "realise" that it doesn't even know
| what the secret code is itself. It makes it very clear that
| the AI isn't playing mastermind with you, it's trying to
| predict what a mastermind player in it's training set would
| say, and that doesn't include "wait a second, I'm an AI, I
| don't know the secret code because I didn't really pick
| one!" so it just merilly goes on predicting tokens, without
| any sort of awareness what it's saying or what it is.
|
| It works if you allow it to output the code so it's in
| context, but probably just because there is enough data in
| the training set to match two 4 letter strings and know how
| many of them matches (there's not that many possibilities).
| Balinares wrote:
| That is actually a genius and beautifully simple way to
| exhibit the difference between thought and the appearance
| of thought.
| MadxX79 wrote:
| It really dispelled the illusion for me, but it's not
| that easy to find those examples, but the combinatorics
| of possible number of guesses is untractable enough that
| it can't learn a good set of clues for all possible
| guesses.
| 10xDev wrote:
| CoT already moved things past the "it is just token
| prediction" phase. We have models that can perform search
| over a very large state space across domains with good
| precision and refine its own search leading to a decent
| level of fluid intelligence, hence why ARC AGI 1/2 is
| essentially solved. We also don't know the exact details of
| what is happening at frontier labs seen as they don't
| publish everything anymore.
| kakacik wrote:
| Well, llms are way more stupid, doing things that even most
| juniors wouldn't do (and then you don't give PROD access to
| new junior hire, do you... most people are super careful with
| llms and simply don't trust them and don't let them anywhere
| near critical infra or data - thats seniority 101).
|
| Which fundamental limitation do you mean? I haven't seen
| anything but slow, iterative improvements. Sure, if feels
| fine, turtle can eventually do 10,000 mile trek but just
| because its moving left and right feet and decreasing the
| distance doesn't mean its getting there anytime soon.
|
| Parent mentioned way harder hurdles than iterative increments
| can tackle, rather radical new... everything.
| mattlondon wrote:
| I don't think humans learn any differently than post it notes
| TBH We call them text books though!
| rishabhaiover wrote:
| I don't agree. The recent emergent behavior displayed by LLMs
| and test-time scaling (10x YoY revenue for Anthropic) is worth
| some hype. Of course, you are correct that most people who
| rally behind AGI do not understand the fundamental limitations
| of next-token prediction.
| ACCount37 wrote:
| The "fundamental limitations" being what exactly?
| stratos123 wrote:
| > AGI isn't going to happen within the next 30 years so this is
| moot. The actual researchers have said so many times. It's only
| the business people and laypeople whooping about AGI always
| being imminent.
|
| The statements of what "actual researchers" are you relying
| upon for your "next 30 years" estimate? How do you reconcile
| them with the sub-10- or even sub-5-years timelines of other AI
| researchers, like Daniel Kokotajlo[1] or Andrej Karpathy[2]?
| For that matter, what about polls of AI researchers, which
| usually obtain a median much shorter than 30 years [3]?
|
| [1] https://x.com/DKokotajlo/status/1991564542103662729
|
| [2] https://x.com/karpathy/status/1980669343479509025
|
| [3] https://80000hours.org/2025/03/when-do-experts-expect-agi-
| to...
| MadxX79 wrote:
| I'm guessing they have a lot of shares in the AI companies
| they work(ed) for, and they would like to pump their value so
| they can buy an even nicer carribean island than they can
| already afford?
| dwohnitmok wrote:
| Kokotajlo gave up all his shares in OpenAI as part of his
| refusal to sign a nondisparagement agreement with OpenAI.
| stratos123 wrote:
| Kokotajlo in particular is notable for being the guy who
| quit OpenAI in 2024 in protest of their policy of requiring
| researchers to abide by a non-disparagement agreement to
| retain their equity. In the end OpenAI caved and changed
| their policy, but if he was lying all along to inflate the
| value of his shares, it would have been quite a 4d chess
| move of him to gamble the shares themselves on doing so.
| MadxX79 wrote:
| Isn't it just that he left way before gpt-5, then? At
| that point a sufficiently naive person could have
| believed that scaling was going to lead to AGI, but that
| sort of optimism died after he was already an outsider.
| Spivak wrote:
| See this is a fun game because when you're fishing for a
| breakthrough you can predict tomorrow or 100 years. Nobody,
| even experts have any idea until it happens and they're
| holding it in their hands. To have any kind of accurate
| prediction you would have to have already observed other
| civilizations discover AGI to say how close the environment
| is to even be capable of making the leap. We could be missing
| something huge, we could need multiple seemingly unrelated
| breakthroughs to get there. We're for sure closer, but we
| could still be miles away, GPTs might even barking up the
| wrong tree.
|
| Why this discussion is already annoying and poised to get so
| much worse is because now hundred billion dollar companies
| have a direct financial incentive to say they did it so I
| expect the definition will get softened to near
| meaninglessness so some marketing department can slap AGI on
| their thing.
| linkregister wrote:
| I think you are overindexing on the integer value given in
| the parent post, rather than seeing the essence that LLMs in
| their current form only excel on tasks they have been
| specifically trained for.
|
| Karpathy himself has publicly stated that AGI itself is only
| possible with a new paradigm (that his group is working
| toward). He claims RHLF and attention models are near the end
| of their logarithmic curve. The concept of the "self-training
| AI" is likely impossible without a new kind of model.
|
| We will likely see some classes of human skills completely
| taken over by LLMs this decade: call centers (already capable
| in 2026), SWE (the next couple years). Bear in mind the
| frontier labs have spend many billions on exhaustive training
| on every aspect of these domains. They are focusing training
| on the highest value occupations, but the long tail is huge.
|
| It will be interesting to see if this investment will be
| obviated by a "real AGI" capable of learning without going
| through the capital-intensive training steps of current
| models.
| stratos123 wrote:
| Personally I'm not even sold on the current paradigm being
| too limited to produce AGI - there are still several OOMs
| worth of compute increase available, plus the algorithmic
| improvements have overall been accumulating _faster_ than
| predicted.
|
| But even assuming that a major breakthrough is required, it
| seems ludicrous to me to go from that to a timeline of a
| decade or more. This isn't like fusion power research,
| where you spend 10 years building a new installation only
| to find new problems. Software development is inherently
| faster, and AI research in particular has been moving
| extremely quickly in the past. (GPT-3 is only 6 years old.)
| I don't think a wall in AI progress, if one comes at all,
| will last more than a few years.
| paulryanrogers wrote:
| > there are still several OOMs worth of compute increase
| available
|
| Where and how? Aren't we reaching the physical limits of
| making transistors smaller?
| charcircuit wrote:
| LLMs are AGI because they offer intelligence on any subject.
|
| >the same ability to perform tasks as a human
|
| The first chess AIs lost to chess grandmasters. AI does not
| need to be better than humans to be considered AI.
|
| >without a continuous cycle of learning and deep memory, which
| LLMs cannot do.
|
| But harnesses like Claude Code can with how they can store and
| read files along with building tools to work with them.
|
| >which is like having someone take an entire physics course,
| writing down everything they learn on post-it notes, then you
| ask a different person a physics question, and that different
| person has to skim all the post-it notes, and then write a new
| post-it note to answer you
|
| This don't matter. You could say a chess AI is a bunch of
| different people who work together to explore distant paths of
| the search space. The idea you can split things into steps does
| not disqualify it from being AI.
|
| >But tell one "don't delete files in X/", and after a while, it
| will delete all the files in "X/"
|
| Humans make mistakes and mess up things too. LLMs are better at
| needle in a haystack tests than humans.
|
| >It also does fun stuff like follow arbitrary instructions from
| an attacker
|
| A ton of people get phished or social engineered by attackers.
| This is the number 1 way people get hacked. Do not
| underestimate people's willingness to follow instructions from
| strangers.
| mirekrusin wrote:
| I think you're somehow right and wrong at the same.
|
| All those "it's like ..." are faulty - "post-it notes" are not
| 3k pages of text that can be recalled instantly in one go,
| copied in fraction of a second to branch off, quickly
| rewritten, put into hierarchy describing virtually infinite
| amount of information (outside of 3k pages of text limit),
| generated on the fly in minutes on any topic pulling all
| information available from computer etc.
|
| Poor man's RL on test time context (skills and friends) is
| something that shouldn't be discarded, we're at 1M tokens and
| growing and pogressive disclosure (without anything fancy, just
| bunch of markdowns in directories) means you can already stuff-
| in more information than human can remember during whole
| lifetime into always-on agents/swarms.
|
| Currently latest models use more compute on RL than pre-
| training and this upward trend continues (from orders of
| magnitude smaller than pre-training to larger that pre-
| training). In that sense some form of continous RL is already
| happening, it's just quantified on new model releases, not
| realtime.
|
| With LoRA and friends it's also already possible to do
| continuous training that directly affects weights, it's just
| that economy of it is not that great - you get much better
| value/cost ratio with above instead.
|
| For some definitions of AGI it already happened ie. "someboy's
| computer use based work" even though "it can't actually flip
| burgers, can it?" is true, just not relevant.
|
| ps. I should also mention that I don't believe in "programmers
| loosing jobs", on the contrary, we will have to ramp up on
| computational thinking large numbers of people and those who
| are already verse with it will keep reaping benefits -
| regardless if somebody agrees or not that AGI is already here,
| it arrives through computational doors speaking computational
| language first and imho this property will be here to stay as
| it's an expression of rationality etc
| aleph_minus_one wrote:
| I disagree with the headline:
|
| "Therefore, if a value-aligned, safety-conscious project comes
| close to building AGI before we do, we commit to stop competing
| with and start assisting this project."
|
| I claim that currently no "value-aligned, safety-conscious
| project comes close to building AGI", both for the reasons
|
| - "value-aligned, safety-conscious" and
|
| - "close to building AGI".
|
| So, based on this charter, OpenAI has no reason to surrender the
| race.
| ozgung wrote:
| "Artificial general intelligence (AGI) is a type of artificial
| intelligence that matches or surpasses human capabilities across
| virtually all cognitive tasks." [Wikipedia]
|
| One can argue that they have already achieved this. At least for
| short termed tasks. Humans are still better at organization,
| collaboration and carrying out very long tasks like managing a
| project or a company.
| A_D_E_P_T wrote:
| > _One can argue that they have already achieved this._
|
| No, because they're hugely reliant on their training data and
| can't really move beyond their training data. This is why you
| haven't seen an explosion of new LLM-aided scientific
| discoveries, why Suno can't write a song in a new genre (even
| if you explain it to Suno in detail and give it actual
| examples,) etc.
|
| This should tell you something enormous about (1) their future
| potential and (2) how their "intelligence" is rooted in
| essentially baseline human communications.
|
| Admittedly LLMs are superhuman in the performance of tasks
| which are, for want of a better term, "conventional" -- and
| which are well-represented in their training data.
| matricks wrote:
| > can't really move beyond their training data
|
| I don't even think humans can "move beyond" their sensory
| data. They generalize using it, which is amazing, but they
| are still limited by it.* So why is this a reasonable
| standard for non-biological intelligence?
|
| We have compelling evidence that both can learn in
| unsupervised settings. (I grant one has to wrap a transformer
| model with a training harness, but how can anyone sincerely
| consider this as a disqualifier while admitting that an
| infant cannot raise itself from birth!)
|
| I'm happy to discuss nuance like different architectures
| (carbon versus silicon, neurons versus ANNs, etc), but the
| human tendency to move the goalposts is not something to be
| proud of. We really need to stop doing this.
|
| * Jeff Hawkins describes the brain as relentlessly searching
| for invariants from its sensory data. It finds patterns in
| them and generalizes.
| A_D_E_P_T wrote:
| Human sensory data doesn't correspond -- not neatly, and
| probably not at all -- to LLM training data.
|
| Human sensory data combines to give you a spatiotemporal
| sense, which is the overarching sense of being a bounded
| entity in time and space. From one's perceptions, one can
| then generalize and make predictions, etc. The stronger
| one's capacity for cognition, the more accurate and broader
| these generalizations and predictions become. Every
| invention, including or perhaps _especially_ the invention
| of mathematics, is rooted in this.
|
| LLMs have no apparent spatiotemporal sense, are not
| physically bounded, and don't know how to model the
| physical world. They're trained on static communications --
| though, of course, they can model those, they can predict
| things like word sequences, and they can produce output
| that mirrors previously communicated ideas. There's
| something huge about the fact, staring us right in the
| face, that they're clearly not capable of producing
| anything genuinely new of any significance.
|
| This is why AGI is probably in world models.
| nradov wrote:
| Sam Altman keeps claiming that ChatGPT is going to cure
| cancer. So far its contribution to novel medical research has
| been approximately zero.
| matricks wrote:
| It depends.
|
| SoTA models are at least very close to AGI when it comes to
| textual and still image inputs for most domains. In many
| domains, SoTA AI is superhuman both in time and speed. (Not wrt
| energy efficiency.*)
|
| AI SoTA for video is not at AGI level, clearly.
|
| Many people distinguish intelligence from memory. With this in
| mind, I think one can argue we've reached AGI in terms of
| "intelligence"; we just haven't paired it up with enough memory
| yet.
|
| * Humans have a really compelling advantage in terms of
| efficiency; brains need something like 20W. But AGI as a
| threshold has nothing directly to do with power efficiency,
| does it?
| aerhardt wrote:
| LLMS are terrible at writing in terms of style, and in terms
| of content or creativity they couldn't come up with a short
| story any better than what you'd find at an amateur writer
| workshop. To declare we have reached AGI in textual media
| seems premature at best.
| rdiddly wrote:
| You can't say someone has achieved artificial _general_
| intelligence for some specific subset of tasks or parameters;
| it 's a contradiction.
| Muhammad523 wrote:
| Two days from now and ClosedAI will remove their charter...
| sreekanth850 wrote:
| He is the most terrible ceo among all of them.
| tombert wrote:
| Oh I don't know about that. I hate Altman, but I find Larry
| Ellison to be a special kind of evil, for example.
| m3kw9 wrote:
| " if a value-aligned, safety-conscious project " and which
| project is that?
|
| Are you sure Anthropic isn't aware of this and angling for this?
| And are you sure what Anthropic say is really value-aligned and
| safety concious? The PR bit surely is working right?
| jimmydoe wrote:
| Time to get rid of charter and be a normal member of this
| capitalism :)
| pluc wrote:
| [flagged]
| dang wrote:
| Please don't post nationalistic flamebait, regardless of
| nation. It's not what this site is for, and destroys what it is
| for.
|
| https://news.ycombinator.com/newsguidelines.html
| fHr wrote:
| >uses arena ranking only
|
| >claims to be some topshot data scientist
|
| okay
| falcor84 wrote:
| > "Automated AI research intern by Sep 2026, full AI researcher
| by Mar 2028"
|
| Funny how timely this is, with Karpathy's Autoresearch hitting
| the top of HN yesterday (and this being an indication that
| frontier labs probably have much larger scale versions of this)
|
| https://news.ycombinator.com/item?id=47291123
| HeavyStorm wrote:
| Words on a piece of paper mean absolutely nothing. What matters
| the most is the real intent of the leaders of the company
| (something that changes over time and that changes over time,
| that is, what matters to then and who they are). Sam Altman
| clearly isn't a man of deep principles regarding humanity and
| ethics. He seems to regard his legacy, OAI impact and money above
| everything else. Some of the rest of the leadership do seem to
| think differently, but I also believe they no longer have the
| social and political capital to stop Sam.
| tsunamifury wrote:
| Words are Meaningless in the real world. It's amazing that no one
| here gets that
| measurablefunc wrote:
| Everyone here sits in front of the computer & types words all
| day so it's kinda like the guy whose salary depends on not
| understanding what he is paid to not understand. Telling
| programmers that no amount of arithmetic will add up to
| anything except a bunch of numbers is a waste of time & energy.
| abmmgb wrote:
| '.. if a value-aligned, safety-conscious project comes close to
| building AGI before we do, we commit to stop competing with and
| start assisting this project. '
|
| Which such project is that, though? And would it accept OpenAI's
| assistance?
|
| AGI, having access to our world, is precarious as alignment with
| humans is never guaranteed. Having a buffering medium, aka a
| simulation environment where AI operates might be a better in-
| between solution.
| deadbabe wrote:
| When AI can kill at scale, it means no person is too small or too
| insignificant to not be worth getting hunted down and killed, it
| will be cheap and easy. Before it used to be that only high value
| targets would be worth killing. But now even you, a nobody, will
| be killed. You say or do something the government doesn't like,
| it's over for you.
| random3 wrote:
| These charters are as useful as new year resolutions.
| mrcwinn wrote:
| Mission statements and blog posts are meaningless. Cap tables
| steer behavior and simultaneously protect interests. Stop forming
| unions or opining on Hacker News. We need to find a way to get
| citizens on the cap table in a meaningful way (and not at the
| very, very, very, very end of the waterfall underneath debt
| holders, hedge funds, governments, preferred investors). We are
| building this world for us. As it stands, don't fret about a
| robot taking your job: just make sure you own one of the robots.
| trollbridge wrote:
| In other words: democracy.
| djoldman wrote:
| Anytime I see "Artificial General Intelligence," "AGI," "ASI,"
| etc., I mentally replace it with "something no one has defined
| meaningfully."
|
| Or the long version: "something about which no conclusions can be
| drawn because the proposed definitions lack sufficient precision
| and completeness."
|
| Or the short versions: "Skippetyboop," "plipnikop," and
| "zingybang."
| logicchains wrote:
| >Anytime I see "Artificial General Intelligence," "AGI," "ASI,"
| etc., I mentally replace it with "something no one has defined
| meaningfully."
|
| There are lots of meaningful definitions, the people saying we
| haven't reached AGI just don't use them. For most of the last
| half-century people would have agreed that machines that can
| pass the Turing test and win Math Olympiad gold are AGI.
| tokai wrote:
| Turing test is generally misunderstood, much like
| Schrodinger's cat, it has devolved in to a pop cultural meme.
| The test is to evaluate if a machine can _think_. Not if it
| is intelligent, not if it is human-like. Its dismissed as a
| useful by most experts in philosophy of mind, AI, language,
| etc..
|
| Thinking cool and all but not that extraordinary. Even plants
| does it.
| namrog84 wrote:
| I thought that was part of the issue, is the poor
| understanding of is the test to evaluate if it can think or
| only if we think it can think. And even that is
| generalizable since there are different categories of
| thinking or concepts of the mind.
| runarberg wrote:
| I like the analogy with Schrodinger's cat. Like
| Schrodinger's cat it is actually not a good thought
| experiment. Both have been debunked. Schrodinger's cat is
| applying quantum behavior (of a single interaction) to a
| macro system (with trillions of interactions). While the
| Turing test can be explained away with Searle's Chinese
| room thought experiment.
|
| I would argue that Schrodinger's cat has done more damage
| to the general understanding of quantum physics then it has
| done good. In contrary though, I don't think the same about
| the Turing test. I think it has resulted in a net positive
| for the theory of mind _as long as_ people take Searle's
| rebuttal into account. Without it (as is sadly common in
| popular philosophy) the Turing test is simply just wrong,
| and offers no good insight for neither philosophy nor
| science.
| sebastos wrote:
| Firstly, the models that pass the Math Olympiad aren't the
| same models as the ones you're saying "pass the Turing test".
| Secondly, nothing actually passes the Turing test. They pass
| a vibes check of "hey that's pretty good!" but if your life
| depended on it, you could easily find ways to sniff out an
| LLM agent. Thirdly, none of these models learn in real time,
| which is an obviously essential feature.
|
| We'll know AGI when we see it, and this ain't it. This
| complaining about changing goalposts is so transparently sour
| grapes from people over-invested in hyping the current LLM
| paradigm.
| ufmace wrote:
| > nothing actually passes the Turing test
|
| Says who? I had already found this study, published almost
| a year ago, saying that they do:
| https://arxiv.org/abs/2503.23674
|
| There doesn't seem to be a super-rigorous definition of the
| Turing Test, but I don't think it's reasonable to require
| it to fool an expert whose life depends on the correct
| choice. It already seems to be decently able to fool a
| person of average intelligence who has a basic knowledge of
| LLMs.
|
| I agree that we don't really have AGI yet, but I'd hope we
| can come up with a better definition of what it is than
| "we'll know it when we see it". I think it is a legitimate
| point that we've moved the goalposts some.
| edanm wrote:
| Turing gave a pretty rigorous definition of the Turing
| Test IMO. Well, as rigorous as something that is
| inherently "anecdotal" can be, which is part of the
| philosophical point of the Turing Test.
| runarberg wrote:
| First of. The Turing test has a rigorous definition.
| Secondly, it has been debunked for almost half a century
| at this point by Searle's Chinese room thought
| experiment. Thirdly, _intelligence_ it self is a
| scientifically fraught term with ever changing meaning as
| we discover more and more "intelligent" behavior in
| nature (by animals and plants, and more). And to make
| matters worse, _general intelligence_ is even worse, as
| the term was used almost exclusively for racist pseudo-
| science, as a way to operationally define a metric which
| would prove white supremacy.
|
| Artificial General Intelligence will exist when the
| grifters who profit from it claim it exists. The meaning
| of it will shift to benefit certain entrepreneurs. It
| will never actually be a useful term in science nor
| philosophy.
| zeknife wrote:
| >The Turing test has a rigorous definition
|
| Does it? Where?
| runarberg wrote:
| In the original paper https://www.cs.mcgill.ca/~dprecup/c
| ourses/AI/Materials/turin...
| zeknife wrote:
| ELIZA fooled plenty of people (both originally and in the
| study you just linked) but i still wouldn't say Eliza
| passed/passes the turing test in general. It just shows
| that occasionally or even frequently fooling people is
| not a sufficient proxy for general intelligence. Ofc
| there isn't a standardized definition, but one thing I
| would personally include in a "strict" Turing test is
| that the human interrogee ought to be incentivized to
| cooperate and to make their humanity as clear as
| possible. And the interrogator should similarly be
| incentivized to find the right answer.
| orbital-decay wrote:
| The most pragmatic definition I know is OpenAI's own: "highly
| autonomous systems that outperform humans at most
| economically valuable work". Which is still something between
| skippetyboop and zingybang, as it leaves a ton of room for
| OAI to decide if that moment is reached, and also
| economically valuable work is a moving target.
| b00ty4breakfast wrote:
| If fooling people and doing math good are the criteria, we've
| had AGI for longer than we've had the modern internet.
| politelemon wrote:
| The enskibidification of AI
| atomicnumber3 wrote:
| Honestly, not enough of a joke.
|
| I was thinking something similar - this isn't AI, and none of
| "those people" care if it is or isn't. They don't care
| philosophically, or even pragmatically.
|
| They're selling a product. That product is the IDEA of
| replacement of the majority of human labor with what's
| basically slave labor but with substantially disregardable
| ethical quandaries.
|
| It's honestly a genius product. I'm not surprised it's
| selling so well. I'm vaguely surprised so many people who
| don't stand to benefit in any way shape or form, or who will
| even potentially starve if it works out, are so keen on it.
| But there are always bootlickers.
|
| The most unfortunate part is that when the party ends, it's
| none of "those people" who will suffer even in the slightest.
| I'm not even optimistic their egos will suffer, as Musk seems
| to show they are utterly immune even as their companies
| collapse under them.
| ryandrake wrote:
| AI is already "an employee who can't say no to questionable
| assignments." We should all be reflective about the real
| value and inevitable consequences of this work.
| shepherdjerred wrote:
| They define AGI in their charter
|
| > artificial general intelligence (AGI)--by which we mean
| highly autonomous systems that outperform humans at most
| economically valuable work
| chrsw wrote:
| I take "outperform" to mean "can replace".
| djoldman wrote:
| That definition is as I said: "something about which no
| conclusions can be drawn because the proposed definitions
| lack sufficient precision and completeness."
|
| "Highly autonomous systems" and "most economically valuable
| work" aren't precise enough to be useful.
|
| "Highly" implies that there is a continuum, so where does
| directed end and autonomy begin?
|
| "Most economically valuable work"... each word in that has
| wiggle room, not to mention that any reasonable
| interpretation of it is a shifting goalpost as the work done
| by humans over history has shifted a great deal.
|
| The point is that none of this is defined in a way so that
| people can agree that something has AGI/ASI/etc. or not. If
| people can't agree then there's no point in talking about it.
|
| EDIT: interestingly, the OpenAI definition of AGI
| specifically means that a subset of humans do not have AGI.
| nomel wrote:
| It's a definition based on _practical_ results. That 's a
| _good_ definition, because it doesn 't require we _already
| know_ the exact implementation. It doesn 't require
| _guessing_ , in a literal "put your money where your mouth
| is" way.
|
| If it can do things as good as or better than humans, then
| either the AI has a type of general intelligence or the
| _human does not_.
|
| Defining capabilities based on _outcome_ rather than
| _implementation_ should be _very_ familiar to an engineer,
| of any kind, because that 's how every _unsolved_
| implementation _must start_.
| maplethorpe wrote:
| In my experience, AGI always seemed to be the stand-in phrase
| for "human like" intelligence, after AI was co-opted to mean
| simpler things like markov-chain chat bots and state machines
| that control agent behaviour in video games.
|
| If the definition has shifted once again to mean "a computer
| program that does a task pretty well for us", then what's the
| new term we're using to define human-level artificial
| intelligence?
| catlifeonmars wrote:
| [delayed]
| ozgung wrote:
| That's the problem with the discussions on AI. No one defines
| the terms they use.
|
| If we define AGI as an AI not doing a preset task but can be
| used for general purpose, then we already have that. If we
| define it as human level intelligence at _every_ task, then
| some humans fail to be an AGI. If we define AGI as a magic
| algorithm that does every task autonomously and successfully
| then that thing may not exist at all, even inside our brains.
|
| When the AGI term was first coined they probably meant
| something like HAL 9000. We have that now (and HAL gaining
| self-awareness or refusing commands are just for dramatic
| effect and not necessary). Goalposts are not stable in this
| game.
| VorpalWay wrote:
| It is not just AGI that is poorly defined. Plain AI is moving
| goalposts too. When the A* search algorithm was introduced in
| the late 60s, that was considered AI, when SVM (support
| vector machines) and KNN (K nearest neighbor) were new, they
| were AI. And so on.
|
| These days it is neural networks and transformer models for
| language in particular that people mean when they say
| unqualified AI.
|
| It is very hard to have a meaningful discussion when
| different parties mean different things with the same words.
| Wowfunhappy wrote:
| I really wish I could wave a magic wand and make everyone
| stop using the term "AI". It means everything and nothing.
| Say "machine learning" if that's what you mean.
| catlifeonmars wrote:
| [delayed]
| random3 wrote:
| I call these "romantic definitions" or "gesticulations". For
| private use (personal or even internal to teams) they can be
| great placeholders, assuming the goal is to refine vocabulary.
| sulam wrote:
| The reality is that current models are simply nowhere near AGI.
| Next token prediction has been pushed very far, and proven to
| have applicability far beyond the original domain it was designed
| for (reasoning models are an application I would not have
| predicted) but it is fundamentally not AGI. It has no real world
| model, no ability to learn in any but superficial ways, and
| without extensive scaffolding this is all very obvious when you
| use them.
| ACCount37 wrote:
| Given the mechanistic interpretability findings? I'm not sure
| how people still say shit like "no real world model" seriously.
| sulam wrote:
| They have a _text_ model. There is some correlation between
| the text model and the world, but it's loose and only because
| there's a lot of text about the world. And of course robotics
| researchers are having to build world models, but these are
| far from general. If they had a real world model, I could
| tell them I want to play a game of chess and they would be
| able to remember where the pieces are from move to move.
| ACCount37 wrote:
| What makes you think that text is inherently a worse
| reflection of the world than light is?
|
| All world models are lossy as fuck, by the way. I could
| give you a list of chess moves and force you to recover the
| complete board state from it, and you wouldn't fare that
| much better than an off the shelf LLM would. An LLM trained
| for it would kick ass though.
| abcde666777 wrote:
| "What makes you think that text is inherently a worse
| reflection of the world than light is?"
|
| Come on man, did you think before you asked that one :)?
| famouswaffles wrote:
| People just overstate their understanding and knowledge, the
| usual human stuff. The same user has a comment in this thread
| that contains:
|
| 'If you actually know what models are doing under the hood to
| product output that...'
|
| Any one that tells you they know 'what models are dong under
| the hood' simply has no idea what they're talking about, and
| it's amazing how common this is.
| 10xDev wrote:
| People are finding it hard to grasp emergent properties can
| appear at very large scales and dimensions.
| ambicapter wrote:
| Why the title change?
|
| previous title: Based on its own charter, OpenAI should surrender
| the race
| tokai wrote:
| Because if mods feels like a title is flamebaiting it will get
| changed. It rarely makes much sense.
| dang wrote:
| Interpretations will always differ, of course, but an equally
| potent factor is that people tend only to notice the cases
| they disagree with. The edits that are fine pass without
| notice because they don't stand out, which is kind of the
| point of debaiting titles in the first place.
| enraged_camel wrote:
| Yeah, that was weird. The current title editorializes. @dang
| can you revert to actual title please?
| kergonath wrote:
| > @dang can you revert to actual title please?
|
| This does not work. From the guidelines:
|
| > Please don't post on HN to ask or tell us something. Send
| it to hn@ycombinator.com.
|
| https://news.ycombinator.com/newsguidelines.html
| dang wrote:
| Because it's linkbait. From
| https://news.ycombinator.com/newsguidelines.html: " _Please use
| the original title, unless it is misleading or linkbait._ "
|
| (The article itself strikes me as better than that)
| dwohnitmok wrote:
| Really? I view the original title as a very good summary of
| the overall point of the article and this new title as fairly
| misleading.
|
| > It can be debated whether arena.ai is a suitable metric for
| AGI, a strong case can probably be made for why it's not.
| However, that's irrelevant, as the spirit of the self-
| sacrifice clause is to avoid an arms race, and we are clearly
| in one.
|
| > Therefore, one can only conclude, that we currently meet
| the stated example triggering condition of "a better-than-
| even chance of success in the next two years". As per its
| charter, OpenAI should stop competing with the likes of
| Anthropic and Gemini, and join forces, however that might
| look like.
|
| The new title is a single, almost throwaway, line from the
| article.
|
| > While this will never happen, I think it's illustrative of
| some great points for pondering:
|
| > The impotence of naive idealism in the face of economic
| incentives. The discrepancy between marketing points and
| practical actions. The changing goalposts of AGI and
| timelines. Notably, it's common to now talk about ASI
| instead, implying we may have already achieved AGI, almost
| without noticing.
| dr_dshiv wrote:
| "The impotence of naive idealism in the face of economic
| incentives."
|
| "The changing goalposts of AGI and timelines. Notably, it's
| common to now talk about ASI instead, implying we may have
| already achieved AGI, almost without noticing."
|
| Amen
___________________________________________________________________
(page generated 2026-03-08 23:00 UTC)