[HN Gopher] Predictions from the METR AI scaling graph are based...
___________________________________________________________________
Predictions from the METR AI scaling graph are based on a flawed
premise
Author : nsoonhui
Score : 43 points
Date : 2025-05-04 07:01 UTC (15 hours ago)
(HTM) web link (garymarcus.substack.com)
(TXT) w3m dump (garymarcus.substack.com)
| Nivge wrote:
| TL;DR - the benchmark depends on its specific dataset, and it
| isn't a perfect representation to evaluate AI progress. That
| doesn't mean it doesn't make sense, or doesn't have value.
| hatefulmoron wrote:
| I had assumed that the Y axis was corresponding to some
| measurement of the LLM's ability to actually work/mull over a
| task in a loop while making progress. In other words, I thought
| it meant something like "you can leave Sonnet 3.7 for a whole
| hour and it will meaningfully progress on a problem", but the
| reality is less impressive. Serves me right for not looking at
| the fine print.
| dist-epoch wrote:
| > Abject failure on a task that many adults could solve in a
| minute
|
| Maybe author should check before pressing "Publish" if the info
| in the post is not already outdated.
|
| ChatGPT passed the image generation test mentioned:
| https://chatgpt.com/share/68171e2a-5334-8006-8d6e-dd693f2cec...
| frotaur wrote:
| Even excluding the fact that this image is simply to
| illustrate, and it's really not the main point of the article,
| in the chat you posted, ChatGPT actually failed again, because
| the r's are not circled.
| comex wrote:
| That's true, but it illustrates a point about 'jagged
| intelligence'. Just like there's a tendency to cherry-pick
| the tasks AI is best at and equate it with general
| intelligence, there's a counter-tendency to cherry-pick the
| tasks AI is worst at and equate it with a general lack of
| intelligence.
|
| This case is especially egregious because of how there were
| probably two different models involved. I assume Marcus'
| images came from some AI service that followed what until
| very recently was the standard pattern: you ask an LLM to
| generate an image; the LLM goes and fluffs out your text,
| then passes it to a completely separate diffusion-based image
| generation model, which has only a rudimentary understanding
| of English grammar. So of course his request for "words and
| nothing else" was ignored. This is a real limitation of the
| image generation model, but that has no relevance to the
| strengths and weaknesses of the LLM itself. And 'AI will
| replace humans' scenarios typically focus on text-based tasks
| that use the LLM itself.
|
| Arguably AI services are responsible for encouraging users to
| think of what are really two separate models (LLM and image
| generation) as a single 'AI'. But Marcus should know better.
|
| And so it's not surprising that ChatGPT was able to produce
| dramatically better results now that it has "native" image
| generation, which supposedly uses the native multimodal
| capabilities of the LLM (though rumors are that that
| description is an oversimplification). The results are still
| not correct. But it's a major advancement that the model now
| respects grammar; it no longer just spots the word "fruit"
| and generates a picture of fruit. Illustration or no, Marcus
| is misrepresenting the state of the art by not including this
| advancement.
|
| If Marcus had used a recent ChatGPT output instead, the
| comparison would be more fair, but still somewhat misleading.
| Even with native capabilities, LLMs are simply worse at both
| understanding and generating images than they are at
| understanding and generating text. But again, text capability
| matters much more. And you can't just assume that a model's
| poor performance on images will correlate with poor
| performance on text.
|
| The thing is, I tend to agree with the substance of Marcus's
| post, _including_ the part where portrayals of current AI
| capabilities are suspect because they don 't pass the 'sniff
| test', or in other words, because they don't take into
| account how LLMs continue to fall down on some very basic
| tasks. I just think the proper tasks for this evaluation
| should be text-based. I'd say the original "count the number
| of 'r's in strawberry" task is a decent example, even if it's
| been patched, because it really showcases the 'confidently
| wrong' issue that continues to plague LLMs.
| croes wrote:
| So OpenAI fixed that, but the next simple task on which AI
| fails is just around the corner.
|
| The problem is AI doesn't think and if a task is totally new it
| doesn't produce the correct answer.
|
| https://news.ycombinator.com/item?id=43800686
| yorwba wrote:
| > you could probably put together one reasonable collection of
| word counting and question answering tasks with average human
| time of 30 seconds and another collection with an average human
| time of 20 minutes where GPT-4 would hit 50% accuracy on each.
|
| So do this and pick the one where humans do best. I doubt that
| doing so would show all progress to be illusory.
|
| But it would certainly be interesting to know what the easiest
| thing is that a human can do but current AIs struggle with.
| xg15 wrote:
| > _But it would certainly be interesting to know what the
| easiest thing is that a human can do but current AIs struggle
| with._
|
| Still "Count the R's" apparently.
| K0balt wrote:
| The problem , really, is human cognitive dissonance. We draw
| false conclusions that competence at some tasks implies
| competence at another. It's not a universal human problem, we
| intuit that a front end loader , just because it can dig really
| well, is not therefore good at all other tasks. But when it
| comes down to cognition, our models break down quickly.
|
| I suspect this is because our proxies are predicated on a task
| set that inherently includes the physical world, which at some
| level connects all tasks and creates links between capabilities
| that generally pervade our environment. LLMs do not exist in
| this physical world, and are therefore not within the set of
| things that can be reasoned about with those proxies.
|
| This will probably gradually change with robotics, as the
| competencies required to exist and function in the physical
| world will (I postulate) generalize to other tasks in such a
| way that it more closely matches the pattern that our
| assumptions are based on.
|
| Of course, if we segregate intelligence into isolated modules
| for motility and cognition, this will not be the case as we
| will not be taking advantage of that generalization. I think
| that would be a big mistake, especially in light of the
| hypotheses that the massive leap in capabilities of LLMs came
| more from the training on things we weren't specifically trying
| to achieve- the bulk of seemingly irrelevant data that unlocked
| simple language processing into reasoning and world modeling.
| the8472 wrote:
| > LLMs do not exist in this physical world, and are therefore
| not within the set of things that can be reasoned about with
| those proxies.
|
| Perhaps not the mainstream models, but deepmind has been
| working on robotics models with simulated and physical RL for
| years https://deepmind.google/discover/blog/rt-2-new-model-
| transla...
| mentalgear wrote:
| what you are describing are world models and physical AI,
| which has recently become much more mainstream after the
| recent nvidia GDC.
| Sharlin wrote:
| > Unfortunately, literally none of the tweets we saw even
| considered the possibility that a problematic graph specific to
| software tasks might not generalize to literally all other
| aspects of cognition.
|
| How am I not surprised?
___________________________________________________________________
(page generated 2025-05-04 23:01 UTC)