[HN Gopher] Predictions from the METR AI scaling graph are based...
       ___________________________________________________________________
        
       Predictions from the METR AI scaling graph are based on a flawed
       premise
        
       Author : nsoonhui
       Score  : 43 points
       Date   : 2025-05-04 07:01 UTC (15 hours ago)
        
 (HTM) web link (garymarcus.substack.com)
 (TXT) w3m dump (garymarcus.substack.com)
        
       | Nivge wrote:
       | TL;DR - the benchmark depends on its specific dataset, and it
       | isn't a perfect representation to evaluate AI progress. That
       | doesn't mean it doesn't make sense, or doesn't have value.
        
       | hatefulmoron wrote:
       | I had assumed that the Y axis was corresponding to some
       | measurement of the LLM's ability to actually work/mull over a
       | task in a loop while making progress. In other words, I thought
       | it meant something like "you can leave Sonnet 3.7 for a whole
       | hour and it will meaningfully progress on a problem", but the
       | reality is less impressive. Serves me right for not looking at
       | the fine print.
        
       | dist-epoch wrote:
       | > Abject failure on a task that many adults could solve in a
       | minute
       | 
       | Maybe author should check before pressing "Publish" if the info
       | in the post is not already outdated.
       | 
       | ChatGPT passed the image generation test mentioned:
       | https://chatgpt.com/share/68171e2a-5334-8006-8d6e-dd693f2cec...
        
         | frotaur wrote:
         | Even excluding the fact that this image is simply to
         | illustrate, and it's really not the main point of the article,
         | in the chat you posted, ChatGPT actually failed again, because
         | the r's are not circled.
        
           | comex wrote:
           | That's true, but it illustrates a point about 'jagged
           | intelligence'. Just like there's a tendency to cherry-pick
           | the tasks AI is best at and equate it with general
           | intelligence, there's a counter-tendency to cherry-pick the
           | tasks AI is worst at and equate it with a general lack of
           | intelligence.
           | 
           | This case is especially egregious because of how there were
           | probably two different models involved. I assume Marcus'
           | images came from some AI service that followed what until
           | very recently was the standard pattern: you ask an LLM to
           | generate an image; the LLM goes and fluffs out your text,
           | then passes it to a completely separate diffusion-based image
           | generation model, which has only a rudimentary understanding
           | of English grammar. So of course his request for "words and
           | nothing else" was ignored. This is a real limitation of the
           | image generation model, but that has no relevance to the
           | strengths and weaknesses of the LLM itself. And 'AI will
           | replace humans' scenarios typically focus on text-based tasks
           | that use the LLM itself.
           | 
           | Arguably AI services are responsible for encouraging users to
           | think of what are really two separate models (LLM and image
           | generation) as a single 'AI'. But Marcus should know better.
           | 
           | And so it's not surprising that ChatGPT was able to produce
           | dramatically better results now that it has "native" image
           | generation, which supposedly uses the native multimodal
           | capabilities of the LLM (though rumors are that that
           | description is an oversimplification). The results are still
           | not correct. But it's a major advancement that the model now
           | respects grammar; it no longer just spots the word "fruit"
           | and generates a picture of fruit. Illustration or no, Marcus
           | is misrepresenting the state of the art by not including this
           | advancement.
           | 
           | If Marcus had used a recent ChatGPT output instead, the
           | comparison would be more fair, but still somewhat misleading.
           | Even with native capabilities, LLMs are simply worse at both
           | understanding and generating images than they are at
           | understanding and generating text. But again, text capability
           | matters much more. And you can't just assume that a model's
           | poor performance on images will correlate with poor
           | performance on text.
           | 
           | The thing is, I tend to agree with the substance of Marcus's
           | post, _including_ the part where portrayals of current AI
           | capabilities are suspect because they don 't pass the 'sniff
           | test', or in other words, because they don't take into
           | account how LLMs continue to fall down on some very basic
           | tasks. I just think the proper tasks for this evaluation
           | should be text-based. I'd say the original "count the number
           | of 'r's in strawberry" task is a decent example, even if it's
           | been patched, because it really showcases the 'confidently
           | wrong' issue that continues to plague LLMs.
        
         | croes wrote:
         | So OpenAI fixed that, but the next simple task on which AI
         | fails is just around the corner.
         | 
         | The problem is AI doesn't think and if a task is totally new it
         | doesn't produce the correct answer.
         | 
         | https://news.ycombinator.com/item?id=43800686
        
       | yorwba wrote:
       | > you could probably put together one reasonable collection of
       | word counting and question answering tasks with average human
       | time of 30 seconds and another collection with an average human
       | time of 20 minutes where GPT-4 would hit 50% accuracy on each.
       | 
       | So do this and pick the one where humans do best. I doubt that
       | doing so would show all progress to be illusory.
       | 
       | But it would certainly be interesting to know what the easiest
       | thing is that a human can do but current AIs struggle with.
        
         | xg15 wrote:
         | > _But it would certainly be interesting to know what the
         | easiest thing is that a human can do but current AIs struggle
         | with._
         | 
         | Still "Count the R's" apparently.
        
         | K0balt wrote:
         | The problem , really, is human cognitive dissonance. We draw
         | false conclusions that competence at some tasks implies
         | competence at another. It's not a universal human problem, we
         | intuit that a front end loader , just because it can dig really
         | well, is not therefore good at all other tasks. But when it
         | comes down to cognition, our models break down quickly.
         | 
         | I suspect this is because our proxies are predicated on a task
         | set that inherently includes the physical world, which at some
         | level connects all tasks and creates links between capabilities
         | that generally pervade our environment. LLMs do not exist in
         | this physical world, and are therefore not within the set of
         | things that can be reasoned about with those proxies.
         | 
         | This will probably gradually change with robotics, as the
         | competencies required to exist and function in the physical
         | world will (I postulate) generalize to other tasks in such a
         | way that it more closely matches the pattern that our
         | assumptions are based on.
         | 
         | Of course, if we segregate intelligence into isolated modules
         | for motility and cognition, this will not be the case as we
         | will not be taking advantage of that generalization. I think
         | that would be a big mistake, especially in light of the
         | hypotheses that the massive leap in capabilities of LLMs came
         | more from the training on things we weren't specifically trying
         | to achieve- the bulk of seemingly irrelevant data that unlocked
         | simple language processing into reasoning and world modeling.
        
           | the8472 wrote:
           | > LLMs do not exist in this physical world, and are therefore
           | not within the set of things that can be reasoned about with
           | those proxies.
           | 
           | Perhaps not the mainstream models, but deepmind has been
           | working on robotics models with simulated and physical RL for
           | years https://deepmind.google/discover/blog/rt-2-new-model-
           | transla...
        
           | mentalgear wrote:
           | what you are describing are world models and physical AI,
           | which has recently become much more mainstream after the
           | recent nvidia GDC.
        
       | Sharlin wrote:
       | > Unfortunately, literally none of the tweets we saw even
       | considered the possibility that a problematic graph specific to
       | software tasks might not generalize to literally all other
       | aspects of cognition.
       | 
       | How am I not surprised?
        
       ___________________________________________________________________
       (page generated 2025-05-04 23:01 UTC)