[HN Gopher] We're running out of benchmarks to upper bound AI ca...
       ___________________________________________________________________
        
       We're running out of benchmarks to upper bound AI capabilities
        
       Author : gmays
       Score  : 14 points
       Date   : 2026-04-10 20:16 UTC (2 hours ago)
        
 (HTM) web link (www.lesswrong.com)
 (TXT) w3m dump (www.lesswrong.com)
        
       | WarmWash wrote:
       | Start front loading the models with 5k, 10k, 50k, 100k tokens of
       | messy quasi related context, and then run the benchmarks.
       | 
       | These models are ridiculously powerful with a blank slate. It's
       | when they get loaded down with all the necessary (and inevitably
       | unnecessary) context to complete the task that they really start
       | to crumble and fold.
        
         | jballanc wrote:
         | We need benchmarks that can distinguish between continuous
         | learning and long-context extrapolation.
        
       | nikisweeting wrote:
       | We can definitely make harder evals, the problem is a good eval
       | set is indistinguishable from good training data / market edge,
       | so no one is incentivized to share their best eval sets publicly.
        
       | UltraSane wrote:
       | This is the least true thing ever. All LLMs are terrible at ARC-
       | AGI-3. Every video game can be used as a benchmark. You could
       | rank LLMs on how long they can keep a game of Dwarf Fortress
       | running or how fast they can beat GTA5.
        
       ___________________________________________________________________
       (page generated 2026-04-10 23:01 UTC)