[HN Gopher] We're running out of benchmarks to upper bound AI ca...
___________________________________________________________________
We're running out of benchmarks to upper bound AI capabilities
Author : gmays
Score : 14 points
Date : 2026-04-10 20:16 UTC (2 hours ago)
(HTM) web link (www.lesswrong.com)
(TXT) w3m dump (www.lesswrong.com)
| WarmWash wrote:
| Start front loading the models with 5k, 10k, 50k, 100k tokens of
| messy quasi related context, and then run the benchmarks.
|
| These models are ridiculously powerful with a blank slate. It's
| when they get loaded down with all the necessary (and inevitably
| unnecessary) context to complete the task that they really start
| to crumble and fold.
| jballanc wrote:
| We need benchmarks that can distinguish between continuous
| learning and long-context extrapolation.
| nikisweeting wrote:
| We can definitely make harder evals, the problem is a good eval
| set is indistinguishable from good training data / market edge,
| so no one is incentivized to share their best eval sets publicly.
| UltraSane wrote:
| This is the least true thing ever. All LLMs are terrible at ARC-
| AGI-3. Every video game can be used as a benchmark. You could
| rank LLMs on how long they can keep a game of Dwarf Fortress
| running or how fast they can beat GTA5.
___________________________________________________________________
(page generated 2026-04-10 23:01 UTC)