[HN Gopher] LMArena is a cancer on AI
___________________________________________________________________
LMArena is a cancer on AI
Author : jumploops
Score : 92 points
Date : 2026-01-07 04:40 UTC (18 hours ago)
(HTM) web link (surgehq.ai)
(TXT) w3m dump (surgehq.ai)
| observationist wrote:
| There's something deeply ironic about this being written by AI.
| Baitception, even.
| dust42 wrote:
| Oh my goodness yes, I almost missed it that the text is
| (mostly?) AI written. That said I agree that LMArena elo scores
| are pushing models in the wrong direction. They move more
| towards McDonald's than quality food.
| dk8996 wrote:
| Seems like they just raised 150m at 1.7B valuation. Crazy.
| koakuma-chan wrote:
| Who? LMArena? That's actually crazy.
| echelon wrote:
| Are they selling:
|
| A. model improvement tests, suites, and benchmarks
|
| B. data on competitors' evals
|
| C. test answer keys
|
| D. alpha to VC firms
|
| E. all of the above
|
| ???
| koakuma-chan wrote:
| Apparently they are selling model evaluations, powered by
| their volunteer users.
| Y_Y wrote:
| I'm taking the Red Cross public next. With the price of
| healthcare these days my earnings projections are uber-
| extreme.
| minimaxir wrote:
| Source: https://techcrunch.com/2026/01/06/lmarena-
| lands-1-7b-valuati...
| keketi wrote:
| We need a service that ranks AI model ranking services. Maybe
| powered by AI instead of humans?
| echelon wrote:
| Just look at Open(ugh)Router. That's a good, though not fully
| accurate, view of where dollars are going.
|
| It'd be nice if it were actually open and we could inspect all
| the statistics.
| a-dub wrote:
| maybe it would work if they could encourage end users to be
| rigorous? (ie, detect if they have the capability to rate well
| and then reward them when they do by comparing them against other
| highly rated raters of the same phenotype)
| sharkjacobs wrote:
| Any metric that can be targeted can be gamed
| kelseyfrog wrote:
| Then target it with metrics worth solving[1].
|
| 1. Ex https://mppbench.com/
| falcor84 wrote:
| But that seems to be measuring "superintelligence" rather
| than just AI, no?
| g947o wrote:
| > Voila: bold text, emojis, and plenty of sycophancy - every
| trick in the LMArena playbook! - to avoid answering the question
| it was asked.
|
| This is hard to swallow.
|
| I don't believe a single word this article says. Apparently the
| "real author" (the human being who wrote the original prompt to
| generate this article) only intend to use this article to
| generate clicks and engagement but don't care at all about what's
| in there.
| atleastoptimal wrote:
| The general conceit of this article, which is something that many
| frontier labs seem to be beginning to realize, is that the
| average human is no longer smart enough to provide sufficient
| signal to improve AI models.
| cyanydeez wrote:
| They need to spend money on actual experts to curate their data
| to improve.
|
| Instead, finance bros are convinced by the argument that number
| goes up.
| Terr_ wrote:
| Sometimes it feels like: def
| is_it_true(question): return
| profit_if_true(question) > profit_if_false(question)
|
| AI will make it cheaper, faster, better, no problem. You can
| eat the cake now _and_ save it for later.
| aspenmartin wrote:
| Wait you know that frontier labs do actually do this right?
| 8f2ab37a-ed6c wrote:
| Is that not exactly what https://www.mercor.com/ does?
| Y_Y wrote:
| But when you're a moron how can you distinguish?
|
| I'm being (mostly) serious, suppose you're a stuffed ahort
| trying to boost your valuation, how can you work out who's
| smart enough to train your LLM? (Never mind how to get them to
| work for you!)
| aspenmartin wrote:
| I do a lot of human evaluations. Lots of Bayesian /
| statistical models that can infer rater quality without
| ground truth labels. The other thing about preference data
| you have to worry about (which this article gets at) is:
| preferences of _who_? Human raters are a significantly biased
| population of people, different ages, genders, religions,
| cultures, etc all inform preferences. Lots of work being done
| to leverage and model this.
|
| Then for LMArena there is the host of other biases /
| construct validity: people are easily fooled, even PhD
| experts; in many cases it's easier for a model to learn how
| to persuade than actually learn the right answers.
|
| But a lot of dismissive comments as if frontier labs don't
| know this, they have some of the best talent in the world.
| They aren't perfect but they in a large sene know what
| they're doing and what the tradeoffs of various approaches
| are.
|
| Human annotations are an absolute nightmare for quality which
| is why coding agents are so nice: they're verifiable and so
| you can train them in a way closer to e.g. alphago without
| the ceiling of human performance
| fc417fc802 wrote:
| > in many cases it's easier for a model to learn how to
| persuade than actually learn the right answers
|
| So we should expect the models to eventually tend toward
| the same behaviors that politicians exhibit?
| atleastoptimal wrote:
| that's why Mercor is worth 2billion
| wongarsu wrote:
| Sure, on the surface judging the judge is just as hard as
| being the judge
|
| But at least the two examples of judging AI provided in the
| article can be solved by any moron by expending enough
| effort. Any moron can tell you what Dorothy says to Toto when
| entering Oz by just watching the first thirty minutes of the
| movie. And while validating answer B in the pan question
| takes some ninth-grade math (or a short trip to wikipedia),
| figuring out that a nine inch diameter circle is in fact not
| the same area as a 9x13 inch square is not rocket science.
| And with a bit of craft paper you could evaluate both answers
| even without math knowledge
|
| So the short answer is: with effort. You spend lots of effort
| on finding a good evaluator, so the evaluator can judge the
| LLM for you. Or take "average humans" and force them to spend
| more effort on evaluating each answer
| Yizahi wrote:
| Yep, it's like getting a commoner from the street evaluate a
| literature PhD in their native language. Sure, both know the
| language, but the depth difference of a specialist vs a
| generalist is too large. And neither we can't use AI to
| automatically evaluate this literature genius because real AI
| doesn't exist (yet), hence the programs can't understand the
| contents of text they output or input. Whoops. :)
| thorum wrote:
| Aside from Meta is there any reason to think the big AI labs are
| still using LMArena data for training? The weaknesses are well
| understood and with the shift to RL there are so many better ways
| to design a reward function.
| dk8996 wrote:
| Such as?
| nl wrote:
| I don't think anyone has _ever_ used it as training. But yes
| labs still do seem to target it as goal (which is a different
| thing).
| aucisson_masque wrote:
| > They're not reading carefully. They're not fact-checking, or
| even trying.
|
| It's not how I do, and I suppose how many people do. I
| specifically ask questions related to niche subjects that I know
| perfectly well and that is very easy for me to spot mistakes.
|
| The first time I used it, that's what came naturally to my mind.
| I believe it's the same for others.
| stared wrote:
| When they released GPT-4.5, it was miles ahead of others when it
| comes to its linguistic skills and insight. Yet, it was never at
| top of the arena - it felt that not everone was able to
| appreciate the edge.
| usef- wrote:
| When the Meta cheating scandal happened I was surprised how
| little of the attention was on this.
|
| Meta "cheated" on lmarena not by using a smarter model but by
| using one that was more verbose and friendly with excessive
| emojis.
| mirekrusin wrote:
| True and what you can realize/read between the lines is something
| deeper.
|
| LLMs are fallible. Humans are fallible. LLMs improve (and improve
| fast). Humans do not (overall, ie. "group of N experts in X", "N
| random internet people").
|
| All those "turing tests" will start flipping.
|
| Today it's "N random internet humans" score too low on those
| benchmarks, tomorrow it'll be "group of N expert humans in X"
| score too low.
___________________________________________________________________
(page generated 2026-01-07 23:00 UTC)