[HN Gopher] LMArena is a cancer on AI
       ___________________________________________________________________
        
       LMArena is a cancer on AI
        
       Author : jumploops
       Score  : 92 points
       Date   : 2026-01-07 04:40 UTC (18 hours ago)
        
 (HTM) web link (surgehq.ai)
 (TXT) w3m dump (surgehq.ai)
        
       | observationist wrote:
       | There's something deeply ironic about this being written by AI.
       | Baitception, even.
        
         | dust42 wrote:
         | Oh my goodness yes, I almost missed it that the text is
         | (mostly?) AI written. That said I agree that LMArena elo scores
         | are pushing models in the wrong direction. They move more
         | towards McDonald's than quality food.
        
       | dk8996 wrote:
       | Seems like they just raised 150m at 1.7B valuation. Crazy.
        
         | koakuma-chan wrote:
         | Who? LMArena? That's actually crazy.
        
           | echelon wrote:
           | Are they selling:
           | 
           | A. model improvement tests, suites, and benchmarks
           | 
           | B. data on competitors' evals
           | 
           | C. test answer keys
           | 
           | D. alpha to VC firms
           | 
           | E. all of the above
           | 
           | ???
        
             | koakuma-chan wrote:
             | Apparently they are selling model evaluations, powered by
             | their volunteer users.
        
               | Y_Y wrote:
               | I'm taking the Red Cross public next. With the price of
               | healthcare these days my earnings projections are uber-
               | extreme.
        
         | minimaxir wrote:
         | Source: https://techcrunch.com/2026/01/06/lmarena-
         | lands-1-7b-valuati...
        
       | keketi wrote:
       | We need a service that ranks AI model ranking services. Maybe
       | powered by AI instead of humans?
        
         | echelon wrote:
         | Just look at Open(ugh)Router. That's a good, though not fully
         | accurate, view of where dollars are going.
         | 
         | It'd be nice if it were actually open and we could inspect all
         | the statistics.
        
       | a-dub wrote:
       | maybe it would work if they could encourage end users to be
       | rigorous? (ie, detect if they have the capability to rate well
       | and then reward them when they do by comparing them against other
       | highly rated raters of the same phenotype)
        
       | sharkjacobs wrote:
       | Any metric that can be targeted can be gamed
        
         | kelseyfrog wrote:
         | Then target it with metrics worth solving[1].
         | 
         | 1. Ex https://mppbench.com/
        
           | falcor84 wrote:
           | But that seems to be measuring "superintelligence" rather
           | than just AI, no?
        
       | g947o wrote:
       | > Voila: bold text, emojis, and plenty of sycophancy - every
       | trick in the LMArena playbook! - to avoid answering the question
       | it was asked.
       | 
       | This is hard to swallow.
       | 
       | I don't believe a single word this article says. Apparently the
       | "real author" (the human being who wrote the original prompt to
       | generate this article) only intend to use this article to
       | generate clicks and engagement but don't care at all about what's
       | in there.
        
       | atleastoptimal wrote:
       | The general conceit of this article, which is something that many
       | frontier labs seem to be beginning to realize, is that the
       | average human is no longer smart enough to provide sufficient
       | signal to improve AI models.
        
         | cyanydeez wrote:
         | They need to spend money on actual experts to curate their data
         | to improve.
         | 
         | Instead, finance bros are convinced by the argument that number
         | goes up.
        
           | Terr_ wrote:
           | Sometimes it feels like:                   def
           | is_it_true(question):              return
           | profit_if_true(question) > profit_if_false(question)
           | 
           | AI will make it cheaper, faster, better, no problem. You can
           | eat the cake now _and_ save it for later.
        
           | aspenmartin wrote:
           | Wait you know that frontier labs do actually do this right?
        
           | 8f2ab37a-ed6c wrote:
           | Is that not exactly what https://www.mercor.com/ does?
        
         | Y_Y wrote:
         | But when you're a moron how can you distinguish?
         | 
         | I'm being (mostly) serious, suppose you're a stuffed ahort
         | trying to boost your valuation, how can you work out who's
         | smart enough to train your LLM? (Never mind how to get them to
         | work for you!)
        
           | aspenmartin wrote:
           | I do a lot of human evaluations. Lots of Bayesian /
           | statistical models that can infer rater quality without
           | ground truth labels. The other thing about preference data
           | you have to worry about (which this article gets at) is:
           | preferences of _who_? Human raters are a significantly biased
           | population of people, different ages, genders, religions,
           | cultures, etc all inform preferences. Lots of work being done
           | to leverage and model this.
           | 
           | Then for LMArena there is the host of other biases /
           | construct validity: people are easily fooled, even PhD
           | experts; in many cases it's easier for a model to learn how
           | to persuade than actually learn the right answers.
           | 
           | But a lot of dismissive comments as if frontier labs don't
           | know this, they have some of the best talent in the world.
           | They aren't perfect but they in a large sene know what
           | they're doing and what the tradeoffs of various approaches
           | are.
           | 
           | Human annotations are an absolute nightmare for quality which
           | is why coding agents are so nice: they're verifiable and so
           | you can train them in a way closer to e.g. alphago without
           | the ceiling of human performance
        
             | fc417fc802 wrote:
             | > in many cases it's easier for a model to learn how to
             | persuade than actually learn the right answers
             | 
             | So we should expect the models to eventually tend toward
             | the same behaviors that politicians exhibit?
        
           | atleastoptimal wrote:
           | that's why Mercor is worth 2billion
        
           | wongarsu wrote:
           | Sure, on the surface judging the judge is just as hard as
           | being the judge
           | 
           | But at least the two examples of judging AI provided in the
           | article can be solved by any moron by expending enough
           | effort. Any moron can tell you what Dorothy says to Toto when
           | entering Oz by just watching the first thirty minutes of the
           | movie. And while validating answer B in the pan question
           | takes some ninth-grade math (or a short trip to wikipedia),
           | figuring out that a nine inch diameter circle is in fact not
           | the same area as a 9x13 inch square is not rocket science.
           | And with a bit of craft paper you could evaluate both answers
           | even without math knowledge
           | 
           | So the short answer is: with effort. You spend lots of effort
           | on finding a good evaluator, so the evaluator can judge the
           | LLM for you. Or take "average humans" and force them to spend
           | more effort on evaluating each answer
        
         | Yizahi wrote:
         | Yep, it's like getting a commoner from the street evaluate a
         | literature PhD in their native language. Sure, both know the
         | language, but the depth difference of a specialist vs a
         | generalist is too large. And neither we can't use AI to
         | automatically evaluate this literature genius because real AI
         | doesn't exist (yet), hence the programs can't understand the
         | contents of text they output or input. Whoops. :)
        
       | thorum wrote:
       | Aside from Meta is there any reason to think the big AI labs are
       | still using LMArena data for training? The weaknesses are well
       | understood and with the shift to RL there are so many better ways
       | to design a reward function.
        
         | dk8996 wrote:
         | Such as?
        
         | nl wrote:
         | I don't think anyone has _ever_ used it as training. But yes
         | labs still do seem to target it as goal (which is a different
         | thing).
        
       | aucisson_masque wrote:
       | > They're not reading carefully. They're not fact-checking, or
       | even trying.
       | 
       | It's not how I do, and I suppose how many people do. I
       | specifically ask questions related to niche subjects that I know
       | perfectly well and that is very easy for me to spot mistakes.
       | 
       | The first time I used it, that's what came naturally to my mind.
       | I believe it's the same for others.
        
       | stared wrote:
       | When they released GPT-4.5, it was miles ahead of others when it
       | comes to its linguistic skills and insight. Yet, it was never at
       | top of the arena - it felt that not everone was able to
       | appreciate the edge.
        
       | usef- wrote:
       | When the Meta cheating scandal happened I was surprised how
       | little of the attention was on this.
       | 
       | Meta "cheated" on lmarena not by using a smarter model but by
       | using one that was more verbose and friendly with excessive
       | emojis.
        
       | mirekrusin wrote:
       | True and what you can realize/read between the lines is something
       | deeper.
       | 
       | LLMs are fallible. Humans are fallible. LLMs improve (and improve
       | fast). Humans do not (overall, ie. "group of N experts in X", "N
       | random internet people").
       | 
       | All those "turing tests" will start flipping.
       | 
       | Today it's "N random internet humans" score too low on those
       | benchmarks, tomorrow it'll be "group of N expert humans in X"
       | score too low.
        
       ___________________________________________________________________
       (page generated 2026-01-07 23:00 UTC)