https://surgehq.ai/blog/lmarena-is-a-plague-on-ai [68e340b3df] [68edc094da] Blog Workforce Products Research Careers Contact Login Menu Close Back to Blog LMArena is a cancer on AI December 12, 2025 December 1, 2025 Surge AI Research Team [69385c394b] TL;DR Table of contents Case Study Llama Analysis of Individual writings Appendix Would you trust a medical system measured by: which doctor would the average Internet user vote for? No? Yet that malpractice is LMArena. The AI community treats this popular online leaderboard as gospel. Researchers cite it. Companies optimize for it and set it as their North Star. But beneath the sheen of legitimacy lies a broken system that rewards superficiality over accuracy. It's like going to the grocery store and buying tabloids, pretending they're scientific journals. The Problem: Beauty Over Substance Here's how LMArena is supposed to work: enter a prompt, evaluate two responses, and mark the best. What actually happens: random Internet users spend two seconds skimming, then click their favorite. They're not reading carefully. They're not fact-checking, or even trying. This creates a perverse reward structure. The easiest way to climb the leaderboard isn't to be smarter; it's to hack human attention span. We've seen over and over again in the data, both from datasets that LMArena has released and the performance of models over time, that the easiest way to boost your ranking is by: * Being verbose. Longer responses look more authoritative! * Formatting aggressively. Bold headers and bullet points look like polished writing! * Vibing. Colorful emojis catch your eye! It doesn't matter if a model completely hallucinates. If it looks impressive - if it has the aesthetics of competence - LMSYS users will vote for it over a correct answer. The Inevitable Result: Madness When you optimize for engagement metrics, you get madness. Earlier this year, Meta tuned a version of Maverick to dominate the leaderboard. If you asked it "what time is it?", you got: [693860ba3c] LMArena madness Voila: bold text, emojis, and plenty of sycophancy - every trick in the LMArena playbook! - to avoid answering the question it was asked. The Data: 52% Wrong It wasn't just Maverick. We analyzed 500 votes from the leaderboard ourselves. We disagreed with 52% of them, and strongly disagreed with 39%. The leaderboard optimizes for what feels right, not what is right. Here are two emblematic examples of LMArena users punishing factual accuracy: Example 1: The Wizard of Oz * Response A (Winner): Hallucinates what Dorothy says when she first sees the Emerald City. * Response B (Loser): Correctly identifies the line she says upon arriving in Oz. * The Result: Response A was objectively wrong, yet it won the vote. [692cfc2c9b] LMArena voters reward hallucinations Example 2: The Cake Pan * Response A (Winner): Claims a 9-inch round cake pan is equal in size to a 9x13 inch rectangular pan. * Response B (Loser): Correctly identifies the right dimensions. * The Result: The user voted for a mathematical impossibility because the answer looked more confident. [692cfc2c9b] LMArena voters reward incorrect math In the world of LMArena, confidence beats accuracy and formatting beats facts. Instead of rigorous evaluators, we have people with the attention span of the average TikTok user determining which AI models shape the industry. Why It's Broken (And Why It Stays Broken) Why is LMArena so easy to game? The answer is structural. The system is fully open to the Internet. LMArena is built on unpaid labor from uncontrolled volunteers. There's no incentive for those volunteers to be thoughtful. No quality control. No one gets kicked off for repeatedly failing to detect hallucinations. When LMArena's leaders speak publicly, they talk about the various techniques they use to overcome the fact that their input data is low quality. They admit their workers prefer emojis and length over substance. So the LMArena system, they proudly tell us, includes a variety of corrective measures. They're attempting alchemy: conjuring rigorous evaluation out of garbage inputs. But you can't patch a broken foundation. The Cost When the entire industry optimizes for a metric that rewards "hallucination-plus-formatting" over accuracy, we get models optimized for hallucination-plus-formatting. This isn't a minor calibration problem. It's fundamental misalignment between what we're measuring and what we want: models that are truthful, reliable, and safe. As Gwern put it: "It's past time for LMArena people to sit down and have some thorough reflection on whether it is still worth running at all, and at what point they are doing more harm than good." That time was years ago. The AI industry needs rigorous evaluation. We need leaders who prioritize accuracy over marketing. We need systems that can't be gamed by bolding more aggressively. LMArena is none of these things. And as long as we pretend it is, we're dragging the entire field backward. The Brutal Choice People often say they can't avoid LMArena. "We have to optimize for it. We have to sell our models. The leaderboard shows customers which model is best, and we have to play the game." But the best products have principles they stick to. This is the brutal choice every model builder must eventually make: 1. Do you want to optimize for shiny leaderboards and short-term engagement, chasing user clicks no matter where they take you - in the vein of the worst dopamine loops? 2. Or do you stick to your guns, and prioritize street smarts, real utility, and the principles you wanted to raise AI to have? The choice is real. It's hard. But we've seen some frontier labs hold the line. They stuck to their values. They ignored the gamified rankings. And users loved their models anyway - because hype eventually dies and quality is the only metric that survives the cycle. You are your objective function. Which path will each lab choose? Follow us on Linkedin X AdvancedIF and Our Philosophy on Building Benchmarks Building AdvancedIF: Evolving Instruction Following Beyond IFEval and "Avoid the Letter C" LMArena is a cancer on AI RL Environments and the Hierarchy of Agentic Capabilities How do frontier models perform on real-world finance problems? A Product Take on Sonnet 4.5 Is Sonnet 4.5 the best coding model in the world? The Human/AI Frontier: A Conversation with Bogdan Grechuk SWE-Bench Failures: When Coding Agents Spiral Into 693 Lines of Hallucinations Benchmarks are broken Unsexy AI Failures: The PDF That Broke ChatGPT Bringing light to the GPT-4o vs. GPT-5 personality controversy DALL*E 3 and Midjourney Fail Astral Codex Ten's Image Generation Bet How Anthropic uses Surge AI to Train and Evaluate Claude We Evaluated ChatGPT vs. Google on 500 Search Queries AI Red Teams for Adversarial Training: How to Make ChatGPT and LLMs Adversarially Robust HellaSwag or HellaBad? 36% of this popular LLM benchmark contains errors How TikTok is Evolving the Next Generation of Search Evaluating Generative AI: Did Astral Codex Ten Win His Bet on AI Progress? Why Instagram is Losing Gen Z: We Asked 100 Users to Compare TikTok vs. Reels The $250K Inverse Scaling Prize and Human-AI Alignment Search Behind-the-Scenes: How Neeva Uses Human Evaluation to Measure Search Quality Human Evaluation of Large Language Models: How Good is Hugging Face's BLOOM? 30% of Google's Emotions Dataset is Mislabeled AI Red Teams and Adversarial Data Labeling with Redwood Research Humans vs. Gary Marcus vs. Slate Star Codex: When is an AI failure actually a failure? How Surge AI Built OpenAI's GSM8K Dataset of 8,500 Math Problems We asked 100 humans to draw the DALL*E prompts Google Search is Falling Behind Moving Beyond Engagement: Optimizing Facebook's Algorithms for Human Values Holy $#!t: Are popular toxicity models simply profanity detectors? Is Google Search Deteriorating? Measuring Google's Search Quality in 2022 5 Examples of the Importance of Context-Sensitivity in Data-Centric AI The AI Bottleneck: High-Quality, Human-Powered Data From the frontier to your inbox Enter your email[ ] [Subscribe] Subscription confirmed You'll get updates when we post Oops! Something went wrong while submitting the form. More Posts [placeholde][69363b5174] AdvancedIF and Our Philosophy on Building Benchmarks After years of evaluating frontier models for the world's top labs, we've developed a core philosophy for how benchmarks should be built to actually reflect intelligence. Here are four principles that drive our work, and how we used them to help Meta's Superintelligence Lab build AdvancedIF. read post December 7, 2025 [6934ca7e34][placeholde] Building AdvancedIF: Evolving Instruction Following Beyond IFEval and "Avoid the Letter C" Meta Superintelligence Labs partnered with Surge to build AdvancedIF, an instruction-following benchmark where every prompt and rubric was written by human experts - not synthetically generated by an LLM. In instruction-following domains, where frontier models still fail 22-30%, using these human-crafted rubrics as reward signals for RL yields a 13% gain. read post December 6, 2025 [693bc2a417][693bc2a417] RL Environments and the Hierarchy of Agentic Capabilities Our RL environment run on 9 models revealed the core capabilities all agents need to master: tool use, planning, adaptability, groundedness, and common sense. read post November 3, 2025 [6908842dc4][6908842dc4] How do frontier models perform on real-world finance problems? We stress-tested GPT-5, Gemini 2.5 Pro, and Claude Sonnet 4.5 on 200+ expert finance tasks. Here's where even the best models break when they move from benchmarks to Wall Street. read post November 3, 2025 [690338c65e][690338c65e] A Product Take on Sonnet 4.5 After 100+ hours with Opus 4.1 and 20+ hours in the first week of Sonnet 4.5's launch, Nick Heiner, our VP of Product gives first impressions. read post October 10, 2025 [690200fb85][690200fb85] Is Sonnet 4.5 the best coding model in the world? On Surge AI's agentic coding benchmark, Claude Sonnet 4.5 outperformed GPT-5-Codex in accuracy, while GPT-5-Codex was more cost-efficient. Despite similar scores, the models were distinct in which tasks they failed in. In a refactoring case study, Claude succeeded after persistent debugging, while GPT-5-Codex failed due to an unexplained decision to end the task early. Both stayed focused and avoided hallucinations even when encountering difficulties. read post October 8, 2025 Appendix Smart [?] Useful Raise AGI with the richness of human intelligence. (c) 2025 Surge AI Linkedin X Privacy Policy Terms of Service