[HN Gopher] Chatbot Arena: Benchmarking LLMs in the Wild with El...
___________________________________________________________________
Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
Author : MMMercy2
Score : 17 points
Date : 2023-05-03 19:26 UTC (3 hours ago)
(HTM) web link (lmsys.org)
(TXT) w3m dump (lmsys.org)
| weichiang wrote:
| Surprised to learn StableLM is worse than plain LLaMA. link to
| their leaderboard: leaderboard.lmsys.org
| circuit10 wrote:
| I've heard that it's really bad for it's size
| lee101 wrote:
| [dead]
| HoshinoAI wrote:
| It's glad to see the old technique is used for new models.
|
| You may also learn from 7.33 dota update which uses a new ranking
| algorithm called Glicko.
| zhisbug wrote:
| could you provide any reference? Is it a variant of ELO?
| HoshinoAI wrote:
| check matchmaking section: https://www.dota2.com/newfrontiers
|
| Valve listed some reason for making the change.
|
| https://en.wikipedia.org/wiki/Glicko_rating_system
___________________________________________________________________
(page generated 2023-05-03 23:02 UTC)