[HN Gopher] Chatbot Arena: Benchmarking LLMs in the Wild with El...
       ___________________________________________________________________
        
       Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
        
       Author : MMMercy2
       Score  : 17 points
       Date   : 2023-05-03 19:26 UTC (3 hours ago)
        
 (HTM) web link (lmsys.org)
 (TXT) w3m dump (lmsys.org)
        
       | weichiang wrote:
       | Surprised to learn StableLM is worse than plain LLaMA. link to
       | their leaderboard: leaderboard.lmsys.org
        
         | circuit10 wrote:
         | I've heard that it's really bad for it's size
        
       | lee101 wrote:
       | [dead]
        
       | HoshinoAI wrote:
       | It's glad to see the old technique is used for new models.
       | 
       | You may also learn from 7.33 dota update which uses a new ranking
       | algorithm called Glicko.
        
         | zhisbug wrote:
         | could you provide any reference? Is it a variant of ELO?
        
           | HoshinoAI wrote:
           | check matchmaking section: https://www.dota2.com/newfrontiers
           | 
           | Valve listed some reason for making the change.
           | 
           | https://en.wikipedia.org/wiki/Glicko_rating_system
        
       ___________________________________________________________________
       (page generated 2023-05-03 23:02 UTC)