[HN Gopher] ASTRA: HackerRank's coding benchmark for LLMs
       ___________________________________________________________________
        
       ASTRA: HackerRank's coding benchmark for LLMs
        
       We help companies hire & upskill developers. A customer recently
       asked: What % of HackerRank problems can LLMs solve? That got us
       thinking--how should hiring evolve when AI can translate natural
       language to code?  Our belief: AI will handle much of code
       generation, so developers will be assessed more on SDLC skills with
       AI assistants.  To explore this, we're benchmarking LLMs on real-
       world software dev scenarios--starting with 65 unseen problems
       across 10 domains. Beyond correctness, we evaluated consistency--an
       often overlooked aspect of AI reliability. We're open-sourcing the
       dataset on Huggingface and expanding it to cover more domains,
       ambiguous specs, and harder challenges.  Would love the HN
       community's take on this!
        
       Author : rvivek
       Score  : 11 points
       Date   : 2025-02-11 17:37 UTC (5 hours ago)
        
 (HTM) web link (www.hackerrank.com)
 (TXT) w3m dump (www.hackerrank.com)
        
       | bobnamob wrote:
       | Seems like a very limited subset of software development to be
       | basing a benchmark on
       | 
       | Where's the kernel dev? Where's the embedded dev? Where's the
       | throwaway python script?
        
       | sosuke wrote:
       | No huggingface models or did I just miss them? Edit: they mention
       | doing open models at some point at the bottom of the page
        
       | rokhayakebe wrote:
       | [delayed]
        
       ___________________________________________________________________
       (page generated 2025-02-11 23:00 UTC)