[HN Gopher] ASTRA: HackerRank's coding benchmark for LLMs
___________________________________________________________________
ASTRA: HackerRank's coding benchmark for LLMs
We help companies hire & upskill developers. A customer recently
asked: What % of HackerRank problems can LLMs solve? That got us
thinking--how should hiring evolve when AI can translate natural
language to code? Our belief: AI will handle much of code
generation, so developers will be assessed more on SDLC skills with
AI assistants. To explore this, we're benchmarking LLMs on real-
world software dev scenarios--starting with 65 unseen problems
across 10 domains. Beyond correctness, we evaluated consistency--an
often overlooked aspect of AI reliability. We're open-sourcing the
dataset on Huggingface and expanding it to cover more domains,
ambiguous specs, and harder challenges. Would love the HN
community's take on this!
Author : rvivek
Score : 11 points
Date : 2025-02-11 17:37 UTC (5 hours ago)
(HTM) web link (www.hackerrank.com)
(TXT) w3m dump (www.hackerrank.com)
| bobnamob wrote:
| Seems like a very limited subset of software development to be
| basing a benchmark on
|
| Where's the kernel dev? Where's the embedded dev? Where's the
| throwaway python script?
| sosuke wrote:
| No huggingface models or did I just miss them? Edit: they mention
| doing open models at some point at the bottom of the page
| rokhayakebe wrote:
| [delayed]
___________________________________________________________________
(page generated 2025-02-11 23:00 UTC)