[HN Gopher] Scaling Karpathy's Autoresearch: What Happens When t...
       ___________________________________________________________________
        
       Scaling Karpathy's Autoresearch: What Happens When the Agent Gets a
       GPU Cluster
        
       Author : hopechong
       Score  : 98 points
       Date   : 2026-03-19 16:55 UTC (6 hours ago)
        
 (HTM) web link (blog.skypilot.co)
 (TXT) w3m dump (blog.skypilot.co)
        
       | kraddypatties wrote:
       | I feel like most of this recent Autoresearch trend boils down to
       | reinventing hyper-parameter tuning. Is the SOTA still Bayesian
       | optimization when given a small cluster? It was ~3 years ago when
       | I was doing this kind of work, haven't kept up since then.
       | 
       | Also, shoutout SkyPilot! It's been a huge help for going multi-
       | cloud with our training and inference jobs (getting GPUs is still
       | a nightmare...)!
        
         | ipsum2 wrote:
         | Hyperparam tuning that has better intuition and can incorporate
         | architecture changes automatically. It won't invent something
         | completely new though.
        
           | kraddypatties wrote:
           | Hm, that's fair. It does feel like there's low hanging fruit
           | in combining "old school" methods for conducting a
           | hyperparameter sweep efficiently _with_ the higher level
           | architecture edit ability of Autoresearch.
           | 
           | Probably would cut the number of runs down by a significant
           | number (as far as I can tell it's doing a grid search once it
           | decides to mess with a knob or section of the architecture).
        
         | karpathy wrote:
         | Wrong and short-sighted take given that the LLM explores
         | serially learning along the way, and can tool use and change
         | code arbitrarily. It seems to currently default to something
         | resembling hyperparameter tuning in absence of more specific
         | instructions. I briefly considered calling the project
         | "autotune" at first but I think "autoresearch" will prove to be
         | the significantly more appropriate name.
        
           | corndoge wrote:
           | Would you say it's fair to describe autoresearch as a form of
           | neural architecture search? I am curious what you think the
           | core differences are between them.
        
           | kraddypatties wrote:
           | I can believe that in the long run.
           | 
           | Does the agent have access to arxiv (a brief skim of the
           | README didn't have an answer)? If not, it could be that the
           | current approach of relying on the model's weights only is
           | resulting in the perceived local optimum of hyperparameter
           | tuning.
           | 
           | Anecdotally, we built a little MCP for arxiv to help with our
           | internal research, noticed a significant boost in the
           | diversity of methods (architecture or otherwise) Claude and
           | friends were able to reference.
        
           | westurner wrote:
           | Is there a cost to converge? And how much does it vary with
           | the random seed?
           | 
           | Re: OpenCogPrime:EconomicAttentionAllocation
           | https://news.ycombinator.com/item?id=45518074 and something
           | about eWASM (edit)
           | https://news.ycombinator.com/item?id=47171887 .. from
           | https://news.ycombinator.com/item?id=46825026 _re: eWASM and
           | costed opcodes for agent efficiency_
        
           | achierius wrote:
           | Out of curiosity, what sort of things have you seen it do
           | that better fit 'autoresearch' than 'autotune' thus far?
           | Optimizations it made that wouldn't be been surfaced by an
           | autotune system, I suppose.
        
             | karpathy wrote:
             | The most recent round of autoresearch (round 2) which
             | decreased "time to GPT-2" from 1.8 hours to 1.65 hours had
             | some examples. I adjusted the program.md to "look at modded
             | nanogpt project and draw inspirations from there for things
             | to try" and it came back with a bunch of tuning, but also
             | tried and implemented new architecture changes, some of
             | which actually helped including the smear gate and the
             | backout skip connection. These are not just
             | hyperparameters, they are new PyTorch code. I'm now working
             | on a more general system that can have a queue of ideas
             | that could be sourced from archive papers, github repos,
             | etc.
        
             | jwilber wrote:
             | I see this critique about autoresearch online often, but I
             | think it's misplaced.
             | 
             | Here's a use case that may illuminate the difference, from
             | my own work at Nvidia. Im currently training some large
             | sparse autoencoders, and there are issues with dead
             | latents. Several solutions exit to help here, such as auxk,
             | which I can certainly include and tune the relevant params
             | as you describe. However, I have several other ideas that
             | are much different, each of which requires editing core
             | code (full evaluation changes, initialization strategies,
             | architecture changes, etc.), including changes to
             | parallelism strategies in the multi-rank environment I'm
             | using. Moreover, based on my ideas and other existing
             | literature, Claude can try a number of new ideas, each
             | potentially involving more code changes.
             | 
             | This automated run-and-discover process is far beyond
             | what's possible with hyperparam search.
        
               | achierius wrote:
               | It wasn't meant as a critique, I'm legitimately
               | interested in knowing more about where it can push
               | boundaries and where it struggles. I agree that in
               | general it's a truism that "Claude can try a number of
               | new ideas" etc., but the question remains as to where in
               | particular it actually takes advantage of this to push
               | the envelope in a way other tools don't -- since that
               | informs when it makes sense to use something like this.
        
           | saberience wrote:
           | Have you actually used LLMs for non trivial tasks? They are
           | still incredibly bad when it comes to actually hard
           | engineering work and they still lie all the time, it's just
           | gotten harder to notice, especially if you're just letting it
           | run all night and generate reams of crap.
           | 
           | Most people are optimizing for terrible benchmarks and then
           | don't really understand what the model did anyone and just
           | assume it did something good. It's the blind leading the
           | blind basically, and a lot of people with an AI-psychosis or
           | delusion.
        
             | nfg wrote:
             | Do you realise who you're replying to?
        
               | _menelaus wrote:
               | lolololol
        
               | emp17344 wrote:
               | Why should we care that he's famous?
        
               | nfg wrote:
               | Fame doesn't enter it - the point is Karpathy has about
               | as strong a claim as anyone to having "actually used LLMs
               | for non trivial tasks".
        
               | CamperBob2 wrote:
               | Reminds me of another famous HN footgun, where some
               | people were arguing about math. One of them backed up his
               | opinion by pointing out that he made it to the Putnam
               | competition, or something like that. The other guy said,
               | "Cool. I won it that year."
               | 
               | Of course that's not a reliable indication of who was
               | right, but still, you never know who you're dissing
               | around here. (Edit: the other poster found it. Even Paul
               | Graham was like, "Damn, son.")
        
               | ericd wrote:
               | Shades of https://news.ycombinator.com/item?id=35079
        
       | zhwu wrote:
       | The most surprising part: the agent had access to both H100s and
       | H200s. Without being told, it noticed H200s scored better and
       | started screening ideas on H100s, then promoting winners to H200s
       | for validation. That strategy emerged entirely on its own.
        
         | Aboutplants wrote:
         | Yeah I thought that was a particularly neat part
        
         | rogerrogerr wrote:
         | Why do we think this emerged "on its own"? Surely this
         | technique has been discussed in research papers that are in the
         | training set.
        
           | fdghrtbrt wrote:
           | Why surely? Have you never seen an LLM try something new?
        
             | rogerrogerr wrote:
             | Is your assertion that no one has ever written "we tried
             | some stuff on the small inexpensive platform first, then
             | moved to the bigger more expensive platform with the more
             | promising options" in a research paper or literally
             | anywhere else?
        
               | fdghrtbrt wrote:
               | No, that's not my assertion. In fact I asserted nothing
               | at all.
        
               | rogerrogerr wrote:
               | You're speaking in riddles; your communication would be
               | more effective if you didn't do that.
        
               | fdghrtbrt wrote:
               | You said "surely", and I asked:
               | 
               | > Why surely? Have you never seen an LLM try something
               | new?
               | 
               | I'm afraid I can't make it any simpler than this.
               | 
               | And I still don't know the answer to how you're so sure.
               | To me there's several explanations, and it seems to you
               | there's only one.
               | 
               | I'm pretty happy with my communication style.
        
               | frank_nitti wrote:
               | Seems to me the commenter was asking: what observations
               | led us to conclude that original affirmative statement
               | that "the AI did this entirely on its own".
               | 
               | Given that this is a common technique and not a novel
               | invention, it's probably present in the training set.
               | 
               | The "surely" reads like it's referring to the presence of
               | that information in the training set. But your response
               | casts it as saying "surely the AI has not invented
               | something on its own".
               | 
               | The original question stands IMO, the burden of proof is
               | on whoever is asserting that the AI has invented
               | something on its own, with or without training data that
               | surely already mentions this approach
        
               | fdghrtbrt wrote:
               | There is no burden of proof on me, because I'm not
               | asserting that AI has invented something on its own. I
               | haven't told you what my view is or whether I ever have a
               | view.
               | 
               | The problem with the reasoning of the person I was
               | responding to is that it's assuming "if X is in the
               | training set and LLM outputs X, then it did so because X
               | is in the training set". That does not follow.
               | Conceivably it's possible that X is in the training set
               | and LLM outputs X, but if X hadn't been in the training
               | set the LLM also would've output X.
               | 
               | Lets look at that phrase again:
               | 
               | > Why do we think this emerged "on its own"? Surely this
               | technique has been discussed in research papers that are
               | in the training set.
               | 
               | This phrase implies "if X was in the training set, then
               | LLM couldn't have come up with X on its own". This is
               | false. In fact, my claim that the implication is false is
               | testable, in the following manner: Have two training
               | sets, T and T'. In T, X is present. In T' you've removed
               | X but left X-adjacent things. Train LLM A on T and A' on
               | T'. Find a prompt that requires that A outputs X. If on
               | the same prompt A' also outputs X, that's an example of
               | my claim. To repeat, my claim is "it's possible that X is
               | in the training set and LLM outputs X, but if X hadn't
               | been in the training set the LLM also would've output X."
               | 
               | In fact, I've just realized I even have a method for
               | constructing (T, T') that guarantees what I've described.
               | Not sure if it's worth a paper on its own though.
        
             | caconym_ wrote:
             | I honestly don't think I have.
             | 
             | In this case, using a cheap(er) signal or heuristic as an
             | initial filter before spending more resources on cases that
             | pass the filter is a pattern that shows up all over the
             | place, and LLMs _are_ good at picking up on patterns like
             | that and generalizing them. AFAICT.
        
               | anon291 wrote:
               | I'm not sure how people say this so confidently. I have a
               | rather esoteric haskell library that I've written and
               | published for years. ChatGPT and Claude both know about
               | it and frequently help me improve it, and propose
               | completely novel approaches. I'm really not sure how
               | people are so confident that they can't think of anything
               | new. This seems like wishful confirmation bias.
        
         | hhh wrote:
         | Why?... The experiment.yaml shows that it is calling h100/200
         | explicitly, it's pretty common for humans to say "number bigger
         | more gooder" for anything... Lie and reverse the values and see
         | what happens. I would put money on a rabbit hole of complaining
         | about it being misconfigured.
        
           | ed wrote:
           | Models are familiar with H100's. They even predate ChatGPT.
        
       | covi wrote:
       | This feels like the chimpanzee with a power drill. An agent is
       | honestly just brute-force search, but guided.
        
         | chaos_emergent wrote:
         | Human-driven research is also brute-force but with a more
         | efficient search strategy. One can think of a parameter that
         | represents research-search-space-navigation efficiency. RL-
         | trained agents will inevitably optimize for that parameter. I
         | agree with your statement insomuch as the value of that
         | efficiency parameter is lower for agents than humans today.
         | 
         | It's really hard to imagine that they __won't__ exceed the
         | human value for that efficiency parameter rather soon given
         | that 1. there are plenty of scalar value functions that can
         | represent research efficiency, of which a subset will result in
         | robust training, and 2. that AI labs have a massive incentive
         | to increase their research efficiency overall, along with
         | billions of dollars and really good human researchers working
         | on the problem.
        
         | groby_b wrote:
         | Is there anything in the research space that doesn't fit
         | "brute-force search, but guided"?
         | 
         | All of science is "gather inputs, make hypothesis, test,
         | analyse" on repeat.
         | 
         | There's plenty to critique in the particular guidance approach,
         | but the overall method is the same.
        
         | gwern wrote:
         | Except the power drill isn't being used to make a better
         | chimpanzee.
        
       | ipsum2 wrote:
       | A cluster is 2 nodes? That's technically true, but not very
       | exciting.
        
       | fabmilo wrote:
       | I am fascinated by this example of using AI to improve AI. I won
       | a small prize using this technique on helion kernels at a pytorch
       | hackathon in SF.
       | 
       | The next step are: - give the agent the whole deep learning
       | literature research and do tree search over the various ideas
       | that have been proposed in the past. - have some distributed
       | notepad that any of these agents can read and improve upon.
        
       | saberience wrote:
       | Wait, "Karpathy's Autoresearch", you mean a loop that prompts the
       | agent to improve a thing given a benchmark?
       | 
       | People have been doing this for a year or more, Ralph loops etc.
       | 
       | I hate the weird strange Twitter world of hero-worship for folks
       | that seems to arise just out of large followings.
       | 
       | Joe no-followers does this six months ago, nobody cares. Karpathy
       | writes a really basic loop and it's now a kind of AI miracle
       | prompting tons of grifters, copy-cats, weird hype.
       | 
       | I do wonder if LLMs have just made everyone seriously, seriously
       | dumber all of a sudden. Most of the "Autoresearch" posts I see
       | are completely rubbish, with AI optimizing for nonsense
       | benchmarks and people failing to understand the graphs they are
       | looking at. So yes, the AI made itself better at a useless
       | benchmark while also making the code worse in 10 other ways you
       | don't actually understand.
        
         | password54321 wrote:
         | The number of refurbished mac minis that are available in my
         | country has suddenly dramatically increased ever since the
         | Clawdbot tweet. People never learn.
        
       | pbkhrv wrote:
       | > How parallelism changed the agent's research strategy > With a
       | single GPU, the agent is stuck doing greedy hill-climbing: try
       | one thing, check the result, pick a direction, try the next
       | thing. With 16 GPUs, the strategy shifts. ...skip... 12
       | experiments in a single 5-minute wave. This makes it much harder
       | to get stuck in local optima and much easier to find interaction
       | effects between parameters.
       | 
       | The agent can theoretically come up with a protocol to run those
       | same 12 experiments one-by-one and only then decide which branch
       | to explore next - which I think would lead to the same outcome?
       | 
       | But in this case, it just happened to have stumbled on this
       | particular outcome only because it didn't get a chance to execute
       | a greedy strategy after the first 1 or 2 results.
       | 
       | Worse experiment design + parallelism = better experiment design
       | + serialized execution ?
        
       | herf wrote:
       | This "early velocity only" approach seems like a problem - how do
       | you know with 5-minute training runs that you aren't affecting
       | the overall asymptote? e.g., what if the AI picks a quantizer
       | that happens to be faster in the first five minutes, but has a
       | big noise floor where it can't make more progress?
        
       ___________________________________________________________________
       (page generated 2026-03-19 23:00 UTC)