[HN Gopher] Scaling Karpathy's Autoresearch: What Happens When t...
___________________________________________________________________
Scaling Karpathy's Autoresearch: What Happens When the Agent Gets a
GPU Cluster
Author : hopechong
Score : 98 points
Date : 2026-03-19 16:55 UTC (6 hours ago)
(HTM) web link (blog.skypilot.co)
(TXT) w3m dump (blog.skypilot.co)
| kraddypatties wrote:
| I feel like most of this recent Autoresearch trend boils down to
| reinventing hyper-parameter tuning. Is the SOTA still Bayesian
| optimization when given a small cluster? It was ~3 years ago when
| I was doing this kind of work, haven't kept up since then.
|
| Also, shoutout SkyPilot! It's been a huge help for going multi-
| cloud with our training and inference jobs (getting GPUs is still
| a nightmare...)!
| ipsum2 wrote:
| Hyperparam tuning that has better intuition and can incorporate
| architecture changes automatically. It won't invent something
| completely new though.
| kraddypatties wrote:
| Hm, that's fair. It does feel like there's low hanging fruit
| in combining "old school" methods for conducting a
| hyperparameter sweep efficiently _with_ the higher level
| architecture edit ability of Autoresearch.
|
| Probably would cut the number of runs down by a significant
| number (as far as I can tell it's doing a grid search once it
| decides to mess with a knob or section of the architecture).
| karpathy wrote:
| Wrong and short-sighted take given that the LLM explores
| serially learning along the way, and can tool use and change
| code arbitrarily. It seems to currently default to something
| resembling hyperparameter tuning in absence of more specific
| instructions. I briefly considered calling the project
| "autotune" at first but I think "autoresearch" will prove to be
| the significantly more appropriate name.
| corndoge wrote:
| Would you say it's fair to describe autoresearch as a form of
| neural architecture search? I am curious what you think the
| core differences are between them.
| kraddypatties wrote:
| I can believe that in the long run.
|
| Does the agent have access to arxiv (a brief skim of the
| README didn't have an answer)? If not, it could be that the
| current approach of relying on the model's weights only is
| resulting in the perceived local optimum of hyperparameter
| tuning.
|
| Anecdotally, we built a little MCP for arxiv to help with our
| internal research, noticed a significant boost in the
| diversity of methods (architecture or otherwise) Claude and
| friends were able to reference.
| westurner wrote:
| Is there a cost to converge? And how much does it vary with
| the random seed?
|
| Re: OpenCogPrime:EconomicAttentionAllocation
| https://news.ycombinator.com/item?id=45518074 and something
| about eWASM (edit)
| https://news.ycombinator.com/item?id=47171887 .. from
| https://news.ycombinator.com/item?id=46825026 _re: eWASM and
| costed opcodes for agent efficiency_
| achierius wrote:
| Out of curiosity, what sort of things have you seen it do
| that better fit 'autoresearch' than 'autotune' thus far?
| Optimizations it made that wouldn't be been surfaced by an
| autotune system, I suppose.
| karpathy wrote:
| The most recent round of autoresearch (round 2) which
| decreased "time to GPT-2" from 1.8 hours to 1.65 hours had
| some examples. I adjusted the program.md to "look at modded
| nanogpt project and draw inspirations from there for things
| to try" and it came back with a bunch of tuning, but also
| tried and implemented new architecture changes, some of
| which actually helped including the smear gate and the
| backout skip connection. These are not just
| hyperparameters, they are new PyTorch code. I'm now working
| on a more general system that can have a queue of ideas
| that could be sourced from archive papers, github repos,
| etc.
| jwilber wrote:
| I see this critique about autoresearch online often, but I
| think it's misplaced.
|
| Here's a use case that may illuminate the difference, from
| my own work at Nvidia. Im currently training some large
| sparse autoencoders, and there are issues with dead
| latents. Several solutions exit to help here, such as auxk,
| which I can certainly include and tune the relevant params
| as you describe. However, I have several other ideas that
| are much different, each of which requires editing core
| code (full evaluation changes, initialization strategies,
| architecture changes, etc.), including changes to
| parallelism strategies in the multi-rank environment I'm
| using. Moreover, based on my ideas and other existing
| literature, Claude can try a number of new ideas, each
| potentially involving more code changes.
|
| This automated run-and-discover process is far beyond
| what's possible with hyperparam search.
| achierius wrote:
| It wasn't meant as a critique, I'm legitimately
| interested in knowing more about where it can push
| boundaries and where it struggles. I agree that in
| general it's a truism that "Claude can try a number of
| new ideas" etc., but the question remains as to where in
| particular it actually takes advantage of this to push
| the envelope in a way other tools don't -- since that
| informs when it makes sense to use something like this.
| saberience wrote:
| Have you actually used LLMs for non trivial tasks? They are
| still incredibly bad when it comes to actually hard
| engineering work and they still lie all the time, it's just
| gotten harder to notice, especially if you're just letting it
| run all night and generate reams of crap.
|
| Most people are optimizing for terrible benchmarks and then
| don't really understand what the model did anyone and just
| assume it did something good. It's the blind leading the
| blind basically, and a lot of people with an AI-psychosis or
| delusion.
| nfg wrote:
| Do you realise who you're replying to?
| _menelaus wrote:
| lolololol
| emp17344 wrote:
| Why should we care that he's famous?
| nfg wrote:
| Fame doesn't enter it - the point is Karpathy has about
| as strong a claim as anyone to having "actually used LLMs
| for non trivial tasks".
| CamperBob2 wrote:
| Reminds me of another famous HN footgun, where some
| people were arguing about math. One of them backed up his
| opinion by pointing out that he made it to the Putnam
| competition, or something like that. The other guy said,
| "Cool. I won it that year."
|
| Of course that's not a reliable indication of who was
| right, but still, you never know who you're dissing
| around here. (Edit: the other poster found it. Even Paul
| Graham was like, "Damn, son.")
| ericd wrote:
| Shades of https://news.ycombinator.com/item?id=35079
| zhwu wrote:
| The most surprising part: the agent had access to both H100s and
| H200s. Without being told, it noticed H200s scored better and
| started screening ideas on H100s, then promoting winners to H200s
| for validation. That strategy emerged entirely on its own.
| Aboutplants wrote:
| Yeah I thought that was a particularly neat part
| rogerrogerr wrote:
| Why do we think this emerged "on its own"? Surely this
| technique has been discussed in research papers that are in the
| training set.
| fdghrtbrt wrote:
| Why surely? Have you never seen an LLM try something new?
| rogerrogerr wrote:
| Is your assertion that no one has ever written "we tried
| some stuff on the small inexpensive platform first, then
| moved to the bigger more expensive platform with the more
| promising options" in a research paper or literally
| anywhere else?
| fdghrtbrt wrote:
| No, that's not my assertion. In fact I asserted nothing
| at all.
| rogerrogerr wrote:
| You're speaking in riddles; your communication would be
| more effective if you didn't do that.
| fdghrtbrt wrote:
| You said "surely", and I asked:
|
| > Why surely? Have you never seen an LLM try something
| new?
|
| I'm afraid I can't make it any simpler than this.
|
| And I still don't know the answer to how you're so sure.
| To me there's several explanations, and it seems to you
| there's only one.
|
| I'm pretty happy with my communication style.
| frank_nitti wrote:
| Seems to me the commenter was asking: what observations
| led us to conclude that original affirmative statement
| that "the AI did this entirely on its own".
|
| Given that this is a common technique and not a novel
| invention, it's probably present in the training set.
|
| The "surely" reads like it's referring to the presence of
| that information in the training set. But your response
| casts it as saying "surely the AI has not invented
| something on its own".
|
| The original question stands IMO, the burden of proof is
| on whoever is asserting that the AI has invented
| something on its own, with or without training data that
| surely already mentions this approach
| fdghrtbrt wrote:
| There is no burden of proof on me, because I'm not
| asserting that AI has invented something on its own. I
| haven't told you what my view is or whether I ever have a
| view.
|
| The problem with the reasoning of the person I was
| responding to is that it's assuming "if X is in the
| training set and LLM outputs X, then it did so because X
| is in the training set". That does not follow.
| Conceivably it's possible that X is in the training set
| and LLM outputs X, but if X hadn't been in the training
| set the LLM also would've output X.
|
| Lets look at that phrase again:
|
| > Why do we think this emerged "on its own"? Surely this
| technique has been discussed in research papers that are
| in the training set.
|
| This phrase implies "if X was in the training set, then
| LLM couldn't have come up with X on its own". This is
| false. In fact, my claim that the implication is false is
| testable, in the following manner: Have two training
| sets, T and T'. In T, X is present. In T' you've removed
| X but left X-adjacent things. Train LLM A on T and A' on
| T'. Find a prompt that requires that A outputs X. If on
| the same prompt A' also outputs X, that's an example of
| my claim. To repeat, my claim is "it's possible that X is
| in the training set and LLM outputs X, but if X hadn't
| been in the training set the LLM also would've output X."
|
| In fact, I've just realized I even have a method for
| constructing (T, T') that guarantees what I've described.
| Not sure if it's worth a paper on its own though.
| caconym_ wrote:
| I honestly don't think I have.
|
| In this case, using a cheap(er) signal or heuristic as an
| initial filter before spending more resources on cases that
| pass the filter is a pattern that shows up all over the
| place, and LLMs _are_ good at picking up on patterns like
| that and generalizing them. AFAICT.
| anon291 wrote:
| I'm not sure how people say this so confidently. I have a
| rather esoteric haskell library that I've written and
| published for years. ChatGPT and Claude both know about
| it and frequently help me improve it, and propose
| completely novel approaches. I'm really not sure how
| people are so confident that they can't think of anything
| new. This seems like wishful confirmation bias.
| hhh wrote:
| Why?... The experiment.yaml shows that it is calling h100/200
| explicitly, it's pretty common for humans to say "number bigger
| more gooder" for anything... Lie and reverse the values and see
| what happens. I would put money on a rabbit hole of complaining
| about it being misconfigured.
| ed wrote:
| Models are familiar with H100's. They even predate ChatGPT.
| covi wrote:
| This feels like the chimpanzee with a power drill. An agent is
| honestly just brute-force search, but guided.
| chaos_emergent wrote:
| Human-driven research is also brute-force but with a more
| efficient search strategy. One can think of a parameter that
| represents research-search-space-navigation efficiency. RL-
| trained agents will inevitably optimize for that parameter. I
| agree with your statement insomuch as the value of that
| efficiency parameter is lower for agents than humans today.
|
| It's really hard to imagine that they __won't__ exceed the
| human value for that efficiency parameter rather soon given
| that 1. there are plenty of scalar value functions that can
| represent research efficiency, of which a subset will result in
| robust training, and 2. that AI labs have a massive incentive
| to increase their research efficiency overall, along with
| billions of dollars and really good human researchers working
| on the problem.
| groby_b wrote:
| Is there anything in the research space that doesn't fit
| "brute-force search, but guided"?
|
| All of science is "gather inputs, make hypothesis, test,
| analyse" on repeat.
|
| There's plenty to critique in the particular guidance approach,
| but the overall method is the same.
| gwern wrote:
| Except the power drill isn't being used to make a better
| chimpanzee.
| ipsum2 wrote:
| A cluster is 2 nodes? That's technically true, but not very
| exciting.
| fabmilo wrote:
| I am fascinated by this example of using AI to improve AI. I won
| a small prize using this technique on helion kernels at a pytorch
| hackathon in SF.
|
| The next step are: - give the agent the whole deep learning
| literature research and do tree search over the various ideas
| that have been proposed in the past. - have some distributed
| notepad that any of these agents can read and improve upon.
| saberience wrote:
| Wait, "Karpathy's Autoresearch", you mean a loop that prompts the
| agent to improve a thing given a benchmark?
|
| People have been doing this for a year or more, Ralph loops etc.
|
| I hate the weird strange Twitter world of hero-worship for folks
| that seems to arise just out of large followings.
|
| Joe no-followers does this six months ago, nobody cares. Karpathy
| writes a really basic loop and it's now a kind of AI miracle
| prompting tons of grifters, copy-cats, weird hype.
|
| I do wonder if LLMs have just made everyone seriously, seriously
| dumber all of a sudden. Most of the "Autoresearch" posts I see
| are completely rubbish, with AI optimizing for nonsense
| benchmarks and people failing to understand the graphs they are
| looking at. So yes, the AI made itself better at a useless
| benchmark while also making the code worse in 10 other ways you
| don't actually understand.
| password54321 wrote:
| The number of refurbished mac minis that are available in my
| country has suddenly dramatically increased ever since the
| Clawdbot tweet. People never learn.
| pbkhrv wrote:
| > How parallelism changed the agent's research strategy > With a
| single GPU, the agent is stuck doing greedy hill-climbing: try
| one thing, check the result, pick a direction, try the next
| thing. With 16 GPUs, the strategy shifts. ...skip... 12
| experiments in a single 5-minute wave. This makes it much harder
| to get stuck in local optima and much easier to find interaction
| effects between parameters.
|
| The agent can theoretically come up with a protocol to run those
| same 12 experiments one-by-one and only then decide which branch
| to explore next - which I think would lead to the same outcome?
|
| But in this case, it just happened to have stumbled on this
| particular outcome only because it didn't get a chance to execute
| a greedy strategy after the first 1 or 2 results.
|
| Worse experiment design + parallelism = better experiment design
| + serialized execution ?
| herf wrote:
| This "early velocity only" approach seems like a problem - how do
| you know with 5-minute training runs that you aren't affecting
| the overall asymptote? e.g., what if the AI picks a quantizer
| that happens to be faster in the first five minutes, but has a
| big noise floor where it can't make more progress?
___________________________________________________________________
(page generated 2026-03-19 23:00 UTC)