[HN Gopher] Research-Driven Agents: When an agent reads before i...
___________________________________________________________________
Research-Driven Agents: When an agent reads before it codes
Author : hopechong
Score : 103 points
Date : 2026-04-09 16:58 UTC (6 hours ago)
(HTM) web link (blog.skypilot.co)
(TXT) w3m dump (blog.skypilot.co)
| hopechong wrote:
| Coding agents that read papers before writing code find
| optimizations that code-only agents miss.
|
| We added a literature review phase to Karpathy's autoresearch
| loop and pointed it at llama.cpp. The agent autonomously read
| arxiv papers, studied competing forks and spun up VMs to run
| parallel experiments.
| dataviz1000 wrote:
| Sorry to spam, I'm working on this also from a different angle.
| Hopefully sharing adds to the conversation.
|
| First, about the loop, Claude's (coding agent) context and
| attention is big enough to self-reflect. Agent Tuning shows a
| technique that not only demonstrates this but a way quantify it.
| [0] The difference is autoresearch's val_bpb measures what the
| agent built; Agent Tuning's p measures the agent itself.
|
| > Claude's attention doesn't distinguish between "instructions
| I'm writing" and "instructions I'm following" -- they're both
| just tokens in context.
|
| Second, doing research, finding academic research to add to
| context helps. Here is an example of an implementation that
| creates trading strategies by reading research and recreating
| them in creative new ways. [1]
|
| The biggest problem is the coding agents don't "Fail fast and
| loud". They fail deceivingly.
|
| [0] https://github.com/adam-s/agent-tuning
|
| [1] https://github.com/adam-s/alphadidactic
| mkagenius wrote:
| > The biggest problem is the coding agents don't "Fail fast and
| loud". They fail deceivingly.
|
| GPT 2 and 3 used to fail fast (and loud coz we could easily see
| it lying)
| dataviz1000 wrote:
| My next exploration will be "Coding Agents: fail slow,
| silent, and deceivingly".
|
| After one month working on using Claude to create trading
| strategies, the one thing I learned; if the strategy looks
| like it can profit, it is a lie. The trading strategy agent
| doesn't find trading strategies that work, it is really a bug
| hunting agent.
| phendrenad2 wrote:
| This is obvious, right? If you want to build a Facebook clone,
| you wouldn't tell the agent "build Facebook". You would provide
| it with a description of every page on Facebook, behaviors,
| interactions, UI, etc.
| faeyanpiraat wrote:
| Have you even read the TL;DR in the linked article??
| phendrenad2 wrote:
| You mean this part?
|
| > TL;DR: Coding agents generate better optimizations when
| they read papers and study competing projects before touching
| code
|
| What made you think I hadn't read the article, let alone that
| TL;DR? I'm really curious. Jumping to an insulting "have you
| read the article" is a big step, so it'll be really
| interesting to see where your mind went.
| KingOfCoders wrote:
| I use #PPPCDC for prompting: plan,plan,plan then verify with:
| Compare the plan to the existing Code. Reread and compare the
| plan to the Docs. Fix the areas you're not Confident about.
| hungryhobbit wrote:
| I think anyone who uses Claude knows that it works smarter when
| you have it make a plan first, and ask it to research the
| existing code as much as possible first ... so the results in
| this article doesn't surprise me at all.
|
| However, I'd be curious to hear back from others who have tried
| adding the shell script (at the end of the article) to their
| flow: does it (really) improve Claude?
| doctorpangloss wrote:
| The skypilot devs need to focus on decoupling their offering, so
| that their very valuable "find the cheapest cloud" functionality
| isn't married to a glitchy reinvention of Kubernetes JobSet and
| MLflow
| simlevesque wrote:
| I've been making skills from arxiv papers for a while. I have a
| one for multi-object tracking for example. It has a SKILL.md
| describing all important papers (over 30) on the subject and a
| folder with each paper's full content as reStructuredText.
|
| To feed Arxiv papers to LLMs I found that RST gives the best
| token count/fidelity ratio. Markdown lacks precision. LateX is
| too verbose. I have a script with the paper's urls, name and date
| that downloads the LateX zips from Arxiv, extracts it, transforms
| them to RST and then adds them to the right folder. Then I ask a
| LLM to make a summary from the full text, then I give other LLMs
| the full paper again with the summary and ask them to improve on
| and and proofread them. While this goes on I read the papers
| myself and at the end I read the summaries and if I approve them
| I add it to the skill. I also add for each paper info on how well
| the algorithms described do in common benchmarks.
|
| I highly recommend doing something similar if you're working in a
| cutting-edge domain. Also I'd like to know if anyone has
| recommendations to improve what I do.
| alex000kim wrote:
| sounds similar to "LLM Knowledge Bases"
| https://xcancel.com/karpathy/status/2039805659525644595
| MrLeap wrote:
| What is RST?
| simlevesque wrote:
| reStructuredText: https://www.sphinx-
| doc.org/en/master/usage/restructuredtext/...
| paulluuk wrote:
| This sounds like it would work, but honestly if you've already
| read all 30 papers fully, what do you still need to llm to do
| for you? Just the boilerplate?
| simlevesque wrote:
| I'm trying to make a go library that implements a wide ranges
| of MOT algorithms and can gather metrics for all of them.
|
| Reading all the papers once isn't the same as this. I find it
| very useful.
|
| I can ask an LLM to do the basic implementations, then I can
| refine them (make the code better, faster, cut on memory
| use), then I can ask the LLM if I'm still implementing the
| algorithms as they're described in the paper.
| ctoth wrote:
| I've been working on ctoth/research-papers-plugin, the pipeline
| to actually get LLMs to extract the notes. I really like your
| insight re RST over Markdown! It sounds like we're working on
| similar stuff and I'll absolutely reach out :)
| simlevesque wrote:
| I'm gonna look at your plugin. My email is in my profile.
|
| Honestly I think that Markdown with LateX code blocks would
| be the most efficient representation but when doing it with
| Pandoc I kept having issues with loss of information and
| sometimes even syntax error.
| satvikpendem wrote:
| Does that even fit in the context? It seems like 30 papers
| worth of content would just overflow it.
| ctoth wrote:
| For each paper, have your agent extract a three sentence
| description, create a description.md, then concat those with
| the paper names into an INDEX.md which it should consult to
| find appropriate papers. Also: have your agent tag papers,
| then autogenerate your tagged collection on the filesystem.
| Then you get nice things like
| https://github.com/ctoth/Qlatt/tree/master/papers/tagged
|
| Then something in your {CLAUDE,AGENTS}.md that says: when
| working on something with relevant context supplied by
| papers, read the papers before doing the work. You can find
| all papers plus their descriptions in ./papers/INDEX.md and
| papers by tag in ./papers/tagged
| austinbaggio wrote:
| Research step makes sense, can also confirm that running multiple
| agents with diverse strategies also compound results more quickly
| than single agents
| alex000kim wrote:
| I am sure this would works well in general. There is a
| challenge wrt to how to make them communicate effectively to
| e.g. 1) avoid duplicative work and 2) allow them to
| combine/overlay each others' findings to yield even better
| results
| outside1234 wrote:
| A research step (gather insights from across the codebase and
| internet for how to accomplish the next step), planning step (how
| should I sequence implementation given that research), an
| implementation step, and a verification step (code review of the
| implementation) is super effective workflow for me.
| alex000kim wrote:
| yup, as the blog says
|
| > The full setup works with any project that has a benchmark
| and test suite.
|
| so having a clear and measurable verification step is key.
| Meaning you can't simply give an AI agent a vague goal e.g.
| "improve the quality of the codebase" because it's too general.
| ctoth wrote:
| I've been very interested in this recently. I'm pretty sure that
| _every_ project should have a . /papers directory of annotated
| papers in it like I do in Qlatt[0].
|
| Literally every project. If it's something that's been done a
| million times then that means it has good literature on it? If
| not, then even more important to find related stuff! And not just
| crunchy CS stuff like databases or compilers or whatever. Are you
| creating a UI? There's probably been great UI research you can
| base off of! Will this game loop be fun in the game you're
| building? There's probably been research about it!
|
| [0]: https://github.com/ctoth/Qlatt/blob/master/papers/
| alex000kim wrote:
| That directory is huge already! I guess the index.md helps the
| agent find what it needs, but even the markdown file is very
| long - this would consume a ton of tokens.
|
| Also I wonder who/what decides what papers go in there.
|
| In the blog post, the agent is allowed to do its own search.
| ctoth wrote:
| Check out the Researcher and Process Leads skill in
| ctoth/research-papers-plugin. I have basically completely
| automated the literature review.
| pstuart wrote:
| Having a "indexed global data collection" of the markdown
| would be a kumbaya moment for AI. There's so much data out
| there but finite disk space. Maybe torrents or IPFS could
| work for this?
| ctoth wrote:
| I'm actually sort of working on this!
| https://github.com/ctoth/propstore -- it's like Cyc, but
| there is no one answer. Plus knowledge bases are literally
| git repos that you can fork/merge. Research-papers-plugin
| is the frontend, we extract the knowledge, then we need
| somewhere to put it :)
| pstuart wrote:
| Awesome! TIL about Cyc, and it's quite intriguing. I'd
| been thinking about how being able to integrate Prolog or
| similar tools might be a valuable endeavor (although I've
| yet to write anything in Prolog myself).
| zzleeper wrote:
| Wow this is amazing. Did you write all those MD files by hand,
| or used an LLM for the simple stuff like extracting abstracts?
| ctoth wrote:
| I used https://github.com/ctoth/research-papers-plugin to
| produce the annotations. The thing that's really cool is how
| they surface the cross-links in the collection, for instance
| look at https://github.com/ctoth/Qlatt/blob/master/papers/Fan
| t_1988_...
|
| Claude is much faster and better at reading papers than Codex
| (some of this is nested skill dispatch) but they both work
| quite incredibly for this. Compile your set of papers, queue
| it up and hit /ingest-collection and go sleep, and come back
| to a remarkable knowledge base :)
| maCDzP wrote:
| I have a ML project. I usually set up a team of agents, where I
| have a leader, archivist, research assistant, researcher,
| developer and tester. The team generates hypothesis based on
| papers, test it, and iterate over that. Everything is documented
| using a lab notebook. It burns tokens but I have found some
| promising strategies that I am testing.
| kaycebasques wrote:
| Gemini has a Deep Research API: https://ai.google.dev/gemini-
| api/docs/deep-research
| tomi_dev wrote:
| This is interesting.Do you see a noticeable difference in output
| quality when the agent reads context first vs going straight into
| generation?
|
| Feels like most tools skip that step.
| tomi_dev wrote:
| This is interesting.
|
| Do you see a noticeable difference in output quality when the
| agent reads context first vs going straight into generation?
|
| Feels like most tools skip that step.
| jbergqvist wrote:
| When I want to solve a new problem with an agent, I always ask it
| to search broadly for prior work in the given area online, and
| then analyze if we can build our solution using it as
| inspiration.
|
| I see it as the solution being out there in "idea space", and by
| having the agent search beforehand we can more efficiently
| explore this space before converging on the final solution.
| prats226 wrote:
| A good experiment would be to also try giving it access to
| latency traces so it can identify issues? Wrt coding agents,
| giving access to observability tools often improve
| coding/debugging ability for me
___________________________________________________________________
(page generated 2026-04-09 23:00 UTC)