[HN Gopher] Research-Driven Agents: When an agent reads before i...
       ___________________________________________________________________
        
       Research-Driven Agents: When an agent reads before it codes
        
       Author : hopechong
       Score  : 103 points
       Date   : 2026-04-09 16:58 UTC (6 hours ago)
        
 (HTM) web link (blog.skypilot.co)
 (TXT) w3m dump (blog.skypilot.co)
        
       | hopechong wrote:
       | Coding agents that read papers before writing code find
       | optimizations that code-only agents miss.
       | 
       | We added a literature review phase to Karpathy's autoresearch
       | loop and pointed it at llama.cpp. The agent autonomously read
       | arxiv papers, studied competing forks and spun up VMs to run
       | parallel experiments.
        
       | dataviz1000 wrote:
       | Sorry to spam, I'm working on this also from a different angle.
       | Hopefully sharing adds to the conversation.
       | 
       | First, about the loop, Claude's (coding agent) context and
       | attention is big enough to self-reflect. Agent Tuning shows a
       | technique that not only demonstrates this but a way quantify it.
       | [0] The difference is autoresearch's val_bpb measures what the
       | agent built; Agent Tuning's p measures the agent itself.
       | 
       | > Claude's attention doesn't distinguish between "instructions
       | I'm writing" and "instructions I'm following" -- they're both
       | just tokens in context.
       | 
       | Second, doing research, finding academic research to add to
       | context helps. Here is an example of an implementation that
       | creates trading strategies by reading research and recreating
       | them in creative new ways. [1]
       | 
       | The biggest problem is the coding agents don't "Fail fast and
       | loud". They fail deceivingly.
       | 
       | [0] https://github.com/adam-s/agent-tuning
       | 
       | [1] https://github.com/adam-s/alphadidactic
        
         | mkagenius wrote:
         | > The biggest problem is the coding agents don't "Fail fast and
         | loud". They fail deceivingly.
         | 
         | GPT 2 and 3 used to fail fast (and loud coz we could easily see
         | it lying)
        
           | dataviz1000 wrote:
           | My next exploration will be "Coding Agents: fail slow,
           | silent, and deceivingly".
           | 
           | After one month working on using Claude to create trading
           | strategies, the one thing I learned; if the strategy looks
           | like it can profit, it is a lie. The trading strategy agent
           | doesn't find trading strategies that work, it is really a bug
           | hunting agent.
        
       | phendrenad2 wrote:
       | This is obvious, right? If you want to build a Facebook clone,
       | you wouldn't tell the agent "build Facebook". You would provide
       | it with a description of every page on Facebook, behaviors,
       | interactions, UI, etc.
        
         | faeyanpiraat wrote:
         | Have you even read the TL;DR in the linked article??
        
           | phendrenad2 wrote:
           | You mean this part?
           | 
           | > TL;DR: Coding agents generate better optimizations when
           | they read papers and study competing projects before touching
           | code
           | 
           | What made you think I hadn't read the article, let alone that
           | TL;DR? I'm really curious. Jumping to an insulting "have you
           | read the article" is a big step, so it'll be really
           | interesting to see where your mind went.
        
       | KingOfCoders wrote:
       | I use #PPPCDC for prompting: plan,plan,plan then verify with:
       | Compare the plan to the existing Code. Reread and compare the
       | plan to the Docs. Fix the areas you're not Confident about.
        
       | hungryhobbit wrote:
       | I think anyone who uses Claude knows that it works smarter when
       | you have it make a plan first, and ask it to research the
       | existing code as much as possible first ... so the results in
       | this article doesn't surprise me at all.
       | 
       | However, I'd be curious to hear back from others who have tried
       | adding the shell script (at the end of the article) to their
       | flow: does it (really) improve Claude?
        
       | doctorpangloss wrote:
       | The skypilot devs need to focus on decoupling their offering, so
       | that their very valuable "find the cheapest cloud" functionality
       | isn't married to a glitchy reinvention of Kubernetes JobSet and
       | MLflow
        
       | simlevesque wrote:
       | I've been making skills from arxiv papers for a while. I have a
       | one for multi-object tracking for example. It has a SKILL.md
       | describing all important papers (over 30) on the subject and a
       | folder with each paper's full content as reStructuredText.
       | 
       | To feed Arxiv papers to LLMs I found that RST gives the best
       | token count/fidelity ratio. Markdown lacks precision. LateX is
       | too verbose. I have a script with the paper's urls, name and date
       | that downloads the LateX zips from Arxiv, extracts it, transforms
       | them to RST and then adds them to the right folder. Then I ask a
       | LLM to make a summary from the full text, then I give other LLMs
       | the full paper again with the summary and ask them to improve on
       | and and proofread them. While this goes on I read the papers
       | myself and at the end I read the summaries and if I approve them
       | I add it to the skill. I also add for each paper info on how well
       | the algorithms described do in common benchmarks.
       | 
       | I highly recommend doing something similar if you're working in a
       | cutting-edge domain. Also I'd like to know if anyone has
       | recommendations to improve what I do.
        
         | alex000kim wrote:
         | sounds similar to "LLM Knowledge Bases"
         | https://xcancel.com/karpathy/status/2039805659525644595
        
         | MrLeap wrote:
         | What is RST?
        
           | simlevesque wrote:
           | reStructuredText: https://www.sphinx-
           | doc.org/en/master/usage/restructuredtext/...
        
         | paulluuk wrote:
         | This sounds like it would work, but honestly if you've already
         | read all 30 papers fully, what do you still need to llm to do
         | for you? Just the boilerplate?
        
           | simlevesque wrote:
           | I'm trying to make a go library that implements a wide ranges
           | of MOT algorithms and can gather metrics for all of them.
           | 
           | Reading all the papers once isn't the same as this. I find it
           | very useful.
           | 
           | I can ask an LLM to do the basic implementations, then I can
           | refine them (make the code better, faster, cut on memory
           | use), then I can ask the LLM if I'm still implementing the
           | algorithms as they're described in the paper.
        
         | ctoth wrote:
         | I've been working on ctoth/research-papers-plugin, the pipeline
         | to actually get LLMs to extract the notes. I really like your
         | insight re RST over Markdown! It sounds like we're working on
         | similar stuff and I'll absolutely reach out :)
        
           | simlevesque wrote:
           | I'm gonna look at your plugin. My email is in my profile.
           | 
           | Honestly I think that Markdown with LateX code blocks would
           | be the most efficient representation but when doing it with
           | Pandoc I kept having issues with loss of information and
           | sometimes even syntax error.
        
         | satvikpendem wrote:
         | Does that even fit in the context? It seems like 30 papers
         | worth of content would just overflow it.
        
           | ctoth wrote:
           | For each paper, have your agent extract a three sentence
           | description, create a description.md, then concat those with
           | the paper names into an INDEX.md which it should consult to
           | find appropriate papers. Also: have your agent tag papers,
           | then autogenerate your tagged collection on the filesystem.
           | Then you get nice things like
           | https://github.com/ctoth/Qlatt/tree/master/papers/tagged
           | 
           | Then something in your {CLAUDE,AGENTS}.md that says: when
           | working on something with relevant context supplied by
           | papers, read the papers before doing the work. You can find
           | all papers plus their descriptions in ./papers/INDEX.md and
           | papers by tag in ./papers/tagged
        
       | austinbaggio wrote:
       | Research step makes sense, can also confirm that running multiple
       | agents with diverse strategies also compound results more quickly
       | than single agents
        
         | alex000kim wrote:
         | I am sure this would works well in general. There is a
         | challenge wrt to how to make them communicate effectively to
         | e.g. 1) avoid duplicative work and 2) allow them to
         | combine/overlay each others' findings to yield even better
         | results
        
       | outside1234 wrote:
       | A research step (gather insights from across the codebase and
       | internet for how to accomplish the next step), planning step (how
       | should I sequence implementation given that research), an
       | implementation step, and a verification step (code review of the
       | implementation) is super effective workflow for me.
        
         | alex000kim wrote:
         | yup, as the blog says
         | 
         | > The full setup works with any project that has a benchmark
         | and test suite.
         | 
         | so having a clear and measurable verification step is key.
         | Meaning you can't simply give an AI agent a vague goal e.g.
         | "improve the quality of the codebase" because it's too general.
        
       | ctoth wrote:
       | I've been very interested in this recently. I'm pretty sure that
       | _every_ project should have a . /papers directory of annotated
       | papers in it like I do in Qlatt[0].
       | 
       | Literally every project. If it's something that's been done a
       | million times then that means it has good literature on it? If
       | not, then even more important to find related stuff! And not just
       | crunchy CS stuff like databases or compilers or whatever. Are you
       | creating a UI? There's probably been great UI research you can
       | base off of! Will this game loop be fun in the game you're
       | building? There's probably been research about it!
       | 
       | [0]: https://github.com/ctoth/Qlatt/blob/master/papers/
        
         | alex000kim wrote:
         | That directory is huge already! I guess the index.md helps the
         | agent find what it needs, but even the markdown file is very
         | long - this would consume a ton of tokens.
         | 
         | Also I wonder who/what decides what papers go in there.
         | 
         | In the blog post, the agent is allowed to do its own search.
        
           | ctoth wrote:
           | Check out the Researcher and Process Leads skill in
           | ctoth/research-papers-plugin. I have basically completely
           | automated the literature review.
        
           | pstuart wrote:
           | Having a "indexed global data collection" of the markdown
           | would be a kumbaya moment for AI. There's so much data out
           | there but finite disk space. Maybe torrents or IPFS could
           | work for this?
        
             | ctoth wrote:
             | I'm actually sort of working on this!
             | https://github.com/ctoth/propstore -- it's like Cyc, but
             | there is no one answer. Plus knowledge bases are literally
             | git repos that you can fork/merge. Research-papers-plugin
             | is the frontend, we extract the knowledge, then we need
             | somewhere to put it :)
        
               | pstuart wrote:
               | Awesome! TIL about Cyc, and it's quite intriguing. I'd
               | been thinking about how being able to integrate Prolog or
               | similar tools might be a valuable endeavor (although I've
               | yet to write anything in Prolog myself).
        
         | zzleeper wrote:
         | Wow this is amazing. Did you write all those MD files by hand,
         | or used an LLM for the simple stuff like extracting abstracts?
        
           | ctoth wrote:
           | I used https://github.com/ctoth/research-papers-plugin to
           | produce the annotations. The thing that's really cool is how
           | they surface the cross-links in the collection, for instance
           | look at https://github.com/ctoth/Qlatt/blob/master/papers/Fan
           | t_1988_...
           | 
           | Claude is much faster and better at reading papers than Codex
           | (some of this is nested skill dispatch) but they both work
           | quite incredibly for this. Compile your set of papers, queue
           | it up and hit /ingest-collection and go sleep, and come back
           | to a remarkable knowledge base :)
        
       | maCDzP wrote:
       | I have a ML project. I usually set up a team of agents, where I
       | have a leader, archivist, research assistant, researcher,
       | developer and tester. The team generates hypothesis based on
       | papers, test it, and iterate over that. Everything is documented
       | using a lab notebook. It burns tokens but I have found some
       | promising strategies that I am testing.
        
       | kaycebasques wrote:
       | Gemini has a Deep Research API: https://ai.google.dev/gemini-
       | api/docs/deep-research
        
       | tomi_dev wrote:
       | This is interesting.Do you see a noticeable difference in output
       | quality when the agent reads context first vs going straight into
       | generation?
       | 
       | Feels like most tools skip that step.
        
       | tomi_dev wrote:
       | This is interesting.
       | 
       | Do you see a noticeable difference in output quality when the
       | agent reads context first vs going straight into generation?
       | 
       | Feels like most tools skip that step.
        
       | jbergqvist wrote:
       | When I want to solve a new problem with an agent, I always ask it
       | to search broadly for prior work in the given area online, and
       | then analyze if we can build our solution using it as
       | inspiration.
       | 
       | I see it as the solution being out there in "idea space", and by
       | having the agent search beforehand we can more efficiently
       | explore this space before converging on the final solution.
        
       | prats226 wrote:
       | A good experiment would be to also try giving it access to
       | latency traces so it can identify issues? Wrt coding agents,
       | giving access to observability tools often improve
       | coding/debugging ability for me
        
       ___________________________________________________________________
       (page generated 2026-04-09 23:00 UTC)