[HN Gopher] Show HN: Showboat and Rodney, so agents can demo wha...
       ___________________________________________________________________
        
       Show HN: Showboat and Rodney, so agents can demo what they've built
        
       Author : simonw
       Score  : 80 points
       Date   : 2026-02-10 17:52 UTC (5 hours ago)
        
 (HTM) web link (simonwillison.net)
 (TXT) w3m dump (simonwillison.net)
        
       | eliben wrote:
       | Very interesting! I encountered the problems these tools are
       | trying to tackle just recently while trying to guide an agent
       | into creating an in-browser tool for me. Closing the loop on a
       | web interface isn't as simple as CLI-only tools. I should give
       | this a try.
       | 
       | It's also interesting that you've shifted to Go for your agent-
       | coded CLI tools, Simon.
        
         | simonw wrote:
         | I'm dabbling with Go at the moment for small tools, mainly as
         | an excuse to learn a new language but also because having a
         | single standalone binary is convenient for shuttling these tiny
         | little tools around.
         | 
         | ... but then I'm mostly running them with "uvx name-of-tool"
         | because it turns out Python's packaging infrastructure for
         | binary tools is so good!
        
           | markusw wrote:
           | If you're coming from the Python world, definitely. I find
           | `go install github.com/simonw/rodney@latest` equally easy. :D
           | Although you need the Go tooling installed, of course. But so
           | much agree, Go is great for CLIs!
        
           | eliben wrote:
           | Right, standalone binaries for CLI tools is great. And if one
           | has Go installed, they can just `go run ...` any tool from
           | its GitHub path, all installation/build/caching happens
           | automagically (meaning the execution is immediate after the
           | first run).
           | 
           | But I can definitely see how someone with `uv` muscle memory
           | wants everything in the same command.
           | 
           | `uv` is the best thing that happened to the Python ecosystem
           | since... I don't know... maybe Numpy.
        
       | saberience wrote:
       | Sounds like both of these tools could be one shot by either
       | Claude or Codex.
       | 
       | Or alternatively, just be a skill versus a tool.
       | 
       | My "agents" already demo stuff all the time by just being
       | prompted to do so. I have notations in my standard Agents.md for
       | how I want my documentation, testing etc.
        
         | simonw wrote:
         | They kind of _were_ one-shotted by Claude. The value is in
         | coming up with a consistent design and good enough --help that
         | you can prompt:                 Run uvx showboat --help and
         | uvx rodney --help and use those       tools to demo the feature
         | you built
         | 
         | The help text effectively doubles as a skill.
        
           | markusw wrote:
           | I guess it would still make sense to have "demo" and
           | "browser-use" skills, so that the agent can reach for them
           | proactively? I always try to remove as much friction as
           | possible for myself, one little bit at a time.
        
             | simonw wrote:
             | My problem is that I work in dozens of different repos
             | generally using Claude Code for web, which doesn't have a
             | way to install extra global skills yet.
             | 
             | I don't want to duplicate my skills into all those repos
             | (and keep them updated) so I prefer the "uvx tool --help"
             | pattern.
        
               | markusw wrote:
               | That's actually one of the things that has kept me from
               | using Claude Code web (that, and I often need a Chrome
               | browser for the agent). But they must be working on it.
               | 
               | I saw an MCP I've set up on claude.ai show up in my local
               | Claude Code MCP list the other day, it seems inevitable
               | that there will be skills integration across environments
               | as well at some point.
        
               | simonw wrote:
               | In working on Rodney I found out that the Claude Code for
               | web environment has a Chrome browser installed already.
               | It's a shame you can't see its output directly - even if
               | it takes a screenshot there's no easy way to view it
               | other than having it commit and push that to a branch in
               | GitHub.
        
       | tardismechanic wrote:
       | See also (the confusingly named) playwright-cli
       | 
       | https://github.com/microsoft/playwright-cli
       | 
       | Different from the cli used for running tests etc that comes
       | bundled with PlayWright
       | 
       | Sample use:                 playwright-cli open
       | https://demo.playwright.dev/todomvc/ --headed       playwright-
       | cli type "Buy groceries"       playwright-cli press Enter
       | playwright-cli type "Water flowers"       playwright-cli press
       | Enter       playwright-cli check e21       playwright-cli check
       | e35       playwright-cli screenshot
        
         | simonw wrote:
         | Yeah that's an excellent option for this kind of thing too.
        
         | markusw wrote:
         | Oh, I hadn't seen that one either, thanks for sharing. Here I
         | am still using the Chrome Devtools MCP like a caveman. :D
        
       | toastal wrote:
       | If agents can generate text so easily, why would they be limited
       | to Markdown instead of reStructuredText, AsciiDoc, or LaTeX which
       | have rich features that help users understand text? I can
       | understand developers refusing to adopt proper formats for
       | documentation, but this seems odd for the bots. It doesn't even
       | generate the correct syntax block in Markdown using "bash"
       | instead of "sh-session".
        
         | giancarlostoro wrote:
         | I think its primarily because that is the most common
         | formatting in every editor now? I could be wrong. Markdown has
         | become the standard for README files for over a decade now.
        
           | toastal wrote:
           | Winning a popularity contest doesn't mean it's good. That is
           | the worst part of about these things as they just generate
           | the most common denominator type code/tooling while also
           | repeating anti-patterns/mistakes like the bash vs. sh-
           | session/console issue I pointed out. Garbage in has been so
           | much garbage out unfortunately.
        
             | giancarlostoro wrote:
             | Never said it was good, just making an observation that
             | Markdown is most likely to be available to render OOTB in
             | more editors. I don't think Markdown is bad necessarily
             | either. It's "good enough" for simple document.
        
         | simonw wrote:
         | Markdown has the widest tool compatibility - GitHub renders it,
         | so does VS Code and many other editors and file hosts.
         | 
         | I didn't know about sh-session, is that documented anywhere?
        
         | bee_rider wrote:
         | I dunno. I've written a bit of LaTeX but does it really shine
         | in this context? IMO the real advantage it has is that it can
         | allow the user to express more complicated intents than
         | Markdown (weird phrasing--my natural instinct was to call LaTeX
         | more precise than Markdown, but Markdown is pretty precise for
         | describing the type of file that it is good at...).
         | 
         | Anyway LLMs don't have underlying intent so maybe it is fine to
         | just let them express what they can in Markdown?
        
       | giancarlostoro wrote:
       | I'll be sure to try these out. I've been building my own
       | alternative to Beads with a concept called "gates" which do not
       | let you close tasks as complete until a gate passes. Would love
       | to throw these in as "gates" for my current workflow.
        
       | Hansenq wrote:
       | I was a bit confused as to how everything works until I read it
       | in detail. Really cool tools, but I think one thing that would
       | help in the introduction is: saying explicitly that the generated
       | .md document is for you (the user) to read through, observe the
       | output of the CLI call, and ensure that the output matches what
       | you would expect.
       | 
       | It's basically an automated test, but at a higher abstraction
       | level and with manual verification--using CLI tools rather than a
       | test harness. Really great work!
        
       | nzoschke wrote:
       | go-rod has been instrumental to my agentic coding loops too. Some
       | uses:
       | 
       | - E2E testing of browser components
       | 
       | - Taking screenshots before and after and having Claude look at
       | them to double check things
       | 
       | - Driving it with an API and CLI as a headless browser
       | 
       | Will definitely give Rodney a look.
        
       | measurablefunc wrote:
       | Google's antigravity does this automatically by creating Task &
       | Walkthrough artifacts.
        
       | johnfn wrote:
       | Out of curiosity, what is the advantage of using Rodney when
       | Playwright has the same set of features and AI understands how to
       | write a Playwright script very well?
        
         | simonw wrote:
         | Maybe not a lot.
         | 
         | Showboat documents look neater if there are single one-line
         | commands that do something useful. Dumping a full Playwright
         | script into a cell is less readable.
         | 
         | Showboat also has a special feature where you can embed an
         | image directly in the document by running:
         | showboat image doc.md 'rodney screenshot'
         | 
         | The command you call should return a path to an image file as
         | the last line of output. Rodney does exactly that.
         | 
         | It may well turn out that Rodney is unnecessary and people find
         | better patterns using Showboat with existing tools like
         | playwright-cli - in which case it won't matter because Showboat
         | and Rodney aren't coupled to each other at all.
         | 
         | Showboat is definitely the more significant of the two
         | projects.
        
       | water-drummer wrote:
       | Wait, why should an LLM simply not just write directly to the
       | markdown file instead of going through the extra step of using a
       | cli tool which is basically `echo 'something' >> file.md` but
       | with templates that should really be in a prompt instead of a
       | being in a compiled binary? Did Claude come up with the idea for
       | this as well?
       | 
       | Also, I am sure you must already know about Playwright mcp so why
       | this? If your goal isn't to make the cli human-friendly, which is
       | the only advantage clis have over mcps doing the same thing, then
       | why not just use the mcp? It doesn't even handle multiple
       | sessions and has a single global state file--this is slop.
        
         | simonw wrote:
         | Because I don't want it to write to the markdown file directly.
         | I want it to tell me the command it runs and I then run that
         | command and write both the command and the output to the file.
         | 
         | Otherwise it's just writing a document, not building a demo you
         | can review.
         | 
         | As far as I can tell you can't hook MCPs up to Claude Code for
         | web.
         | 
         | I originally planned to support separate sessions but decided
         | to leave that out for the initial release. I've opened an issue
         | for that here: https://github.com/simonw/rodney/issues/6
        
       | mentalgear wrote:
       | A bit like jupyter notebooks, isn't it?
        
         | simonw wrote:
         | Yes, very much so. It's a much thinner, less feature-rich
         | alternative.
         | 
         | It would be interesting to experiment with Jupyter notebooks as
         | an alternative that could work in Claude Code for web.
         | 
         | I had a poke around just now and couldn't find an existing CLI
         | tool that lets you build those up a section at a time in the
         | same way as Showboat. I did find this Python library though:
         | uv run --with nbformat python -c '         import nbformat
         | nb = nbformat.v4.new_notebook()
         | nb.cells.append(nbformat.v4.new_markdown_cell("# NBTerm
         | Exploration"))
         | nb.cells.append(nbformat.v4.new_code_cell("import
         | sys\nprint(f\"Python {sys.version}\")"))
         | nb.cells.append(nbformat.v4.new_code_cell("x = [i**2 for i in
         | range(10)]\nprint(x)"))
         | nb.cells.append(nbformat.v4.new_code_cell("sum(x)"))
         | with open("demo.ipynb", "w") as f:
         | nbformat.write(nb, f)         '
         | 
         | So you could tell the agent to run code like that and then
         | inspect the `demo.ipynb` notebook later on. It doesn't show the
         | result of evaluating the cells though, you need to run this
         | afterwards to have that happen:                   uv run --with
         | nbformat --with nbclient --with ipykernel python -c '
         | import nbformat         from nbclient import NotebookClient
         | nb = nbformat.read("demo.ipynb", as_version=4)         client =
         | NotebookClient(nb, timeout=60)         client.execute()
         | nbformat.write(nb, "demo_executed.ipynb")         '
        
           | mentalgear wrote:
           | Cool, I have to say I find the idea intriguing as a
           | tracability tool in they that LLMs can show you step be step
           | how a program is assembled / an output was generated.
        
         | samuelson wrote:
         | I think it's more about the interface than the output. The
         | agent can add stuff to a markdown file with simple cli commands
         | rather than a more complex editor or file interface.
        
       | sNyZZzzz wrote:
       | Using Markdown as both docs and executable output is cool, but
       | I'm curious how it scales when agents hit more complex ui.
        
       | Sharlin wrote:
       | I can't wait for tools that allow agents to hold stand-ups,
       | retrospectives and sprint planning sessions, all facilitated by
       | an agentic scrum master.
        
         | eclipxe wrote:
         | My clawdbot setup does just that. No joke.
        
       | simlevesque wrote:
       | rodney seems to be pretty much the same as agent-browser:
       | https://github.com/vercel-labs/agent-browser
        
         | simonw wrote:
         | Hah! I hadn't seen that one before. Yeah, the CLI design is
         | very similar.
         | 
         | Main difference is Rodney can be installed as a single Go
         | binary or via uv/pip, agent-browser is Rust and npm.
         | 
         | Looks like agent-browser was first released at the start of
         | January, it's very new.
        
           | simlevesque wrote:
           | Yeah it's pretty new indeed. It's very effective at doing
           | pretty much any browser automation task and I have to say
           | that using it with the included skill is pretty seamless.
        
       | samuelson wrote:
       | I love your content, but I wish you'd make your blog theme
       | responsive for wider screens/non-mobile. I prefer to read content
       | like this on a large screen.
       | 
       | Showboat seems like it could actually be quite useful for humans
       | too, just for making quick notes from a CLI without opening an
       | editor. The "pop" command makes me wonder if there would be a
       | benefit to also having an array-like in addition to the stack-
       | like interface. It seems like it would be fairly trivial to
       | generate an index of markdown blocks so that they could be edited
       | individually.
       | 
       | I like the idea of Rodney, but I wonder if you might actually
       | have better results by asking the agent to generate equivalent
       | Selenium scripts instead. I'm specifically suggesting Selenium
       | because it's been around so long so I assume there's a lot of
       | Selenium in the LLMs training data, but there are other options
       | that might work too.
        
         | simonw wrote:
         | First time someone's asked for the site to be wider! I have it
         | setup so on a wide screen the text is still a readable width,
         | do you think it needs to bump up that max width a bit more?
         | 
         | I've found the models are so good at Playwright that I don't
         | consider Selenium any more. Rodney is my first experiment _not_
         | using Playwright.
        
           | TheKnack wrote:
           | I second the request to make the site responsive. When I load
           | the page the CSS constrains the main content to 560px and the
           | whole page is constrained to 940px. Here's how it displays on
           | my system:
           | 
           | https://i.postimg.cc/zDMD9nYD/Simon.png
        
             | dwb wrote:
             | Your tastes are your own, and there is an argument for just
             | filling the window, but you won't find a typographic
             | authority that advocates setting body text much wider than
             | that (and I would agree with them).
        
             | simonw wrote:
             | Can you take a screenshot of some other site that is wider
             | but has a comfortable reading width for you?
        
       | cadamsdotcom wrote:
       | Great to see you doing red/green TDD Simon!
       | 
       | Passing tests in your repo are great documentation of the tool at
       | a microscopic level. And rerunning tests only burns tokens on
       | failures (since passed tests just print a dot) so it's token
       | efficient too.
       | 
       | Some other neat tricks:
       | 
       | - For greater efficiency configure your test runner to print
       | nothing (not even a dot/filename) for test successes. Agents
       | don't need progress dots, only the exit code & failure details
       | 
       | - Have your agent implement a 10ms timeout per test. pytest has
       | hooks to do this. The agent will see tests time out and mock out
       | all I/O and third party code - why test what one assumes third
       | parties tested already! Your test suite is CPU-bound without a
       | shared database, has no shared data and no tests that interfere
       | with or depend on each other, so tests can run in parallel.
        
         | simonw wrote:
         | That timeout trick is very neat!
         | 
         | I'm OK with longer running tests because I always have them run
         | against a real database (often SQLite, sometimes PostgreSQL)
         | and real files created in temporary directories but I can see
         | how the time limit might be useful for tests that don't need
         | those kind of components.
        
       ___________________________________________________________________
       (page generated 2026-02-10 23:01 UTC)