[HN Gopher] Show HN: Cua-Bench - a benchmark for AI agents in GU...
___________________________________________________________________
Show HN: Cua-Bench - a benchmark for AI agents in GUI environments
Hey HN, we're excited to share Cua-Bench (
https://github.com/trycua/cua ), an open-source framework for
evaluating and training computer-use agents across different
environments. Computer-use agents show massive performance
variance across different UIs--an agent with 90% success on Windows
11 might drop to 9% on Windows XP for the same task. The problem is
OS themes, browser versions, and UI variations that existing
benchmarks don't capture. The existing benchmarks (OSWorld,
Windows Agent Arena, AndroidWorld) were great but operated in silos
--different harnesses, different formats, no standardized way to
test the same agent across platforms. More importantly, they were
evaluation-only. We needed environments that could generate
training data and run RL loops, not just measure performance. Cua-
Bench takes a different approach: it's a unified framework that
standardizes environments across platforms and supports the full
agent development lifecycle--benchmark, train, deploy. With Cua-
Bench, you can: - Evaluate agents across multiple benchmarks with
one CLI (native tasks + OSWorld + Windows Agent Arena adapters) -
Test the same agent on different OS variations (Windows
11/XP/Vista, macOS themes, Linux, Android via QEMU) - Generate new
tasks from natural language prompts - Create simulated
environments for RL training (shell apps like Spotify, Slack with
programmatic rewards) - Run oracle validations to verify
environments before agent evaluation - Monitor agent runs in real-
time with traces and screenshots All of this works on macOS,
Linux, Windows, and Android, and is self-hostable. To get started:
Install cua-bench: % pip install cua-bench Run a basic
evaluation: % cb run dataset datasets/cua-bench-basic --agent demo
Open the monitoring dashboard: % cb run watch <run_id> For
parallelized evaluations across multiple workers: % cb run dataset
datasets/cua-bench-basic --agent your-agent --max-parallel 8 Want
to test across different OS variations? Just specify the
environment: % cb run task slack_message --agent your-agent --env
windows_xp % cb run task slack_message --agent your-agent --env
macos_sonoma Generate new tasks from prompts: % cb task generate
"book a flight on kayak.com" Validate environments with oracle
implementations: % cb run dataset datasets/cua-bench-basic
--oracle The simulated environments are particularly useful for RL
training--they're HTML/JS apps that render across 10+ OS themes
with programmatic reward verification. No need to spin up actual
VMs for training loops. We're seeing teams use Cua-Bench for: -
Training computer-use models on mobile and desktop environments -
Generating large-scale training datasets (working with labs on
millions of screenshots across OS variations) - RL fine-tuning
with shell app simulators - Systematic evaluation across OS themes
and browser versions - Building task registries (collaborating
with Snorkel AI on task design and data curation, similar to their
Terminal-Bench work) Cua-Bench is 100% open-source under the MIT
license. We're actively developing it as part of Cua
(https://github.com/trycua/cua), our Computer Use Agent SDK, and
we'd love your feedback, bug reports, or feature ideas. GitHub:
https://github.com/trycua/cua Docs: https://cua.ai/docs/cuabench
Technical Report: https://cuabench.ai We'll be here to answer any
technical questions and look forward to your comments!
Author : someguy101010
Score : 33 points
Date : 2026-01-26 17:46 UTC (2 days ago)
(HTM) web link (github.com)
(TXT) w3m dump (github.com)
| visarga wrote:
| Interesting, a computer use environment. I made a CUA benchmark
| too, 200 web tasks with internal code based evaluation. You can
| integrate them if you want.
|
| https://github.com/UiPath/uipath_enterprise_benchmark
|
| https://arxiv.org/abs/2511.17131
| frabonacci wrote:
| Hey visarga - I'm the founder of Cua, we might have met at the
| CUA ICML workshop? The OS-agnostic VNC approach of your
| benchmark is smart and would make integration easy. We're open
| to collaborating - want to shoot me an email at f@trycua.com?
| augusteo wrote:
| The trajectory export feature is smart. Evaluation and training
| data collection in the same tool.
|
| I'm curious how the benchmarks handle non-determinism. Real GUIs
| have loading states, animations, popups that appear sometimes but
| not always. Does cuabench control for that, or is variance just
| part of the measurement?
|
| Also interested in what "Windows Arena" tests specifically.
| Windows has so many edge cases - UAC prompts, driver install
| dialogs, random update notifications. Those feel like the hard
| mode for computer-use agents.
| frabonacci wrote:
| Thanks - trajectory export was key for us since most teams want
| both eval and training data.
|
| On non-determinism: we actually handle this in two ways. For
| our simulated environments (HTML/JS apps like the Slack/CRM
| clones), we control the full render state so there's no
| variance from animations or loading states. For native OS
| environments, we use explicit state verification before scoring
| - the reward function waits for expected elements rather than
| racing against UI timing. Still not perfect, but it filters out
| most flaky failures.
|
| Windows Arena specifically - we're focusing on common
| productivity flows (file management, browser tasks, Office
| workflows) rather than the edge cases you mentioned. UAC
| prompts and driver dialogs are exactly the hard mode scenarios
| that break most agents today. We're not claiming to solve those
| yet, but that's part of why we're open-sourcing this - want to
| build out more adversarial tasks with the community.
| rfw300 wrote:
| Interesting project, but the lack of any actual benchmark results
| on existing models/agents is disappointing.
| frabonacci wrote:
| Fair point - we just open-sourced this so benchmark results are
| coming. We're already working with labs on evals, focusing on
| tasks that are more realistic than OSWorld/Windows Agent Arena
| and curated with actual workers. If you want to run your agent
| on it we'd love to include your results.
___________________________________________________________________
(page generated 2026-01-28 23:01 UTC)