Introducing CUA-Lite, an open platform for developing computer-use agents. It is built around three standardized abstractions — Lite.Gym for environments, Lite.Sample for supervised data, and one harness per model across eval, SFT and RL. 15+ benchmarks, 10+ datasets and 30k+ verifiable tasks are already plugged in.

Code ↗:https://github.com/cua-lite/cua-lite · Project site ↗:https://cua-lite.github.io

Computer-use agents operate desktop, browser and mobile applications, and developing one takes environments with verifiable tasks, supervised data, and models, plus the harness that runs a model in an environment. Open-source efforts already exist for all three, but they remain scattered: each environment has its own interface and runtime; each supervised dataset its own format; each model its own action space, prompt format and rollout code.

Without a shared standard, connecting models to environments means rewriting the same code: an action mapping, a context window, and a rollout loop — so resources cannot be pooled and scaled, nor can evaluation, SFT and RL share one infrastructure.

CUA-Lite closes all three , in one place:

Any agent plugs into any environment through one interface. lite.gym exposes an environment as a Gym-style reset / step / close loop. The loop standardizes one observation format across environments — a screenshot, text, or both per tool call — and one GUI action space per platform, issued as tool calls: click and type on desktop and browser, tap and swipe on mobile, plus extra tools per environment such as bash in our desktop sandboxes. It wraps the environment's own runtime: a VM, a browser, or a mobile emulator.

CUA-Lite also provides optional, lightweight VM-free desktop sandboxes of its own: Docker containers that replicate OSWorld's desktop at much lower cost and hold 30k+ verifiable tasks for training. They need no /dev/kvm , the hardware virtualization VM-based benchmarks require, so they run on any host with Docker — and Lite.OSWorld (ours) runs OSWorld's own tasks and evaluators, unchanged.

The VM-free container isn't just for OSWorld — it's the foundation for CUA-Lite's family of sandboxes. The same base already runs browser and desktop tasks, and real science desktops: GMAT flying spacecraft, PyMOL turning proteins.

Call for sandbox contributors. A sandbox only matters while people run it. Add yours to CUA-Lite, and every agent trains and benchmarks on it — now and later. One integration, and the whole field builds on it.

Lite.* environments ↗:https://github.com/cua-lite/cua-lite/tree/main/lite/gym/envs/lite · Env guide ↗:https://github.com/cua-lite/cua-lite/blob/main/docs/envs.md · Leaderboard ↓:#lb

Convert a dataset once, and every agent can train on it. LiteSample is the one schema, shared across every env, agent, and task type, free on Hugging Face.

One shape for all of it, whatever it came from: messages whose tool calls are the interface's actions and whose tool results are its observations — from a GUI grounding label to a full rollout.

CUA-Lite ships an adapter per model, packing a unified LiteSample into the exact training format each one needs. The figure above shows one, with the building blocks to add your own.

10+ datasets are already on Hugging Face : existing CUA corpora — grounding · understanding · use — preprocessed into Lite.Sample, plus fresh rollouts from frontier CUAs. Browse the corpora:https://huggingface.co/collections/cua-lite/corpora and the rollouts:https://huggingface.co/collections/cua-lite/rollouts; below, one of them, WebGym:

Call for data contributors. Data only matters while models can train on it. Share yours with CUA-Lite, and every agent trains on it — even models that don't exist yet. One conversion, and the whole community trains on it.

Preprocessing guide ↗:https://github.com/cua-lite/cua-lite/blob/main/lite/data/preproc/AGENTS.md · Agent harnesses ↗:https://github.com/cua-lite/cua-lite/tree/main/lite/agents · SFT guide ↗:https://github.com/cua-lite/cua-lite/blob/main/docs/sft.md

Each model has one harness, the code that adapts it to the interface and the format. Through its harness a model runs in every integrated environment, and eval and RL consume the rollouts it produces; with the same harness, any stored LiteSample is rendered into the model's own training format for SFT — here, Qwen3.5's:

Set --model-id for the agent and --env-id for the benchmark:

15+ benchmarks are already integrated — ours is the VM-free runtime, the interface and the integration, not the task suites — and the VM-free desktop sandboxes are an addition, not a replacement: the original OSWorld VM sits right beside Lite.OSWorld, and the mobile benchmarks still need a VM or an emulator:

SFT on CUA-Lite's corpora, then reinforce in its envs — GRPO and beyond, on the Slime:https://github.com/THUDM/slime trainer. Train any open agent on any data and any env:

Lite.Sample, adapted to each model — pick a dataset and a student:

Rollouts scored in the env drive GRPO updates — pick a model and env:

Bring a dataset, an env, or an agent — each one compounds. Or just tell us what's missing — GitHub:https://github.com/cua-lite/cua-lite · Hugging Face:https://huggingface.co/cua-lite · Email:mailto:zhanhui@berkeley.edu.