介绍 CUA-Lite,这是一个用于开发计算机使用代理的开放平台。它围绕三个标准化抽象构建——Lite.Gym 用于环境,Lite.Sample 用于监督数据,以及跨评估(eval)、监督微调(SFT)和强化学习(RL)的每个模型的一个接口。已有 15+ 个基准测试、10+ 个数据集和 30k+ 可验证任务接入。
代码 ↗:https://github.com/cua-lite/cua-lite · 项目网站 ↗:https://cua-lite.github.io
计算机使用代理可以操作桌面、浏览器和移动应用,而开发一个代理需要具备可验证任务的环境、监督数据和模型,以及在环境中运行模型的接口。开源项目已经涵盖了这三方面,但它们仍然分散:每个环境都有自己的接口和运行时;每个监督数据集都有自己的格式;每个模型都有自己的动作空间、提示格式和回滚(rollout)代码。
如果没有共享标准,将模型连接到环境意味着要重写相同的代码:动作映射、上下文窗口和回滚循环——因此资源无法整合和扩展,评估、SFT 和 RL 也无法共享一个基础设施。
CUA-Lite 将这三者整合在一个平台上:
任何代理都可以通过一个接口接入任何环境。lite.gym 将环境暴露为 Gym 风格的 reset / step / close 循环。该循环在环境间标准化一种观察格式——每次工具调用返回截图、文本或两者结合——以及每个平台的统一 GUI 动作空间,以工具调用形式发出:在桌面和浏览器上点击和输入,在移动端上点击和滑动,并为每个环境提供额外工具,例如我们桌面沙箱中的 bash。它封装了环境自身的运行时:虚拟机(VM)、浏览器或移动模拟器。
CUA-Lite 还提供了可选的轻量级无虚拟机桌面沙箱:Docker 容器,以更低成本复现 OSWorld 的桌面,并包含 30k+ 可验证任务用于训练。它们无需 /dev/kvm,这是基于硬件虚拟化的 VM 基准测试所需的,因此可以在任何支持 Docker 的主机上运行——Lite.OSWorld(我们的版本)能够运行 OSWorld 自身的任务和评估器,完全不变。
无虚拟机容器不仅用于 OSWorld——它是 CUA-Lite 沙箱家族的基础。相同的基础已经运行浏览器和桌面任务,以及真实科学桌面:GMAT 飞行航天器,PyMOL 翻转蛋白质。
呼吁沙盒贡献者。沙盒只有在有人运行时才有意义。将你的沙盒添加到 CUA-Lite,每个代理都会在其上进行训练和基准测试——现在和未来。一次集成,全领域都可以建立在其之上。
Lite.* 环境 ↗:https://github.com/cua-lite/cua-lite/tree/main/lite/gym/envs/lite · 环境指南 ↗:https://github.com/cua-lite/cua-lite/blob/main/docs/envs.md · 排行榜 ↓:#lb
数据集只需转换一次,每个代理都可以在其上进行训练。LiteSample 是统一的模式,在每个环境、代理和任务类型间共享,Hugging Face 免费提供。
统一格式,无论来源:消息的工具调用是界面的操作,工具结果是其观察——从 GUI 基础标签到完整的执行。
CUA-Lite 为每个模型提供一个适配器,将统一的 LiteSample 打包成每个模型所需的训练格式。上图显示了一个示例,并展示了添加你自己模块的构建块。
已有 10+ 数据集在 Hugging Face 上:现有 CUA 语料库——基础标签 · 理解 · 使用——已预处理为 Lite.Sample,以及前沿 CUA 的新执行。浏览语料库:https://huggingface.co/collections/cua-lite/corpora 和执行:https://huggingface.co/collections/cua-lite/rollouts;下面是其中之一,WebGym:
呼吁数据贡献者。数据只有在模型可以训练时才有意义。将你的数据分享给 CUA-Lite,每个代理都可以在其上训练——即使是尚不存在的模型。一次转换,全社区都可以训练。
预处理指南 ↗:https://github.com/cua-lite/cua-lite/blob/main/lite/data/preproc/AGENTS.md · 代理适配器 ↗:https://github.com/cua-lite/cua-lite/tree/main/lite/agents · SFT 指南 ↗:https://github.com/cua-lite/cua-lite/blob/main/docs/sft.md
每个模型都有一个适配器,用于将模型适配到接口和格式。通过适配器,模型可以在每个已集成环境中运行,评估和强化学习使用它生成的执行结果;使用相同的适配器,任何存储的 LiteSample 都可以转换为模型自己的 SFT 训练格式——这里展示了 Qwen3.5 的示例:
为代理设置 --model-id,为基准测试设置 --env-id:
已经集成了15+个基准测试——我们的重点是无虚拟机的运行时、接口和集成,而不是任务套件——无虚拟机的桌面沙箱是一个补充,而不是替代:原始的OSWorld虚拟机与Lite.OSWorld并排存在,移动基准测试仍然需要虚拟机或模拟器:
在CUA-Lite的语料库上进行SFT,然后在其环境中强化——GRPO及以后,在Slime上:https://github.com/THUDM/slime 训练器。可以在任何数据和任何环境中训练任何开放代理:
Lite.Sample,适配每个模型——选择一个数据集和一个学生:
在环境中评分的回滚驱动GRPO更新——选择一个模型和环境:
带上一个数据集、一个环境或一个代理——每一个都会累加。或者直接告诉我们缺少什么——GitHub:https://github.com/cua-lite/cua-lite · Hugging Face:https://huggingface.co/cua-lite · 邮箱:mailto:zhanhui@berkeley.edu。
Introducing CUA-Lite, an open platform for developing computer-use agents. It is built around three standardized abstractions — Lite.Gym for environments, Lite.Sample for supervised data, and one harness per model across eval, SFT and RL. 15+ benchmarks, 10+ datasets and 30k+ verifiable tasks are already plugged in.
Code ↗:https://github.com/cua-lite/cua-lite · Project site ↗:https://cua-lite.github.io
Computer-use agents operate desktop, browser and mobile applications, and developing one takes environments with verifiable tasks, supervised data, and models, plus the harness that runs a model in an environment. Open-source efforts already exist for all three, but they remain scattered: each environment has its own interface and runtime; each supervised dataset its own format; each model its own action space, prompt format and rollout code.
Without a shared standard, connecting models to environments means rewriting the same code: an action mapping, a context window, and a rollout loop — so resources cannot be pooled and scaled, nor can evaluation, SFT and RL share one infrastructure.
CUA-Lite closes all three , in one place:
Any agent plugs into any environment through one interface. lite.gym exposes an environment as a Gym-style reset / step / close loop. The loop standardizes one observation format across environments — a screenshot, text, or both per tool call — and one GUI action space per platform, issued as tool calls: click and type on desktop and browser, tap and swipe on mobile, plus extra tools per environment such as bash in our desktop sandboxes. It wraps the environment's own runtime: a VM, a browser, or a mobile emulator.
CUA-Lite also provides optional, lightweight VM-free desktop sandboxes of its own: Docker containers that replicate OSWorld's desktop at much lower cost and hold 30k+ verifiable tasks for training. They need no /dev/kvm , the hardware virtualization VM-based benchmarks require, so they run on any host with Docker — and Lite.OSWorld (ours) runs OSWorld's own tasks and evaluators, unchanged.
The VM-free container isn't just for OSWorld — it's the foundation for CUA-Lite's family of sandboxes. The same base already runs browser and desktop tasks, and real science desktops: GMAT flying spacecraft, PyMOL turning proteins.
Call for sandbox contributors. A sandbox only matters while people run it. Add yours to CUA-Lite, and every agent trains and benchmarks on it — now and later. One integration, and the whole field builds on it.
Lite.* environments ↗:https://github.com/cua-lite/cua-lite/tree/main/lite/gym/envs/lite · Env guide ↗:https://github.com/cua-lite/cua-lite/blob/main/docs/envs.md · Leaderboard ↓:#lb
Convert a dataset once, and every agent can train on it. LiteSample is the one schema, shared across every env, agent, and task type, free on Hugging Face.
One shape for all of it, whatever it came from: messages whose tool calls are the interface's actions and whose tool results are its observations — from a GUI grounding label to a full rollout.
CUA-Lite ships an adapter per model, packing a unified LiteSample into the exact training format each one needs. The figure above shows one, with the building blocks to add your own.
10+ datasets are already on Hugging Face : existing CUA corpora — grounding · understanding · use — preprocessed into Lite.Sample, plus fresh rollouts from frontier CUAs. Browse the corpora:https://huggingface.co/collections/cua-lite/corpora and the rollouts:https://huggingface.co/collections/cua-lite/rollouts; below, one of them, WebGym:
Call for data contributors. Data only matters while models can train on it. Share yours with CUA-Lite, and every agent trains on it — even models that don't exist yet. One conversion, and the whole community trains on it.
Preprocessing guide ↗:https://github.com/cua-lite/cua-lite/blob/main/lite/data/preproc/AGENTS.md · Agent harnesses ↗:https://github.com/cua-lite/cua-lite/tree/main/lite/agents · SFT guide ↗:https://github.com/cua-lite/cua-lite/blob/main/docs/sft.md
Each model has one harness, the code that adapts it to the interface and the format. Through its harness a model runs in every integrated environment, and eval and RL consume the rollouts it produces; with the same harness, any stored LiteSample is rendered into the model's own training format for SFT — here, Qwen3.5's:
Set --model-id for the agent and --env-id for the benchmark:
15+ benchmarks are already integrated — ours is the VM-free runtime, the interface and the integration, not the task suites — and the VM-free desktop sandboxes are an addition, not a replacement: the original OSWorld VM sits right beside Lite.OSWorld, and the mobile benchmarks still need a VM or an emulator:
SFT on CUA-Lite's corpora, then reinforce in its envs — GRPO and beyond, on the Slime:https://github.com/THUDM/slime trainer. Train any open agent on any data and any env:
Lite.Sample, adapted to each model — pick a dataset and a student:
Rollouts scored in the env drive GRPO updates — pick a model and env:
Bring a dataset, an env, or an agent — each one compounds. Or just tell us what's missing — GitHub:https://github.com/cua-lite/cua-lite · Hugging Face:https://huggingface.co/cua-lite · Email:mailto:zhanhui@berkeley.edu.