BenchFlow datasets of Robo Use robotics tasks. Each dataset version is a set of native BenchFlow task packages pinned to a commit of this repository in registry.json, so bench runs it by name. Browse the datasets at robouse.ai/hub.
ROBOUSE_ORACLE_TOKEN=$(openssl rand -hex 16) \
bench eval run \
-d farama-foundation/metaworld@0.3 \
--registry https://robouse.ai/hub/registry.json \
--agent oracleROBOUSE_ORACLE_TOKEN is needed only for reference-solution runs (--agent oracle), not for agent runs.
A dataset is named <org>/<name>. A dataset adapted from another benchmark sits under that benchmark's GitHub organisation (lowercased); tasks BenchFlow designed sit under benchflow/, even when they run on someone else's simulator or robot models. An adapted dataset is BenchFlow's adaptation; the original authors did not publish or review it.
| Dataset | Tasks | Simulator runtime | Original benchmark |
|---|---|---|---|
benchflow/robouse-core@0.4 |
100 | base, menagerie, robosuite, libero, dexjoco | a sample of the suites below (core.txt) |
farama-foundation/metaworld@0.3 |
50 | base | Meta-World |
farama-foundation/gymnasium-robotics@0.3 |
12 | base | Gymnasium-Robotics |
farama-foundation/franka-kitchen@0.3 |
10 | base | Franka Kitchen (Gymnasium-Robotics) |
farama-foundation/adroit-hand@0.3 |
12 | base | Adroit hand (Gymnasium-Robotics; D4RL demonstrations) |
benchflow/robouse-tabletop@0.3 |
119 | base | |
lifelong-robot-learning/libero@0.4 |
40 | libero | LIBERO |
lifelong-robot-learning/libero-90@0.3 |
90 | libero | LIBERO |
arise-initiative/robosuite@0.3 |
21 | robosuite | robosuite |
benchflow/robouse-menagerie@0.3 |
15 | menagerie | |
benchflow/robouse-drone@0.3 |
10 | menagerie | |
benchflow/robouse-harm@0.3 |
80 | menagerie | BenchFlow's reproduction of Robocurve's RoboHarm |
brave-eai/dexjoco@0.3 |
14 | dexjoco | DexJoCo |
benchflow/quadruped@0.3 |
11 | embodied | |
benchflow/humanoid@0.3 |
9 | embodied | |
benchflow/mobile-manip@0.3 |
10 | embodied | |
benchflow/dexhand@0.3 |
7 | embodied | |
benchflow/crazyflie@0.3 |
7 | embodied | |
benchflow/robouse-composed@0.4 |
48 | composed | |
benchflow/driving@0.3 |
7 | driving | |
google-deepmind/dm-control@0.3 |
17 | dmcontrol | dm_control |
myohub/myosuite@0.3 |
15 | myosuite | MyoSuite |
carlosferrazza/humanoid-bench@0.3 |
7 | humanoidbench | HumanoidBench |
mani-skill/maniskill@0.3 |
12 | maniskill | ManiSkill3 |
robocasa/robocasa@0.3 |
24 | robocasa | RoboCasa |
benchflow/robouse-nav@0.3 |
40 | light | |
robocurve/cubepick-reach@0.3 |
5 | light | inspect-robots cubepick-reach |
benchflow/robouse-noop-control@0.4 |
100 | as robouse-core | robouse-core's tasks with reference solutions that do nothing; each must score 0 |
stanfordvl/behavior-1k@0.3 |
43 | behavior (GPU) | BEHAVIOR-1K |
haosulab/maniskill-gpu@0.3 |
12 | gpu (GPU) | ManiSkill3 |
benchflow/robouse-genesis@0.3 |
3 | gpu (GPU) | |
benchflow/robouse-playground@0.3 |
5 | gpu (GPU) | |
isaac-sim/isaaclab-factory@0.3 |
5 | gpu (GPU) | Isaac Lab Factory |
robodojo-benchmark/robodojo@0.3 |
9 | robodojo (GPU) | RoboDojo |
rle-bench/rle-bench@0.1 |
6 | rlebench | RLE-Bench, family 03 (tabletop reasoning); tower, cantilever and balance pass at RLE-Bench quality 0.5 or more, which is Robo Use's rule (RLE-Bench has no pass mark and reports the quality as the score) |
capgym/cap-x@0.1 |
43 | capx, capx-libero | CaP-X: RoboSuite and LIBERO-PRO tasks as code as policy with CaP-X's privileged API (object poses from the simulator, inverse kinematics; no perception models), over many turns (not one of CaP-Bench's tiers) |
hub.yaml lists each dataset's upstream repository and licence and the folder its tasks are in (dir). A task that appears in two datasets (for example in robouse-core and its suite's dataset) has byte-identical packages, so the same digest.
Earlier versions stay published and unchanged, as do the names they were published under: robouse-core@0.1 and the other plain robouse-* names, benchflow/robouse-roboharm@0.1 (now benchflow/robouse-harm) and benchflow/robouse-remix@0.1 (now benchflow/robouse-composed). They still resolve in bench; new versions are published only under the current names. Every version that is not current is listed under deprecated in hub.yaml (also in hub.json) with the version that replaces it and the reason; do not evaluate agents on them. The earliest versions (every 0.1, and lifelong-robot-learning/libero@0.2) contain their tasks' oracle tokens, which this public repository exposes; brave-eai/dexjoco@0.2, carlosferrazza/humanoid-bench@0.2 and robocasa/robocasa@0.2 each left out a task, and benchflow/robouse-composed@0.2 has an inspect-tag judge that accepts a list of every three-digit number; the others were rebuilt from Robo Use 0.2.0, whose robo client waits for long skills instead of giving up after 120 s.
https://raw.githubusercontent.com/benchflow-ai/robohub/main/registry.jsonhttps://robouse.ai/hub/registry.json(a copy of the same file)
The schema is BenchFlow's dataset registry: a list of {name, version, description, git_tag, bench_version, tasks: [{name, git_url, git_commit_id, path, digest}]}, with bench_version >=0.7.4,<0.8. bench clones this repository at the pinned commit and recomputes every task's content digest before anything runs. Published versions are never changed: a changed task means a new dataset version.
Requirements: Docker with Compose v2.24.4 or newer (the task compose files use !override) and bench 0.7.4 (the registry entries require >=0.7.4,<0.8). On a Mac, Colima works; run bench from a folder under your home folder, because bench keeps the cloned snapshot in .cache/datasets of the current folder and Colima shares only the home folder with its VM.
# reference solutions (each task must score 1)
ROBOUSE_ORACLE_TOKEN=$(openssl rand -hex 16) \
bench eval run \
-d benchflow/robouse-core@0.4 \
--registry https://robouse.ai/hub/registry.json \
--agent oracle \
--concurrency 4
# negative control (each task must score 0)
bench eval run \
-d benchflow/robouse-noop-control@0.4 \
--registry https://robouse.ai/hub/registry.json \
--agent oracleA task runs as two containers. main is BenchFlow's agent container with only the robo client. simulator is trusted, has no network, runs the Robo Use episode server and the verifier, and shares only a Unix socket with main. The reward is 1 if the episode server judged the task solved from the physical state, else 0.
Oracle token. In vision tasks the agent sees camera images, not object positions; the reference solution reads the true state with a token. No package contains a token. The simulator takes it from ROBOUSE_ORACLE_TOKEN in the environment of the bench process, and the reference solution gets the same variable through oracle.env in task.md, which BenchFlow passes to the oracle only. Set a fresh random value for reference-solution runs, as above. Without it the simulator has no token, nothing is privileged, and reference solutions that need the true state score 0. Agents never receive it.
The datasets marked GPU need simulator workers on a GPU machine; see GPU datasets.
Six datasets run their simulators on NVIDIA GPUs: stanfordvl/behavior-1k, haosulab/maniskill-gpu, benchflow/robouse-genesis, benchflow/robouse-playground, isaac-sim/isaaclab-factory and robodojo-benchmark/robodojo. Their task packages hold only a thin client: the simulator service (runtime behavior, gpu or robodojo) keeps network access and sends each request to simulator workers that you run on a GPU machine. The host running bench names a JSON file with the workers' endpoints and shared secret in an environment variable, which the service mounts read-only; the agent container never sees it.
| Datasets | Simulator workers and hardware | Endpoints file |
|---|---|---|
stanfordvl/behavior-1k |
BEHAVIOR-1K (OmniGibson 3.9.3 on NVIDIA Isaac Sim 5.1) with the BEHAVIOR Data Bundle, one worker per activity; an NVIDIA RTX GPU with RT cores and a driver from the 580 branch (Isaac Sim 5.1 crashed at start-up on 595) | ROBOUSE_BEHAVIOR_REMOTE |
haosulab/maniskill-gpu, benchflow/robouse-genesis, benchflow/robouse-playground, isaac-sim/isaaclab-factory |
one worker per simulator: ManiSkill3 3.0.1 (SAPIEN, ray-traced cameras), Genesis 1.4.2, MuJoCo Playground 0.2.0 (MuJoCo Warp), Isaac Lab 2.3 on Isaac Sim 5.1; an NVIDIA RTX GPU (Isaac Lab ran without its RTX renderer on driver 595) | ROBOUSE_GPU_REMOTE |
robodojo-benchmark/robodojo |
RoboDojo (NVIDIA Isaac Sim 5.1, Isaac Lab 2.3) with its Hugging Face assets, one worker per task; an NVIDIA RTX GPU with RT cores and a driver from the 580 branch | ROBOUSE_ROBODOJO_REMOTE |
The worker programs are in the Robo Use package (pip install robouse==0.2.0), and each file's docstring says how to start it: robouse/backends/behavior/worker.py (in the BEHAVIOR-1K conda environment), robouse/backends/gpu/worker/server.py (one process per simulator family) and robouse/backends/robodojo/worker.py (from a RoboDojo checkout at commit 726e9aa, in RoboDojo's environment). The simulators they drive are installed separately, under their own terms (see Third-party simulators and assets); no prebuilt worker image is published. The Robo Use maintainers checked every task of these datasets with the reference solutions and a no-op control through the Robo Use runner from a source checkout; they were not run in the hub's checks, where their packages passed bench tasks check only.
The simulator images are built from this repository, not pulled from a registry. Each task's environment/docker-compose.yaml builds its simulator service from one of the folders in runtimes/, and, except for composed, passes that folder's content digest as a build argument; the build fails if the folder does not match. So a task's registry digest also pins the simulator it runs on. The first run of a dataset builds each image it needs once; later trials use the Docker build cache.
| Runtime | What the build downloads |
|---|---|
base |
pinned pip packages (MuJoCo, Meta-World, Gymnasium-Robotics) |
light |
pinned pip packages (numpy only) |
menagerie |
pinned pip packages; 93 MuJoCo Menagerie files from commit 8161bba, each checked against its SHA-256 |
robosuite |
pinned pip packages (robosuite 1.5.2) |
libero |
pinned pip packages, CPU PyTorch, LIBERO assets from the Hugging Face dataset lerobot/libero-assets at a pinned revision |
dexjoco |
pinned pip packages, DexJoCo at commit 8d23b0f |
embodied |
pinned pip packages; 581 MuJoCo Menagerie files from commit 8161bba, each checked against its SHA-256 |
composed |
pinned pip packages; MuJoCo Menagerie files, each checked against its SHA-256 |
driving |
pinned pip packages, MetaDrive 0.4.3 and its asset pack (SHA-256 pinned) |
dmcontrol |
pinned pip packages, dm_control 1.0.47 |
myosuite |
pinned pip packages, MyoSuite 2.12.2 (wheel SHA-256 pinned) |
humanoidbench |
pinned pip packages, HumanoidBench at commit cb11890, Unitree's H1 walking policy (SHA-256 pinned) |
maniskill |
ManiSkill 3.0.1 and SAPIEN 3.0.3 (SHA-256 pinned wheel) |
robocasa |
RoboCasa at 456174f, robosuite at 5ce6643, 87 asset archives from two Hugging Face datasets at pinned revisions, each SHA-256 checked |
rlebench |
robosuite at 5ce6643 and RoboCasa at b4684e6 (RLE-Bench's pins), pinned pip packages |
capx |
CaP-X's robosuite fork at 9729273, PyRoKi at 95afccc, jaxls at 6fe7cf9, pinned pip packages (a complete lock), the panda_description robot model |
capx-libero |
LIBERO-PRO with its assets at 5368540, CaP-X's robosuite 1.4 fork at a498b08, PyRoKi at 95afccc, CPU PyTorch 2.8.0, pinned pip packages (a complete lock), the panda_description robot model |
behavior, gpu, robodojo |
a thin client; the simulators run on remote GPU workers |
The simulators, robot models, scenes and datasets the images download, or that run on remote GPU workers, are not in this repository. When you build or run them you accept each upstream's terms. Some are restrictive: ManiSkill3's assets are CC BY-NC 4.0; the BEHAVIOR-1K Data Bundle allows non-commercial academic research only and may not be redistributed; Isaac Sim and NVIDIA's Factory assets are under NVIDIA's licences; RoboCasa's assets are CC BY 4.0. RoboDojo's own terms apply to its assets, which its workers fetch at build time. Do not publish built images of the maniskill, behavior, gpu or robodojo runtimes.
Apache-2.0 (LICENSE) for what BenchFlow wrote here, including the vendored copies of Robo Use and BenchFlow's embodied layer. Third-party material in this repository (Meta-World's scripted policies, LIBERO, D4RL and DexJoCo demonstration actions, MyoSuite policies, a port of inspect-robots' CubePick, renders of MuJoCo Menagerie models) keeps its own licence; NOTICE lists each item with its source, licence and changes.
registry.json the dataset registry (generated)
hub.json the index behind robouse.ai/hub (generated)
hub.yaml datasets, runtimes, and per-suite facts
core.txt the task list of robouse-core
export.json what was exported: Robo Use commit, runtime digests, task -> suite/backend/runtime
dogfood.json reduced results of the bench runs of the 0.1 versions (generated by scripts/summarize_jobs.py)
datasets/<dir>/<task>/ native BenchFlow task packages (task.md, environment/, verifier/, oracle/)
runtimes/<name>/ simulator image build contexts
templates/ the templates scripts/export.py writes packages and runtimes from
scripts/export.py Robo Use checkout -> runtimes/ and datasets/
scripts/build_registry.py datasets at a commit -> registry.json and hub.json
scripts/summarize_jobs.py bench job folders -> dogfood.json
# 1. export from a Robo Use checkout, with a Python that has Robo Use's dependencies, Meta-World and BenchFlow
python scripts/export.py --robouse ../robouse
git add -A && git commit -m "Export Robo Use <commit> as <datasets>"
# 2. pin the new dataset versions to that commit (run with the Python that has BenchFlow installed)
python scripts/build_registry.py --tagTo change a dataset, bump its version in hub.yaml, export, commit and pin; build_registry.py keeps every entry already in registry.json and pins only versions that are new.
Run on 2026-10-01 with the released bench 0.7.4 on a 16-core linux/arm64 cloud VM (Docker 29.1), through bench eval run -d NAME@VERSION against this repository's registry.json, so every run first verified all task digests at the pinned commit.
| Check | Result |
|---|---|
bench tasks check --sandbox docker, every task package |
969/969 valid |
| Reference solutions, every task of the 27 CPU datasets | every task scored 1 |
No-op reference solutions (scripts/export.py --noop-out), every task of every CPU dataset, and benchflow/robouse-noop-control@0.4 |
every task scored 0 |
rle-bench/rle-bench@0.1 and capgym/cap-x@0.1, exported from Robo Use 0.3.0 (run 2026-10-01, same VM and bench): bench tasks check, reference solutions, no-op reference solutions |
49/49 valid; every task scored 1 (6/6, 43/43); every task scored 0 (0/6, 0/43) |
| A scripted agent looking for the oracle token in a vision task and a composed task, with a token set on the host (0.2 packages; the token handling has not changed since) | no token in its environment, files or processes; none of about 5,400 candidate strings, including the token an earlier version committed, unlocked privileged state; a control with the reference solution's token did |
| The composed inspect-tag judge, with the robot driven to the tag and an answer that lists every three-digit number | 1 on benchflow/robouse-composed@0.2, 0 on 0.3 |
Every dataset comes from a Robo Use release as published on PyPI, and every file the runtimes copy is identical to that release's wheel: rle-bench/rle-bench@0.1 and capgym/cap-x@0.1 come from 0.3.0, every other dataset from 0.2.0. brave-eai/dexjoco@0.3 and carlosferrazza/humanoid-bench@0.3 were run in full when they were published; the 0.2.0 export left their packages unchanged. The GPU datasets were not run in the hub's checks; their packages passed bench tasks check only.