Training Hubs
Publishable packages that expose gym-anything benchmarks to external RL stacks (prime-rl / verifiers today).
Training stacks like prime-rl
consume environments as installable packages that expose a
load_environment() function (the
verifiers protocol). The
extras/hubs/ tree holds those packages: one folder per provider, one
package per benchmark, each a thin shell over the library.
The layers
A hub package contains no logic at all. The work is split so N providers times M benchmarks costs one adapter per provider and one declaration per benchmark:
- Benchmark layout contract (core) —
gym_anything.registryenumerates any benchmark root that follows the standard shape (environments/<env>/env.json+tasks/<id>/,splits/*_split.json,splits/verified.json). Adding a benchmark folder makes it consumable by every downstream surface with no integration-specific binding code. - Protocol adapter (core) —
gym_anything.integrations.prime_rl.verifiersturns any gym-anything env into a verifiersMultiTurnEnvthat runs a real reference agent'sstep()verbatim — the same class--agentselects locally. The one substitution is the agent's model call (self.llm_call): a bridge suspendsstep(), hands the messages the agent just built to the framework as the turn's prompt, and resumes it with the sampled completion. Message construction, history management, and parsing are the agent's own code, so the local and driven harnesses cannot diverge (tests/test_agent_driver_parity.pypins this). The episode is scored by the env's real verifier while the VM is still alive.gym_anything.integrations.prime_rl.hubsupplies the dataset rows (build_task_rows) and the canonicalload_environmentsurface. Installed with theprime-rlextra:pip install "gym-anything[prime-rl]". - Publishable package — CUA-World's package,
cua-world(packaging/cua-world/), carries the corpus and its hub entry points. Itscua_worldmodule is a declaration:make_hub_loader("cua_world", env_id="cua-world", ...).env_args(what prime-rl workers use to re-instantiate the environment) is derived from the canonical signature, so it stays faithful by construction.
Agents mirror --agent / --agent_args on local evaluation: pass agent
(a class name from agents/agents/, e.g. "Qwen3VLAgent") and agent_args
in the env args. Any agent whose step() routes its model call through
self.llm_call and speaks OpenAI messages is drivable; provider-native
agents (Anthropic/Google native tools) are rejected with a clear error.
Because agents may rebuild their prompt each turn (windowed history), use
prime-rl's trajectory_strategy = "branching" when training with such
agents; append-only agents also work with "interleaved".
Quickstart
# From the repo root: editable gym-anything, then the cua-world package
# without deps so it resolves your checkout instead of the main-branch pin.
uv pip install -e ".[modal,prime-rl,benchmark,agents]"
uv pip install --no-deps packaging/cua-world
vf-eval cua-world -m google/gemini-3.5-flash \
-b https://api.pinference.ai/api/v1 -k PRIME_API_KEY \
-n 1 -r 1 \
-a '{"env_names": ["gimp_env"], "task_ids": ["horizontal_mirror"], "max_turns": 10}'With the default runner="modal" each rollout boots a real Linux, Windows,
or Android guest in a Modal VM Sandbox (see
Runners). Pass remote_url instead to run
rollouts on a gym-anything remote cluster; the worker
then picks the runner.
In a prime-rl training config the environment is referenced by id, and
args is forwarded verbatim to load_environment(**args):
[[orchestrator.train.env]]
id = "cua-world"
args = { env_names = ["gimp_env"], split = "train", max_turns = 15 }Rewards
The reward is the benchmark's own. At episode end the task's post_task
hook runs in the VM, the task's verifier.py produces
{passed, score, feedback}, and the score maps to reward exactly as
gym-anything does natively (sparse tasks give 0/1, partial/rubric give
score/100). Zero-weight diagnostic metrics (verifier_passed,
verifier_score, actions_executed, parse_errors) ride along. Pass
verifier_mode="vlm_checklist" (plus vlm_* args) to grade against each
task's vlm_checklist.json with a VLM instead.
Publishing
prime env push --path packaging/cua-world publishes to the Environments
Hub. The package depends on gym-anything from its main branch so it
installs standalone and always gets the latest core (no [tool.uv.sources]
path override: prime env push installs from the source directory, where a
local path would not exist). To publish a new version, bump the package
version and re-push.
Adding a provider or benchmark
A new provider gets extras/hubs/<provider>/, and if it speaks a different
protocol, a sibling adapter module in gym_anything/integrations/. A new
benchmark that follows the layout contract needs no integration code at
all: pass its root or package name to load_benchmark_environment, or add
a make_hub_loader declaration to the benchmark's own package, as
cua-world does, if it should have its own hub listing. Shells must consume core through the
public API only.