Runners
The different ways Gym Anything can start environments.
Runner is the project word for how an environment gets started.
You don't need this page when first trying Gym Anything.
It becomes important when:
- something doesn't start on your machine
- you want a different way to launch environments
- you need a feature that only some launch methods support
Common Options
We support several launch methods, including:
- Docker
- Apptainer direct
- QEMU
- AVD
- AVF on Apple Silicon
- Local
Easiest First Step
If you're unsure what works on your machine, gym-anything doctor is the fastest way to see what's installed and what's missing.
gym-anything doctorWhat Changes Between Launch Methods
Different launch methods support different things.
For example, depending on how you start the environment, support may differ for:
- caching
- recording
- save-and-restore features
- platform-specific behavior
You usually only need to care about these differences once you move beyond the first run.
Quick Feature Table
If your question is simply "does this support caching?" or "can I use savevm here?", use this table:
| Launch method | Caching | use_savevm | Live recording |
|---|---|---|---|
| Docker | Yes | No | Yes |
| Apptainer direct | No | No | No |
| QEMU | Yes | Yes | No |
| Modal | Yes | Yes | No |
| AVD | Yes | No | No |
| AVF | No | No | No |
| use.computer (macOS) | No | No | No |
| Local | No | No | No |
This reflects our current runner behavior.
Sandweave Ubuntu Image
Sandweave runs Linux environments with normal and fast IO. Install
gym-anything[sandweave] on Linux with Python 3.11+ and select
GYM_ANYTHING_RUNNER=sandweave. For remote workers, see
Sandweave with Weave.
The runner's template pins a prepared Ubuntu 22.04 image by release URL and SHA-256 checksums. Sandweave downloads and imports it automatically on each worker, then reuses its local cache. The QCOW2 file is read offline; workers never boot QEMU or need KVM to prepare this image.
The image uses Gym's QEMU Ubuntu provisioning, followed by
sandweave_ubuntu_setup.py. The desktop packages and configuration are retained;
native Firefox replaces Snap Firefox, and snapd is removed. No benchmark
environment or task setup is baked into the base.
An environment can adjust the template through runner_options.template, a
Sandweave template table merged over the runner's own. The runner's base image
(extends) cannot be replaced. A machine that runs Docker containers needs
cgroup v1: under cgroup v2, runc attaches eBPF device programs, which gVisor
rejects, so every container fails to start.
"runner_options": {"template": {"runtime_options": {"cgroup": "v1"}}}Machines that must reach each other join a Sandweave service network
(sandweave 0.2.24 or newer) through runner_options.service_network: the ID
returned by sandweave.create_service_network(), the machine's names on the
network, and the named networks it shares with its peers. Traffic between
members stays on the worker; everything else keeps the machine's normal network
policy. The ID belongs to one episode, so checkpoint keys ignore it.
"runner_options": {"service_network": {"id": "net-...", "aliases": ["mail.example.test"], "networks": ["office"]}}A machine can reserve less memory than its guest limit (sandweave 0.2.28 or
newer), so a worker packs more machines that rarely use their full limit at the
same time. resources.mem_gb stays the guest limit, and
runner_options.memory.reservation is the amount counted for admission.
Sandweave marks this experimental: a reservation is scheduling accounting, not a
cap on real use, and if the machines together use more than the worker's RAM,
the host can kill any sandbox in the allocation. The option therefore requires
"experimental": true. Checkpoint keys ignore the reservation, and restores
apply the current setting.
"runner_options": {"memory": {"reservation": "8GiB", "experimental": true}}Maintaining The Published Image
Only maintainers build the image. On a Linux machine with KVM, Apptainer,
fakeroot, GNU tar, sfdisk and e2fsprogs, use a local disk with at least
64 GiB of build space while keeping 15% free:
export GYM_ANYTHING_QEMU_CACHE=/scratch/gym-image-build/qemu
mkdir -p "$GYM_ANYTHING_QEMU_CACHE"
python -m gym_anything.runtime.runners.build_sandweave_ubuntu \
--source-qcow2 "$GYM_ANYTHING_QEMU_CACHE/base_ubuntu_gnome.qcow2" \
--output /scratch/gym-image-build/releaseIf the source is absent, the builder creates it with
build_base_qcow2_nodocker.py. It prepares a disposable copy, checks the package
changes, removes build-machine keys and state, and packages a new filesystem.
Build-machine host keys are omitted; the guest generates keys at first boot.
The source image is preserved.
Builds require KVM and fail immediately without it; no software-emulation fallback
is used. The output records source, recipe and artifact checksums for auditing.
Publish the ZIP, image.json, SHA256SUMS, package list and build provenance as
assets of a new versioned GitHub release. Update the template with that release's
URL and both checksums from image.json. Validate a fresh worker and the existing
Moodle and Google Earth environments before updating the runner's pin. Rebuilds
can receive newer Firefox packages from the PPA, so each published build gets a
new release rather than replacing an existing artifact.
QEMU Fast I/O
QEMU's host VNC, SSH, ADB and input-agent forwards bind to 127.0.0.1.
Apptainer direct VNC and Docker-published VNC ports also use loopback. To access
a desktop from another machine, use an SSH tunnel to its host port. Services
inside an isolated guest still accept connections from the port forwarder.
QEMU supports fast_io=True for low-latency screenshot observations and fast
keyboard/mouse input:
from gym_anything import from_config
env = from_config(
"benchmarks/cua_world/environments/google_earth_env",
task_id="take_screenshot",
fast_io=True,
)In fast mode, env.capture_observation() and env.step(...) return
obs["screen"]["image"] as a PIL image object. env.step(...) also sends
mouse and keyboard actions through persistent fast paths instead of the guest
SSH/PyAutoGUI path.
Non-fast mode keeps the file-backed obs["screen"]["path"] behavior and the
legacy input path.
Default: QMP Screendump
The default QEMU fast backend is QMP:
export GYM_ANYTHING_QEMU_FAST_IO_BACKEND=qmpThis is the stable default. It asks QEMU for a screendump, exchanges the
frame through local scratch space, and decodes it in-process.
When fast_io=True, the QEMU runner prefers local scratch under
/tmp/gym-anything-qemu-work when there is enough free space. Override this
with:
export GYM_ANYTHING_FAST_IO_WORK_DIR=/local/ssd/gym-anything-qemu-workFast Input
Fast mouse input uses QMP input-send-event. For Linux and Windows guests,
fast_io=True also starts QEMU with an explicit xHCI USB keyboard and USB
tablet. This makes the guest see standard virtual HID devices instead of relying
on QEMU's default PS/2/VMMouse path.
The runner emits absolute pointer coordinates for move, click, double-click, right-click, drag, and scroll actions. For click semantics, pointer movement, button down, and button up are sent as separate QMP commands so the guest sees ordered HID events through the virtual input devices.
Linux keyboard input uses a required guest-side uinput agent by default. The
runner starts gym-anything-fast-inputd, validates /dev/uinput, validates the
GymAnything Fast Keyboard device in the guest input stack, and then sends
keyboard actions over a persistent forwarded TCP connection. If that contract is
not available, fast_io=True fails loudly instead of falling back silently.
The previous QMP keyboard path is still available only as an explicit experimental backend:
export GYM_ANYTHING_QEMU_FAST_KEYBOARD_BACKEND=qmp-experimentalThat backend uses QMP send-key with a 5 ms hold time and a 10 ms gap between
key submissions. These values are temporary safety margins around QEMU's
asynchronous key-release scheduling; they are not the intended production
correctness contract.
The QMP timing margins can be adjusted when testing that experimental backend:
export GYM_ANYTHING_QEMU_QMP_KEY_HOLD_MS=5
export GYM_ANYTHING_QEMU_QMP_KEY_GAP_MS=10Input latency benchmarks report two timed boundaries:
runner_inject_action: starts beforerunner.inject_action(...)and ends after the fast input backend acknowledges the command sequence. For QMP mouse this includes QMP socket write/flush and response. For Linux keyboard this includes the uinput agent acknowledgement.env_step: starts beforeenv.step([action], wait_between_actions=0.0)and ends after the observation is returned. In fast mode this includes fast input dispatch and fast screenshot capture.
Use app-level evidence to prove that the target application consumed the event.
The benchmark under benchmarks/cua_world/tools/measure_input_latency.py
records the guest input device list, XInput event counts, a Google Earth search
field semantic readback, saved gedit file contents for click, double-click,
drag selection, triple-click, text, and hotkey behavior, plus contact-sheet
screenshots for right-click menus and Google Earth zoom/drag behavior.
QEMU savevm snapshots capture device state. A snapshot created before a QEMU
device-topology change can fail to load after fast input devices are added.
Regenerate savevm checkpoints when changing QEMU devices; disk-only
checkpoints do not serialize the live device state in the same way.
fast_io=True removes the legacy fixed two-second post-action sleep and the
default synchronous step-cycle wait from env.step(...). Use explicit wait
actions when a task needs app-level settling:
env.step([{"action": "wait", "time": 0.5}])You can also restore a fixed fast-mode settle delay with environment variables:
export GYM_ANYTHING_FAST_IO_ACTION_SETTLE_MS=25
export GYM_ANYTHING_FAST_IO_STEP_CYCLE_MS=50Optional: D-Bus Display
The D-Bus backend is faster because QEMU pushes display updates to a listener thread and the API returns the latest in-memory frame:
export GYM_ANYTHING_QEMU_FAST_IO_BACKEND=dbusQEMU's D-Bus display backend requires qemu-system-modules-opengl inside the
Apptainer QEMU image. If the configured image does not provide -display dbus,
Gym Anything prepares a cached sandbox automatically on first use. The sandbox
is built from GYM_ANYTHING_QEMU_CONTAINER, installs the missing QEMU display
module and dbus, and is reused from:
~/.cache/gym-anything/qemu/containers/The first D-Bus run can therefore take longer and needs network access plus Apptainer fakeroot support. Disable automatic sandbox preparation if you want the runner to fail fast instead:
export GYM_ANYTHING_QEMU_DBUS_AUTO_INSTALL=0Use a specific managed sandbox cache path with:
export GYM_ANYTHING_QEMU_DBUS_SANDBOX=/path/to/qemu-dbus-sandboxYou can also point directly at a prebuilt image:
export GYM_ANYTHING_QEMU_CONTAINER=docker://your-registry/qemu-dbus:latest
export GYM_ANYTHING_QEMU_FAST_IO_BACKEND=dbusKeep QMP as the default unless you need the lowest screenshot latency and the host can support the D-Bus runtime requirements.
Modal
GYM_ANYTHING_RUNNER=modal runs the QEMU stack inside a
Modal VM Sandbox instead of on your own machine. The
sandbox has a real Linux kernel with a working /dev/kvm, so the same Linux,
Windows, and Android guests that the QEMU runner boots run there unchanged —
this is the OS-agnostic path for hosted, cloud-side environments.
Under the hood ModalRunner starts a QemuNativeRunner inside the sandbox and
proxies to it over a small HTTP shim, so no runner logic is duplicated. Base
QCOW2 images and checkpoints live in a Modal Volume (gym-anything-qemu-cache)
and are built or uploaded once, then reused by every later sandbox.
Requirements: pip install "gym-anything[modal]" and a configured Modal token
(modal token set ...). Relevant environment variables:
GYM_ANYTHING_MODAL_APP— Modal app name (defaultgym-anything-runner)GYM_ANYTHING_MODAL_VOLUME— cache volume name (defaultgym-anything-qemu-cache)GYM_ANYTHING_MODAL_TIMEOUT— sandbox lifetime in seconds (default10800)
This runner is also the default backend when a benchmark runs as a prime-rl / verifiers environment; see Training Hubs.
Modal Native
GYM_ANYTHING_RUNNER=modal_native runs an existing Linux environment directly
inside a Modal VM Sandbox. It does not start QEMU inside the Sandbox. The runner
boots its own Ubuntu desktop image, starts systemd in a PID namespace, exposes
TigerVNC display :1, copies the existing mounts to their configured paths, and
runs the existing environment and task hooks unchanged.
CPU, memory, screen resolution, outbound networking, VNC password, and timeout
are mapped from the same EnvSpec fields used by QemuApptainerRunner wherever
Modal exposes an equivalent control. The image includes the preprovisioned ga
desktop account used by QEMU environments. Mount ro/rw modes remain copied
filesystem inputs rather than live host binds, matching QEMU behavior.
Ubuntu package-to-Snap transitions run only after the VM's systemd and snapd
are available, so existing environment installers continue to use the normal
Ubuntu packages unchanged. Modal's VM kernel does not expose squashfs, so the
runner extracts verified Snap images and presents them through read-only bind
mounts instead of leaving one snapfuse helper process per package. This keeps
application process discovery consistent with a normal Ubuntu VM. Snap apps
work functionally, but Modal's kernel has AppArmor disabled, so snapd reports
devmode rather than strict AppArmor confinement.
Filesystem checkpoint caching is supported for pre_start, post_start, and
task-specific post_task levels. Checkpoint IDs and distributed creation locks
live in a Modal Dict; each checkpoint is a Modal filesystem-snapshot Image. The
runner restarts from the new Image after creating a checkpoint so the first run
has the same disk-only semantics as later restores. Modal VM Sandboxes do not
provide VM memory snapshots, so use_savevm=True is unsupported.
fast_io=True starts a native service next to the Xvnc server. A background
XDamage listener captures changed frames through MIT-SHM into triple-buffered
RGB memory, so screenshot requests return the latest completed frame instead of
asking VNC to encode a new one. Keyboard and mouse actions are translated to one
XTest event batch and acknowledged after a single XSync. The service is
exposed through a persistent token-authenticated Modal TLS tunnel. A client
running inside the Sandbox reads changed RGB frames directly from a seqlock-
protected shared-memory triple buffer; calls made outside the Sandbox still
include WAN transfer time. Fast mode requests at least four vCPUs so the RGB
conversion can use four bounded row workers even when an environment asks for
fewer; non-fast mode continues to honor the configured CPU count and use VNC.
Requirements: pip install "gym-anything[modal_native]" and configured Modal
credentials. Relevant environment variables:
GYM_ANYTHING_MODAL_NATIVE_APP— app name (defaultgym-anything-modal-native-runner)GYM_ANYTHING_MODAL_NATIVE_CHECKPOINT_DICT— checkpoint registry nameGYM_ANYTHING_MODAL_NATIVE_TIMEOUT— Sandbox lifetime seconds (default10800)GYM_ANYTHING_MODAL_NATIVE_DESKTOP_TIMEOUT— desktop startup timeout (default300)GYM_ANYTHING_MODAL_NATIVE_SNAPSHOT_TIMEOUT— snapshot operation timeout (default600)GYM_ANYTHING_MODAL_NATIVE_SNAPSHOT_TTL— snapshot TTL seconds ornone(defaultnone)GYM_ANYTHING_MODAL_NATIVE_REGION— optional Modal regionGYM_ANYTHING_MODAL_NATIVE_INBOUND_CIDR— optional comma-separated VNC ingress CIDRs
This runner currently supports Linux CPU environments only. It fails explicitly for GPU requests and for Windows, Android, or macOS specifications. Audio and UI tree observations remain empty, matching the Linux QEMU runner.
use.computer (macOS)
runner: use_computer (set by the macos preset) runs environments on remote
macOS sandboxes from use.computer: ephemeral M4 Mac
VMs (4 cores / 8 GB) cloned from a base image, with lume as the only login
(passwordless sudo). This is the macOS counterpart to Modal — macOS guests
cannot be self-hosted on non-Apple hardware, so a vendor fleet is the hosted
path.
UseComputerRunner talks to the sandbox through the use.computer SDK:
commands run over gateway-proxied SSH (exec_ssh), screenshots come from the
SDK's screenshot API, and mouse/keyboard actions map to its input API. Spec
mounts are uploaded into the sandbox at start (read-only modes become
plain copies inside the throwaway VM). Checkpoint caching and savevm are
unsupported (the upstream API has no snapshot endpoint), and audio capture
is unavailable.
Requirements: pip install "gym-anything[use_computer]" and an API key.
Relevant environment variables:
USE_COMPUTER_API_KEY— API key (mk_live_*), minted at use.computerUSE_COMPUTER_BASE_URL— optional; defaults to the production fleet
Environments targeting this runner live in benchmarks/cua_world-macos/.