Agents
The reference agents and programs that run them against Gym Anything environments.
The agents/ folder contains reference agents and the programs used to run them.
You don't need this part of the repository to use Gym Anything itself.
You only need it if:
- you want to run one of the reference agents
- you want a starting point for your own agent
- you want to evaluate an agent across benchmark tasks
The Two Main Parts
Most of this section comes down to two folders:
Agents
Agent classes that decide what actions to take
Evaluation
Programs that run agents against environments
What Lives In agents/agents/
agents/agents/ contains the agent classes.
An agent is the code that looks at the current observation and decides what actions to take next.
We currently ship these agents:
ClaudeAgentGemini3AgentGeminiComputerUseAgent,GPT54ComputerUseAgent(provider-native computer-use tools)Qwen3VLAgentKimiAzureAgentClaudeCodeAgent,CodexCliAgent(external CLI harnesses, see below)
These classes are exported through agents.agents, which is why the evaluation commands refer to names such as --agent ClaudeAgent.
What Lives In agents/evaluation/
agents/evaluation/ contains the programs that actually run an agent against environments and tasks.
The two main ones are:
run_single.py: run one agent on one taskrun_batch.py: run one agent across many tasks
If you only want to try an agent once, start with run_single.py.
What An Agent Receives
Our current agent interface is built around four methods:
__init__init(task_description, display_resolution, save_path)step(obs, action_outputs)finish(...)
The important one is step(...).
That method receives:
- the latest observation from the environment
- the outputs from the previous actions
and it returns one or more action groups.
Each action group contains:
tool_idactions
The actions list can contain normal environment actions such as mouse and keyboard input. It can also contain built-in control actions such as:
{"action": "screenshot"}{"action": "wait", "time": 1.5}
Those control actions are handled by the environment layer, not by the evaluation loop.
GeminiComputerUseAgent accepts either the normal obs["screen"]["path"] or
a chronological obs["frames"] list. When several frames are present, every
frame is sent in order in the initial request and in later computer function
responses. Its request deadline and total retry attempts can be set with
request_timeout_seconds and request_attempts in agent_args.
Autonomous Agents (External CLI Harnesses)
Some agents are not step-based. ClaudeCodeAgent and CodexCliAgent wrap
existing coding CLIs (Claude Code, Codex) and let the CLI run the whole
episode on its own. These set a class flag autonomous = True and implement
run_episode(env, task_description) instead of step(...). run_single.py
detects the flag, calls run_episode once, and then runs the same mark_done
verification as every other agent.
The CLI never runs inside the task VM. It runs in a throwaway, GUI-less
scratch container that can reach exactly one thing on the host: an in-process
action gateway. The CLI sends one action as a JSON string (the same
vocabulary the other agents use) through a small act command, the gateway
calls env.step() and returns the post-action screenshot plus the remaining
step budget. So the CLI can only affect the environment through the gateway,
and it can never bypass the mouse-and-keyboard action space by editing files
or the database directly. The shared machinery lives in
agents/shared/cli_harness.py.
The sandbox is backend-agnostic and auto-selected the way env runners are
(GYM_ANYTHING_AGENT_SANDBOX override, then detect), in
agents/shared/agent_sandbox.py: Apptainer (rootless, no daemon, the
primary path on clusters) or Docker (daemon machines, which additionally
get bridge network isolation). There is deliberately no no-isolation fallback:
if neither is available the run fails rather than launch a permission-bypassed
CLI unsandboxed. File and process isolation hold on both backends, and the env
is a separate VM either way, so the agent has no filesystem path into it. On
Docker the bridge also blocks any route to the env's ports; under rootless
Apptainer (shared host network) that guarantee softens to the agent only ever
being handed the gateway and never the env's SSH/VNC credentials.
These agents need Apptainer or Docker on the host and the relevant API key
(ANTHROPIC_API_KEY for Claude Code, OPENAI_API_KEY for Codex), which is
injected into the sandbox per run, never baked into the image.
The Fastest Way To Try One
gym-anything benchmark moodle --task enroll_student --agent ClaudeAgent --model claude-opus-4That command loads the environment, resets it, creates the agent, lets it act until the run finishes, and writes run artifacts.
Run gym-anything agents to see all available agent names.
If You Want To Add Your Own Agent
The simplest workflow is:
- copy a nearby file in
agents/agents/ - implement the same basic methods
- export the new class from
agents/agents/__init__.py - run it with
gym-anything benchmark
If you're starting from scratch, read:
agents/agents/base.py- one concrete agent such as
agents/agents/claude.py agents/evaluation/run_single.py
That gives you the agent interface first, then one implementation, then the program that drives it.