Gym Anything
Reference Agents

Agents

The reference agents and programs that run them against Gym Anything environments.

The agents/ folder contains reference agents and the programs used to run them.

You don't need this part of the repository to use Gym Anything itself.

You only need it if:

  • you want to run one of the reference agents
  • you want a starting point for your own agent
  • you want to evaluate an agent across benchmark tasks

The Two Main Parts

Most of this section comes down to two folders:

What Lives In agents/agents/

agents/agents/ contains the agent classes.

An agent is the code that looks at the current observation and decides what actions to take next.

We currently ship these agents:

  • ClaudeAgent
  • Gemini3Agent
  • GeminiComputerUseAgent, GPT54ComputerUseAgent (provider-native computer-use tools)
  • Qwen3VLAgent
  • KimiAzureAgent
  • ClaudeCodeAgent, CodexCliAgent (external CLI harnesses, see below)

These classes are exported through agents.agents, which is why the evaluation commands refer to names such as --agent ClaudeAgent.

What Lives In agents/evaluation/

agents/evaluation/ contains the programs that actually run an agent against environments and tasks.

The two main ones are:

  • run_single.py: run one agent on one task
  • run_batch.py: run one agent across many tasks

If you only want to try an agent once, start with run_single.py.

What An Agent Receives

Our current agent interface is built around four methods:

  • __init__
  • init(task_description, display_resolution, save_path)
  • step(obs, action_outputs)
  • finish(...)

The important one is step(...).

That method receives:

  • the latest observation from the environment
  • the outputs from the previous actions

and it returns one or more action groups.

Each action group contains:

  • tool_id
  • actions

The actions list can contain normal environment actions such as mouse and keyboard input. It can also contain built-in control actions such as:

  • {"action": "screenshot"}
  • {"action": "wait", "time": 1.5}

Those control actions are handled by the environment layer, not by the evaluation loop.

GeminiComputerUseAgent accepts either the normal obs["screen"]["path"] or a chronological obs["frames"] list. When several frames are present, every frame is sent in order in the initial request and in later computer function responses. Its request deadline and total retry attempts can be set with request_timeout_seconds and request_attempts in agent_args.

Autonomous Agents (External CLI Harnesses)

Some agents are not step-based. ClaudeCodeAgent and CodexCliAgent wrap existing coding CLIs (Claude Code, Codex) and let the CLI run the whole episode on its own. These set a class flag autonomous = True and implement run_episode(env, task_description) instead of step(...). run_single.py detects the flag, calls run_episode once, and then runs the same mark_done verification as every other agent.

The CLI never runs inside the task VM. It runs in a throwaway, GUI-less scratch container that can reach exactly one thing on the host: an in-process action gateway. The CLI sends one action as a JSON string (the same vocabulary the other agents use) through a small act command, the gateway calls env.step() and returns the post-action screenshot plus the remaining step budget. So the CLI can only affect the environment through the gateway, and it can never bypass the mouse-and-keyboard action space by editing files or the database directly. The shared machinery lives in agents/shared/cli_harness.py.

The sandbox is backend-agnostic and auto-selected the way env runners are (GYM_ANYTHING_AGENT_SANDBOX override, then detect), in agents/shared/agent_sandbox.py: Apptainer (rootless, no daemon, the primary path on clusters) or Docker (daemon machines, which additionally get bridge network isolation). There is deliberately no no-isolation fallback: if neither is available the run fails rather than launch a permission-bypassed CLI unsandboxed. File and process isolation hold on both backends, and the env is a separate VM either way, so the agent has no filesystem path into it. On Docker the bridge also blocks any route to the env's ports; under rootless Apptainer (shared host network) that guarantee softens to the agent only ever being handed the gateway and never the env's SSH/VNC credentials.

These agents need Apptainer or Docker on the host and the relevant API key (ANTHROPIC_API_KEY for Claude Code, OPENAI_API_KEY for Codex), which is injected into the sandbox per run, never baked into the image.

The Fastest Way To Try One

gym-anything benchmark moodle --task enroll_student --agent ClaudeAgent --model claude-opus-4

That command loads the environment, resets it, creates the agent, lets it act until the run finishes, and writes run artifacts.

Run gym-anything agents to see all available agent names.

If You Want To Add Your Own Agent

The simplest workflow is:

  1. copy a nearby file in agents/agents/
  2. implement the same basic methods
  3. export the new class from agents/agents/__init__.py
  4. run it with gym-anything benchmark

If you're starting from scratch, read:

  1. agents/agents/base.py
  2. one concrete agent such as agents/agents/claude.py
  3. agents/evaluation/run_single.py

That gives you the agent interface first, then one implementation, then the program that drives it.

On this page