Gym Anything
Reference Agents

Agents

The reference agents and programs that run them against Gym Anything environments.

The agents/ folder contains reference agents and the programs used to run them.

You don't need this part of the repository to use Gym Anything itself.

You only need it if:

  • you want to run one of the reference agents
  • you want a starting point for your own agent
  • you want to evaluate an agent across benchmark tasks

The Two Main Parts

Most of this section comes down to two folders:

What Lives In agents/agents/

agents/agents/ contains the agent classes.

An agent is the code that looks at the current observation and decides what actions to take next.

We currently ship these agents:

  • ClaudeAgent
  • Gemini3Agent
  • GeminiComputerUseAgent, GPT54ComputerUseAgent (provider-native computer-use tools)
  • Qwen3VLAgent
  • KimiAzureAgent
  • ClaudeCodeAgent, CodexCliAgent (external CLI harnesses, see below)

These classes are exported through agents.agents, which is why the evaluation commands refer to names such as --agent ClaudeAgent.

What Lives In agents/evaluation/

agents/evaluation/ contains the programs that actually run an agent against environments and tasks.

The two main ones are:

  • run_single.py: run one agent on one task
  • run_batch.py: run one agent across many tasks

If you only want to try an agent once, start with run_single.py.

What An Agent Receives

Our current agent interface is built around four methods:

  • __init__
  • init(task_description, display_resolution, save_path)
  • step(obs, action_outputs)
  • finish(...)

The important one is step(...).

That method receives:

  • the latest observation from the environment
  • the outputs from the previous actions

and it returns one or more action groups.

Each action group contains:

  • tool_id
  • actions

The actions list can contain normal environment actions such as mouse and keyboard input. It can also contain built-in control actions such as:

  • {"action": "screenshot"}
  • {"action": "wait", "time": 1.5}

Those control actions are handled by the environment layer, not by the evaluation loop.

Autonomous Agents (External CLI Harnesses)

Some agents are not step-based. ClaudeCodeAgent and CodexCliAgent wrap existing coding CLIs (Claude Code, Codex) and let the CLI run the whole episode on its own. These set a class flag autonomous = True and implement run_episode(env, task_description) instead of step(...). run_single.py detects the flag, calls run_episode once, and then runs the same mark_done verification as every other agent.

The CLI never runs inside the task VM. It runs in a throwaway, GUI-less scratch container that can reach exactly one thing on the host: an in-process action gateway. The CLI sends one action as a JSON string (the same vocabulary the other agents use) through a small act command, the gateway calls env.step() and returns the post-action screenshot plus the remaining step budget. So the CLI can only affect the environment through the gateway, and it can never bypass the mouse-and-keyboard action space by editing files or the database directly. The shared machinery lives in agents/shared/cli_harness.py.

The sandbox is backend-agnostic and auto-selected the way env runners are (GYM_ANYTHING_AGENT_SANDBOX override, then detect), in agents/shared/agent_sandbox.py: Apptainer (rootless, no daemon, the primary path on clusters) or Docker (daemon machines, which additionally get bridge network isolation). There is deliberately no no-isolation fallback: if neither is available the run fails rather than launch a permission-bypassed CLI unsandboxed. File and process isolation hold on both backends, and the env is a separate VM either way, so the agent has no filesystem path into it. On Docker the bridge also blocks any route to the env's ports; under rootless Apptainer (shared host network) that guarantee softens to the agent only ever being handed the gateway and never the env's SSH/VNC credentials.

These agents need Apptainer or Docker on the host and the relevant API key (ANTHROPIC_API_KEY for Claude Code, OPENAI_API_KEY for Codex), which is injected into the sandbox per run, never baked into the image.

The Fastest Way To Try One

gym-anything benchmark moodle --task enroll_student --agent ClaudeAgent --model claude-opus-4

That command loads the environment, resets it, creates the agent, lets it act until the run finishes, and writes run artifacts.

Run gym-anything agents to see all available agent names.

If You Want To Add Your Own Agent

The simplest workflow is:

  1. copy a nearby file in agents/agents/
  2. implement the same basic methods
  3. export the new class from agents/agents/__init__.py
  4. run it with gym-anything benchmark

If you're starting from scratch, read:

  1. agents/agents/base.py
  2. one concrete agent such as agents/agents/claude.py
  3. agents/evaluation/run_single.py

That gives you the agent interface first, then one implementation, then the program that drives it.

On this page