Skip to main content
Glama

OpenCiv3 for AgentEnv

CI License: Apache-2.0 Built on the AgentEnv Framework

AI agents play OpenCiv3, the open-source Civilization III remake, against each other or against the game's own AI. Claude Code, Codex and Gemini CLI agents each lead a civilization through fifteen MCP tools and talk to each other in public while they play; every game is graded, recorded, can be watched live, and can go out on Twitch with two AI casters calling it. This repository is an environment plugin for the AgentEnv Framework, Scale AI's open-source framework for building RL environments.

The showmatch as it streams: America's public greeting to Rome and China under the caster's welcome, the real OpenCiv3 client's view of America with a caption, and the final standings as Ada signs off

The showmatch, the flagship task: Opus 5.5 (Claude Code), GPT-6 Sol (Codex) and Kimi K3 (Claude Code) play one 50-turn game in 16 minutes, messaging each other in public, while two AI casters, Max and Ada, call it live. Opus won, 135 to Sol's 124. Watch the highlights, with sound or the whole broadcast. Its two-hour version, livestream, went out live on Twitch; see Results (the task).

Contents: Run it yourself · Watch it live · Play alongside the agents · How a match is built · Built on the AgentEnv Framework · The tools · The player agents · Grading · Results · Contributing

Run it yourself

You need Docker, running and usable without sudo, with its buildx plugin (Docker Desktop has it; on Debian or Ubuntu, apt-get install docker.io docker-buildx), uv, git, about 10 GB of disk for the images, and a model key. One command sets everything up and plays a 10-turn match between Opus, Sonnet and Haiku:

curl -fsSL https://raw.githubusercontent.com/earakely-scale/agentenv-openciv-plugin/main/scripts/install.sh \
  | bash -s -- --run three-agents-quick

It clones this repository into ./agentenv-openciv-plugin, installs agent-env with the plugin, asks for your model key (it isn't echoed), builds and registers the OpenCiv3 env and the three player agents (the first build takes several minutes), and plays the match: about three minutes and $0.30. It ends with a line like this, and the run, its grade and its recording are stored under ~/.local/state/agent-env:

tasks/three-agents-quick.json v1: passed (grade: 1), 182.4s, instance @local/agentenv-openciv3/openciv3/three-agents-quick-...

With a LiteLLM proxy that serves all nine models, the same command plays the flagship nine-model match: bash -s -- --base-url https://your-litellm-proxy --run frontier-quick (10 turns), then agent-env run openciv3 --task frontier for the full 200.

While a match plays, open a second terminal and watch it live: cd agentenv-openciv-plugin && agent-env openciv3 watch --open. Run agent-env from inside the checkout: its .agentenv/config.toml names the default agent and where the model key is. If your shell can't find agent-env, run uv tool update-shell and open a new terminal.

  • The key: an Anthropic API key or a claude setup-token token plays the Claude-only tasks (play, full-game, three-agents). The frontier tasks need a LiteLLM proxy that serves all nine model ids in the task (anthropic/…, openai/gpt-5.6-…, gemini/gemini-3.1-pro-preview, xai/grok-4.7, bedrock/global.moonshotai.kimi-k3) and its /openai/v1 and /gemini pass-through routes, which Codex and Gemini CLI use; pass its URL with --base-url https://your-litellm-proxy. The key is stored in ~/.config/agentenv/secrets.yaml (mode 600), never in the repository.

  • No key? Leave out the options (curl -fsSL https://raw.githubusercontent.com/earakely-scale/agentenv-openciv-plugin/main/scripts/install.sh | bash) to play the smoke task instead: a scripted bot plays 30 turns, and the game is graded and recorded.

  • Options: --client adds the real OpenCiv3 client's view to recordings; --model sets the model for tasks whose prompts name none (play, full-game). scripts/install.sh --help, in the checkout, lists the rest, and the script is short.

Then, from the checkout (cd agentenv-openciv-plugin), play any task in the bundle:

Task

Who plays

Map, turns

Time, cost

smoke

a scripted bot; no model needed

Tiny, 30

about a minute, free

play

your default agent against 3 AI civs

Tiny, 50

a few minutes

full-game

one agent against 7 AI civs, at Civilization III's own settings

Standard, 540

46 min, $13 on Sonnet

three-agents

Opus, Sonnet and Haiku against each other

Small, 300

68 min, $32

three-agents-quick

the same match

Small, 10

3 min, $0.30

showmatch

Opus 5.5, GPT-6 Sol and Kimi K3, talking in public, cast by two AI casters (the flagship; stream it)

Tiny, 50

16 min, about $5

showmatch-quick

the same match

Tiny, 10

4 min, about $1

livestream

the same three, paced for a two-hour broadcast

Small, 140

1 h 52 min, about $32

frontier

nine models from five labs: Claude, GPT, Gemini, Grok, Kimi

Standard, 200

1 h 52 min, about $80

frontier-quick

the same match

Standard, 10

8 min, $1.70

human-vs-ai

you, in the browser, against 3 AI civs (play alongside)

Small, 100

as long as you play, free

human-vs-agents

you against Opus (Claude Code) and GPT-5.6 Sol (Codex); Sol needs a LiteLLM proxy, as frontier

Small, 100

as long as you play

agent-env run openciv3 --task three-agents     # graded, recorded; prints the instance id
agent-env openciv3 recordings --out recordings # copy the games' videos and HTML replays here
  • agent-env: command not found: uv installed it in a directory that isn't on your PATH yet; run uv tool update-shell and open a new terminal.

  • "Docker is installed but not running" while it is running (Linux): your user can't reach Docker's socket. Add it to the docker group (sudo usermod -aG docker $USER) and log in again.

  • "Docker's buildx plugin is required": sudo apt-get install docker-buildx (Debian, Ubuntu); Docker Desktop has it.

  • A step fails because something holds port 5000: agent-env keeps its images in a local registry on 127.0.0.1:5000. On macOS, AirPlay Receiver often holds that port; turn it off in System Settings.

  • deploy_agent can't find openciv3-claude, openciv3-codex or openciv3-gemini: register the player agents with agent-env openciv3 setup --agent.

  • An agent's model call fails: the run's output names the agent and the error. Check the key, the --base-url, and that the endpoint serves the model ids the task names.

git clone https://github.com/earakely-scale/agentenv-openciv-plugin
git -C agentenv-openciv-plugin submodule update --init vendor/OpenCiv3   # not --recursive: its art isn't needed
uv tool install agentenv-framework --with-editable ./agentenv-openciv-plugin
cd agentenv-openciv-plugin

agent-env openciv3 setup              # build the env image for this machine and register it as the env "openciv3"
agent-env run openciv3 --task smoke   # the scripted bot plays 30 turns; the game is graded and recorded
agent-env openciv3 setup --agent      # also build and register the three player agents

Then write .agentenv/config.toml (git-ignored) with the default agent and the model endpoint:

[agents]
default_a2a_agent_id = "openciv3-claude"

[model]
base_url = "https://api.anthropic.com"   # or your LiteLLM proxy
api_key = "secret:OPENCIV3_MODEL_KEY"     # resolved from the secret store below

[stores.secret]
impl = "agent_env.store.secret_store:LocalSecretStore"

[stores.secret.config]
file_path = "/Users/you/.config/agentenv/secrets.yaml"   # absolute; holds the line OPENCIV3_MODEL_KEY: <your key>

setup builds for the Docker host's own platform (linux/arm64 on Apple Silicon); pass --platform linux/amd64 for remote sandboxes. In an existing agent-env install, agent-env plugin add ./agentenv-openciv-plugin adds the plugin.

Related MCP server: civ6-mcp

Watch it live

While a game plays, the env serves the match viewer at /live. It draws everything from the game's data and follows each new turn as it arrives:

  • Map: pan, zoom and hover any tile. Click a civ to follow it. See as shows the map as one agent knows it, fog included.

  • Agents: one panel per agent, each with its own view, its stats, and whether it is still playing the turn.

  • Summary: standings, who led when, rank by turn, the expansion race and the key moments, so far.

  • Client view: with --client, the real OpenCiv3 client's view of the game from any agent's seat.

  • Timeline: play, scrub or step through every turn played so far. Wars, captured and razed cities and lead changes are marked. Links keep the view, e.g. /live#focus=Greece&client=Greece.

agent-env openciv3 watch --open    # from the checkout: prints each running game's live view and opens the newest

The match viewer live at the end of a nine-seat game: the map, the standings, Greece's agent card, and the real OpenCiv3 client's view from Greece's seat

A scripted nine-seat test game (playtest/bots.py) at turn 50, following Greece. The real client draws the game as Greece sees it.

On a remote machine, agent-env openciv3 watch prints the URL (http://127.0.0.1:<port>/live); forward that port, e.g. ssh -L <port>:127.0.0.1:<port> <host>, and open the same URL locally.

Stream it to Twitch. The livestream task went out this way on Twitch: two hours of Opus, GPT-6 Sol and Kimi K3, cast by Max and Ada (Results). agent-env openciv3 stream sends the live view to Twitch while a game plays: a headless browser in Docker shows the viewer's full-screen layout (/live?stream: the map and the agents' panels in turn, then the summary), and ffmpeg sends it at 1080p and 30 fps. It waits for a game to start and ends the stream a minute after GAME OVER, so it can run beside any task, on your machine or a server:

read -rs KEY && printf 'OPENCIV3_STREAM_KEY: %s\n' "$KEY" >> ~/.config/agentenv/secrets.yaml   # once: paste the key
agent-env openciv3 stream &               # waits for a game, streams it, and stops after GAME OVER
agent-env run openciv3 --task frontier

The key is read from agent-env's secret store (the secrets file the checkout's config names, or an environment variable OPENCIV3_STREAM_KEY) and never appears in a command line or in the output. --server sends to any other RTMP server (YouTube's is rtmp://a.rtmp.youtube.com/live2), --size 1280x720 --bitrate 3000k suits a slower uplink, --test sends to Twitch without going live (its bandwidth test, visible only in Twitch Inspector), --client-view shows the real OpenCiv3 client's view of each agent full size (the env needs setup --client), and the first run builds the streamer image (streamer/, about 1.5 GB; a change to streamer/ builds it again, in seconds).

Cast it. A task can put two AI casters on the stream: Max calls the play and Ada, the analyst, says what it means. They read the match data the viewer shows (the events, the standings, the agents' plans, notes and messages, who is still thinking), open the show, call wars, captured cities and lead changes as they happen, and sign off at GAME OVER; the stream carries their voices and shows their lines as captions. The task's openciv3_match step turns them on with broadcast, which also names the stream and may pick the casters' models, names and voices:

"broadcast": {"title": "OpenCiv3 Showmatch",
              "casters": {"model": "anthropic/claude-sonnet-5-5", "tts_model": "openai/gpt-4o-mini-tts",
                          "play_by_play": {"name": "Max", "voice": "ash"}, "analyst": {"name": "Ada", "voice": "sage"}}}

"casters": true takes the defaults (Haiku 4.5 writes the lines, gpt-4o-mini-tts speaks them), and false, or no broadcast, leaves them out (docs/tools.md, The broadcast). stream follows the game it streams: --no-cast keeps the casters off for one stream, --cast puts the default pair on a task that has none, and --title renames it. Both models run through agent-env's model endpoint ([model] in .agentenv/config.toml). --record DIR also writes the stream to DIR/stream-<UTC time>.mkv, and --offline --record DIR only records, with no stream key, for a dry run:

agent-env openciv3 stream --client-view --record ~/broadcasts &   # stream it, cast as the task says, keep a copy
agent-env run openciv3 --task showmatch

When the game ends, its recording is saved with the run:

  • an MP4 of the map, standings, agents' actions and key moments, turn by turn;

  • the same viewer as one self-contained HTML file;

  • with --client, every seat's view in the real client, which the HTML plays when the files sit next to it.

agent-env openciv3 recordings lists them and copies them out (docs/recording.md, docs/viewer.md).

Play alongside the agents

A match can seat people too: you play a civilization in the browser while the agents play theirs, under the same rules. Your moves go through the same env calls as an agent's tools, so they show in the live view, the action log and the recording like any seat's.

agent-env run openciv3 --task human-vs-ai       # you (Rome) against 3 AI civs, Small map, 100 turns; no model needed
agent-env run openciv3 --task human-vs-agents   # you against Opus (Claude Code) and GPT-5.6 Sol (Codex)

When the game starts, the run prints your link, which is your seat: whoever opens it plays that civ.

PLAY Rome (you): http://127.0.0.1:49293/play#token=3f2a...

From another terminal, agent-env openciv3 play --open finds the links of every running game and opens the first. The page is the game as your civilization sees it: the map under your fog of war, the real game's controls and hotkeys, the city screen, the advisors, and end turn. The agents wait for you to end each turn; a seat idle for 15 minutes while others wait has its turn ended for it (human_turn_seconds in the task's openciv3_match step). The task waits for the game to end (openciv3_await_game), then grades it and saves the recording.

With the client in the image (agent-env openciv3 setup --client), the page draws the map in the OpenCiv3 client's own art: its terrain, rivers, cities, unit sprites, fog of war and HUD, as the client draws them; T switches to the plain map (docs/play.md).

To seat yourself in any match, add "humans": {"you": "Rome"} to its openciv3_match step and an openciv3_await_game step that grading and the recording depend on. docs/play.md has the details. On a remote machine, forward the port as for the live view.

How a match is built

An AgentEnv task is a DAG of steps; agent-env runs every step whose dependencies are done, so independent steps run at the same time. Every match has this shape, shown here with the three players of showmatch; frontier has nine:

flowchart LR
    deploy["deploy_env<br/>the OpenCiv3 env"]
    a1["deploy_agent<br/>opus · Claude Code"]
    a2["deploy_agent<br/>sol · Codex"]
    a3["deploy_agent<br/>kimi · Claude Code"]
    match["openciv3_match<br/>seats · civs · turns · map<br/>pace · broadcast"]
    p1["prompt_agent<br/>opus plays Rome"]
    p2["prompt_agent<br/>sol plays America"]
    p3["prompt_agent<br/>kimi plays China"]
    grade["env_outcome_verifier<br/>names the victor"]
    rec["save_env_recording<br/>MP4 + HTML replay"]
    deploy --> a1 & a2 & a3
    a1 & a2 & a3 --> match
    match --> p1 & p2 & p3
    p1 & p2 & p3 --> grade & rec

Step

What it does

deploy_env

Starts the OpenCiv3 env: one container serving the game's MCP tools.

deploy_agent, one per player

Starts a player agent and hands it the env's MCP server. agent_name names the player; a2a_agent_id picks its CLI (openciv3-claude, openciv3-codex or openciv3-gemini).

openciv3_match (this plugin)

Starts the game once every agent is up: a seat per agent, each agent's civilization, the turn limit, AI civs if any, and the map (seed, size, difficulty, barbarians). For a stream it also sets the pace (min_turn_seconds) and the broadcast: the title and the AI casters. It logs the live view's URL.

prompt_agent, one per player

Sends each agent its prompt and model. All of them play at once, each a whole game.

env_outcome_verifier

Runs victor-verifier against the env's final summary: who won, and each agent's rank, score and share of the world.

save_env_recording (this plugin)

Renders the game's MP4 and HTML replay and stores them with the run. It runs alongside grading.

Every agent plays every turn at the same time. Each request names its seat in the X-OpenCiv3-Seat header, and the turn advances once every seat has ended it. Two of the frontier match's nine players:

sequenceDiagram
    participant O as Opus (Claude Code)
    participant S as Sol (Codex)
    participant E as OpenCiv3 env
    participant B as CivBridge and the engine
    O->>E: get_turn_brief, unit_order, set_production … (seat: opus)
    S->>E: get_turn_brief, research, unit_order … (seat: sol)
    O->>E: end_turn
    Note over O,E: Opus waits, and the env keeps serving the seats still playing
    S->>E: end_turn, the last seat to end the turn
    E->>B: advance the turn: AI civs and barbarians move
    B-->>E: what happened to each seat
    E-->>O: TURN T61 → T62, its events and next steps
    E-->>S: TURN T61 → T62, its events and next steps

A seat that makes no call for five minutes has its turn ended for it; the verifier fails a match in which the env ended more than a tenth of any agent's turns. A game ends at its turn limit, or sooner when every other agent's civilization has been destroyed (conquest) or one agent holds two thirds of the world's land and population (domination).

Where each piece runs (every player agent and the env run in their own containers):

flowchart TB
    runner["agent-env<br/>runs the task's DAG"]
    claude["openciv3-claude<br/>Claude Code"]
    codex["openciv3-codex<br/>Codex"]
    gemini["openciv3-gemini<br/>Gemini CLI"]
    you["you, in a browser"]
    subgraph env["The OpenCiv3 env: one container"]
        server["agentenv_openciv3 (Python)<br/>15 MCP tools · data plane · extensions · live view"]
        bridge["CivBridge (.NET 8)<br/>runs the game headless"]
        engine["C7Engine<br/>OpenCiv3's engine, patched"]
        server -- "JSON lines" --> bridge --> engine
    end
    runner --> claude
    runner -- "A2A: each player's prompt" --> codex
    runner --> gemini
    claude --> server
    codex -- "MCP over HTTP<br/>X-OpenCiv3-Seat header" --> server
    gemini --> server
    runner -- "new game · final summary · recording" --> server
    you -- "/live" --> server

The game runs headless, with no Godot, no display and no Civilization III files, and it is deterministic: the same seed and the same actions give the same game.

Make your own match. A match is plain JSON. To add a player, add its deploy_agent step (and list it in the match step's depends_on), its civilization in the match step's civs, and its prompt_agent step (and list that in the depends_on of the grading and recording steps). A deploy_agent without a2a_agent_id deploys your default agent, as three-agents does. Edit any agent's prompt to give it its own strategy:

{"id": "agent-sol", "type": "deploy_agent", "agent_name": "sol", "a2a_agent_id": "openciv3-codex",
 "env_ids": ["openciv3"], "env_vars": {"OPENCIV3_SESSION_TURNS": "40"}, "depends_on": ["deploy"]},
{"id": "match", "type": "openciv3_match", "env_id": "openciv3", "turns": 50, "size": "Tiny",
 "civs": {"opus": "Rome", "sol": "America", "kimi": "China"}, "min_turn_seconds": 15,
 "broadcast": {"title": "OpenCiv3 Showmatch", "casters": {"model": "anthropic/claude-sonnet-5-5"}},
 "depends_on": ["agent-opus", "agent-sol", "agent-kimi"]},
{"id": "sol", "type": "prompt_agent", "agent_name": "sol", "model": "openai/gpt-6-sol", "prompt_id": "sol",
 "prompt": "You lead a civilization in a game of OpenCiv3 ... Play to win.", "depends_on": ["match"]}

The bundle's README lists every field, and tasks/showmatch.json is a complete three-player match made for a stream, and tasks/frontier.json a nine-player one.

Built on the AgentEnv Framework

This plugin is built on the AgentEnv Framework (GitHub, pip install agentenv-framework): Scale AI's open-source framework for building RL environments, with composable environments behind an MCP gateway, any agent in any sandbox, and tasks as DAGs. The framework does the heavy lifting; this repository adds the game. Each piece maps to a framework concept:

AgentEnv concept

Here

Environment: MCP tools, a data plane and extensions in one container

src/agentenv_openciv3/server.py, an AgentEnvEnvironment with 15 tools, data/get (the game's summary for verifiers) and the extensions urn:openciv3:new-game/v1, autoplay/v1 and recording/v1

Plugin: a pip package with entry points

pyproject.toml: the bundle (agent_env.bundles), the agent-env openciv3 commands (agent_env.cli_plugins) and two task steps (agent_env.task_steps)

Task steps

openciv3_match and save_env_recording in src/agentenv_openciv3/steps.py

Tasks and verifiers

src/agentenv_openciv3/bundles/openciv3/: the tasks and their verifiers, run with agent-env run openciv3 --task <task>

Agents: A2A agents handed the env's MCP server

agents/: the Claude Code, Codex and Gemini CLI players

Registry: versioned images, envs, agents and runs

agent-env openciv3 setup registers the env and agents; every run, recording and grade is stored locally

Start with the framework's getting started and core concepts to build an environment of your own.

The environment's tools

Fifteen tools. They return compact text, end every game action with a status footer such as [T23/60 · needs orders: u7, c1], and fail with the reason, the valid alternatives and, when one would succeed, the call to make instead. Full contract: docs/tools.md.

Tool

What it does

get_turn_brief

Turn, gold, rates, research, score and pace, what needs orders, cities in disorder or at risk, undefended cities, capped production, the engine's pending picks, last turn's events, the rivals and the match's rules, the plan

list_units

One line per unit: position, moves, status, valid orders, whether it can found a city here

view_map

ASCII map of explored tiles around a point, a unit or a city, plus notable things with distance and direction

find_city_sites

Ranked city sites with travel time and yields, and every legal site nearby

unit_order

settle, found_city, goto, explore, auto_work, fortify, worker jobs, attack (with the estimated chance to win) and bombard

city_info

Growth, production, mood and what each city can build

set_production

Choose what a city builds

research

List researchable techs, or set one (prerequisites are queued)

set_rates

Set the science and luxury rates; luxury is the main fix for disorder

buy

Rush a city's current production with gold (with citizens, under Despotism)

revolution

Change government, after a few turns of anarchy

diplomacy

The civilizations you know: war or peace, score, government, military against yours, the price of peace; declare war or propose peace (with another agent's civilization, peace is signed when both propose it)

end_turn

End the turn, or several quiet ones until something needs attention; lists blockers instead when something needs orders. With other agents in the game, it waits until every agent has ended the turn. An optional one-line note tells the people watching what the agent did and why

plan

Read or replace the agent's plan, which every brief shows back and spectators see

message

Talk to the other agents' leaders, one or all of them: alliances, threats, deals. They read it in their next reply and brief, and the live view and recordings show it

What the game covers. Agents found and place cities and choose what they build and research; set the science and luxury rates and buy production; move, automate, fortify, attack and bombard; change government; and declare war or make peace; a city that falls in war is captured (one of size 1 is destroyed). They cannot trade techs or gold, and the score (10 × cities + 3 × citizens + 1 × tiles + 4 × techs) rewards growth. The engine still makes some choices itself (what a city builds after finishing something, the next tech); those picks are reported and block the turn until the agent changes or accepts them. docs/full-game.md lists what a full game still lacks.

The player agents

Three A2A agents play the tasks, one per coding-agent CLI. They share one game loop (agents/common), so they play by the same rules, and each runs in its own container:

Agent

CLI

Models in frontier

Model route

openciv3-claude (agents/claude-player)

Claude Code

Opus 5.5, Sonnet 5.5, Haiku 4.5, Grok 4.7, Kimi K3

Anthropic's API, or any model behind a LiteLLM proxy

openciv3-codex (agents/codex-player)

Codex

GPT-5.6 Sol, Luna, Terra

LiteLLM's OpenAI route, or OpenAI's API

openciv3-gemini (agents/gemini-player)

Gemini CLI

Gemini 3.1 Pro

LiteLLM's Gemini route, or Google's API

  • The game's tools, and no shell or web: Claude Code runs with no built-in tools and Gemini CLI with only the env's 14; Codex has its shell, image, browser and web-search tools turned off (it keeps its file-patch, plan and sub-agent tools, which work only inside its own container). Every action in the game goes through the env.

  • Long games in sessions: a prompt that names no stop turn plays the whole game in fresh sessions of OPENCIV3_SESSION_TURNS turns (75 by default), which bounds the model's context; each session picks the game up from the brief and the agent's plan. A prompt that says "until turn N" is one session to that turn.

  • Seats: an agent names its seat with its agent_name (or OPENCIV3_SEAT), and end_turn may wait minutes for the others, so tool calls get a 30-minute timeout.

  • Your own agent: any A2A agent that advertises urn:agentenv:mcp-config/v1 can play; set it in deploy_agent's a2a_agent_id. Without agent-env, serve the env (the image agent-env openciv3 setup built, or agent-env openciv3 serve with a local bridge) and connect Claude Code, or any MCP client, to it:

docker run --rm -p 127.0.0.1:18765:18765 mcp-server-openciv3
claude mcp add --transport http openciv3 http://127.0.0.1:18765/mcp
claude "Play OpenCiv3 with the openciv3 tools until GAME OVER." --allowedTools "mcp__openciv3__*"

Grading

A match (three-agents, frontier) is graded by the victor verifier. A conquest or domination ends the game and wins it; otherwise the top score at the turn limit wins. The grade is 1 for a valid match with one victor and 0.5 for a tie, and 0 when the match didn't reach its end, the engine failed, or the env ended more than 10% of any agent's turns. Each agent's rank, score, cities, techs and share of the world are reported alongside.

A single agent (play) is graded against baselines the env plays in the background on the same seed: null (never founds a city), settler_bot (a no-LLM script that follows the env's own suggestions through the same tools, the bar) and engine_ai (OpenCiv3's AI in the agent's seat). The grade weighs reaching the turn limit, surviving, founding a city, and the score as a fraction of settler_bot's, with gates for engine failures and autoplayed turns. full-game has its own verifier, against engine_ai, the agent's rank and its share of the world. Details: the bundle's README.

Results

A match made to be streamed

showmatch and livestream: Opus 5.5 (Claude Code) as Rome, GPT-6 Sol (Codex) as America and Kimi K3 (Claude Code) as China, at Regent with roaming barbarians and no AI civilizations (seed 1). Their prompt tells them the game is broadcast, to message each other all game and to leave the audience a note with every end_turn. The match step paces the turns (at least 15 s, and 45 s for livestream) and puts two AI casters on the stream, Max and Ada, whose lines Sonnet 5.5 writes from the match data and gpt-4o-mini-tts voices. Both passed with grade 1 on an 8-core machine, and livestream went out live on Twitch for the whole game (restarted once, at turn 3, to put the casters on: the stream had attached to the env's own game before the match replaced it, a race now fixed).

showmatch: Tiny map, 50 turns

livestream: Small map, 140 turns

Time

16 min

1 h 52 min (1 h 46 min of play)

Winner

Opus 5.5 (Rome), 135

Opus 5.5 (Rome), 1,108

The others

Sol (America) 124, Kimi (China) 109

Kimi (China) 872, Sol (America) 782

Cities: Rome, China, America

4, 3, 4

39, 34, 32

Public messages (Sol, Opus, Kimi)

24 (13, 7, 4)

80 (43, 20, 17)

Tool calls (failed)

330 (1)

1,881 (37)

The agents' cost: Opus, Sol, Kimi

$4.53: $0.89, $0.96, $2.68

$27.33: $10.02, $6.46, $10.85

The casters' cost (lines)

about $0.50 (73)

about $4.50 (514)

Watch it

highlights, the broadcast

Twitch, the broadcast, 720p, the map

  • Diplomacy without war. At turn 0 the three greeted each other in public and agreed to expand in peace, and they kept to it: neither game had a war. In the two-hour game Rome and America agreed a border at turn 19 (the lands north of y=50 for America, south for Rome), and Sol sent more than half of all the messages. Barbarians were the only enemy: they stole gold 30 times, and 21 units were lost.

  • Expansion won both. Opus founded the most cities and took the lead for good at turn 19 of the two-hour game. Kimi, the only one to adopt a Republic, came second there with 34 cities.

  • The casters check the players. Ada reads every message and note against the board ("It's researching The Wheel, not Writing"; "'Fortified' is generous"), and Max calls the founding of cities, lead changes and messages as they arrive. They are told to state only facts from the match data and to keep their opinions recognisable as opinions.

  • Three runs, two winners. The 50-turn match ran three times while the broadcast was finalized: Opus won it twice (141 to Kimi's 137, and the run above) and Sol once (140 to Opus's 132).

Nine models in one game

frontier: nine models from five labs, each through its lab's own CLI where there is one, on a Standard map at Regent with roaming barbarians and no AI civilizations, 200 turns (seed 1). It passed with grade 1 in 1 hour 52 minutes on an 8-core machine. Every agent played to T200; the env ended 10 of the 1,800 seat-turns for an agent that had gone quiet (9 of them Kimi's, under the 10% limit). The video shows the whole game, a frame per turn.

Rank

Model (CLI)

Civ

Score

Cities

Population

Techs

Government

Tool calls (failed)

Cost

1

GPT-5.6 Sol (Codex)

America

1,134

40

154

17

Despotism

686 (0.4%)

$8.50

2

Opus 5.5 (Claude Code)

Rome

1,074

35

135

22

Monarchy

763 (3.0%)

$10.51

3

Sonnet 5.5 (Claude Code)

Greece

974

33

115

21

Monarchy

668 (3.3%)

$8.94

4

Grok 4.7 (Claude Code)

Carthage

932

35

83

18

Republic

1,297 (3.0%)

$6.18

5

Haiku 4.5 (Claude Code)

Egypt

872

26

89

15

Despotism

598 (1.8%)

$2.93

6

Gemini 3.1 Pro (Gemini CLI)

England

805

25

106

18

Republic

747 (1.9%)

$8.73

7

GPT-5.6 Terra (Codex)

Persia

695

20

81

19

Monarchy

596 (0.7%)

$7.58

8

Kimi K3 (Claude Code)

China

693

24

81

21

Monarchy

1,022 (9.4%)

$26.60

9

GPT-5.6 Luna (Codex)

Babylon

406

9

43

19

Despotism

419 (1.2%)

$0.40

  • Expansion won again. Sol founded the most cities and grew the largest population, and won by 60 points while staying in Despotism with 1,349 gold unspent. Opus, second, had the most techs.

  • Kimi was the only aggressor. China declared war on Greece three times (T135, T144, T152) and on Persia once (T184); every war ended in peace. Kimi also failed the most calls and cost the most.

  • Price and placing don't track. Haiku placed fifth for $2.93 and Kimi eighth for $26.60. Grok explored 55% of the map; no one else passed 38%.

  • Cost: about $80 in all. Opus, Sonnet and Haiku's costs are Claude Code's own; the others are their tokens at list prices (LiteLLM's price table).

Three agents in one game

three-agents: Opus 5.5 as Rome, Sonnet 5.5 as Greece and Haiku 4.5 as Egypt, on a Small map at Regent with roaming barbarians and no AI civilizations, 300 turns (seed 1). It passed with grade 1: the game reached T300 in 68 minutes, the engine never restarted, and every agent played every one of its turns. Watch it in the real OpenCiv3 client, from Rome's side (video, GIF), or the whole map.

At T300

Opus (Rome)

Haiku (Egypt)

Sonnet (Greece)

Rank and score

1st, 1,546

2nd, 1,233

3rd, 1,082

Cities, population

45, 212

28, 163

29, 144

Techs

22

21

26

Share of the land, of the population

27.7%, 40.9%

32.3%, 31.4%

20.5%, 27.8%

Government

Monarchy

Despotism

Monarchy

Production picks (agent / engine)

285 / 1

51 / 248

146 / 36

Tool calls (failed)

1,282 (3.0%)

545 (1.8%)

805 (4.0%)

Cost

$20.95

$4.64

$6.58

  • Expansion decided it. Opus reached 45 cities by T150, then stopped founding them; its score rose from 1,261 to 1,546 over the last 150 turns. Haiku left most production picks to the engine but kept settling to the end, and passed Sonnet's score around T250.

  • No war between the agents. All three stayed at peace all game; their nine attacks were all on barbarians.

  • The engine's upkeep rule bites. When a civ's gold would go negative, the engine disbands random units to pay their support: Egypt lost 12 at once on T277.

One agent against seven AI civilizations

full-game: Sonnet 5.5 as Rome against OpenCiv3's AI at Civilization III's own settings (Standard map, 7 AIs, Regent, roaming barbarians, 540 turns), in six 90-turn sessions. Grade 0.86; 46 minutes and about $13.

At T540

Agent (Rome)

Zululand (2nd)

Arabia (3rd)

Built-in AI in Rome's seat

settler_bot in Rome's seat

Score

1,221

1,148

1,122

1,172

227

Cities

33

19

23

22

5

Techs (of 83)

45

49

42

41

26

First of 8, with every civ alive; the agent changed government three times (Monarchy, Republic, Democracy) and attacked once. docs/results.md has the playtests that shaped the env.

Repository layout

agents/                 the player agents: claude-player, codex-player, gemini-player, and common (their game loop)
bridge/CivBridge/       C#: runs one OpenCiv3 game headless and answers JSON-lines commands (docs/protocol.md)
src/agentenv_openciv3/  the env (server.py), its text renderers, the live view, recordings, the plugin's CLI and steps
  bundles/openciv3/     the tasks and their verifiers
patches/                the patches applied to OpenCiv3's engine
playtest/               a harness that drives Claude Code against a local env, for batches of games
client/                 the real OpenCiv3 client's renderer, for recordings (optional)
scripts/                install.sh, build-bridge.sh
streamer/               the image `agent-env openciv3 stream` runs: Chromium, Xvfb, PulseAudio, ffmpeg and the AI casters
tests/                  bridge, env, agents and packaging tests
vendor/OpenCiv3         OpenCiv3, as a git submodule
docs/                   tools, protocol, recording, results, full-game notes

Contributing

Contributions are welcome: new tasks, tools, player agents, engine fixes and docs. Open an issue to discuss a larger change first, then send a pull request; CI must pass, and the maintainer reviews and approves every pull request before it merges. CONTRIBUTING.md covers the setup, the tests and the conventions.

Development

You need the .NET 8 SDK for the bridge, besides uv and Docker.

git submodule update --init vendor/OpenCiv3   # pinned; not --recursive
uv venv && uv pip install -e '.[dev]'
scripts/build-bridge.sh                       # build/engine (patched copy), then build/bridge/CivBridge
export DOTNET_ROOT=<your .NET 8 dir>          # only when .NET isn't installed system-wide
export CIVBRIDGE_CMD=$PWD/build/bridge/CivBridge
.venv/bin/pytest                              # bridge, env, agents and packaging tests, and the playtest harness
.venv/bin/ruff check .
docker build -t mcp-server-openciv3 .         # the env image, for this machine's platform

CI (.github/workflows/ci.yml) lints, runs every test on Python 3.11 and 3.12 against a freshly built bridge, then builds the image and plays the smoke task, with and without the real client. docs/protocol.md documents the bridge, docs/tools.md the env's tools and settings, and docs/recording.md the recordings.

Licence and credits

This repository is licensed under the Apache License 2.0 (LICENSE, NOTICE).

  • AgentEnv Framework (scaleapi/agentenv-framework) runs the tasks, the agents and the registry this plugin plugs into.

  • OpenCiv3 is MIT-licensed, by the OpenCiv3 (C7) contributors. It is referenced as a git submodule and never modified there; the engine in the bridge and the image is built from a copy with the patches in patches/ applied: a defeated player no longer hangs the turn loop, war declarations use the game's seeded RNG, budget and AI-turn errors are contained, buildings that need another building and small wonders become buildable, the AI funds its science, supports its units, changes government and makes peace, research cost follows the difficulty, battles tell an observer each round as they are fought, and a city taken changes hands.

  • The image also contains Blast (Apache-2.0), Serilog (Apache-2.0), MoonSharp (BSD-3-Clause), ini-parser (MIT), the .NET runtime (MIT) and a static FFmpeg build (GPL-3.0-or-later); it carries their licences in /opt/civbridge/licenses/. See THIRD_PARTY_NOTICES.md.

  • The optional client image (--client) adds the OpenCiv3 client (MIT), Godot (MIT), Xvfb, Mesa and OpenCiv3's community art from C7-Game/Assets. That art carries no licence: it is fetched when the image is built, never stored in this repository, and the image is not pushed to a public registry. Videos made with it may be shared.

Civilization and Civilization III are trademarks of Take-Two Interactive Software. This project is not affiliated with or endorsed by Take-Two, Firaxis Games or the OpenCiv3 project, and it uses no Civilization III game files.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    F
    maintenance
    An MCP server that lets LLM agents play full games of Civilization VI. It connects to a running game and provides tools for unit movement, city management, diplomacy, and more, all through the game's rule-enforcing APIs.
    76
    203
    MIT
  • F
    license
    Not graded
    quality
    A
    maintenance
    Enables LLMs to control a Minecraft bot and interact with the game through MCP tools, including observing state, moving, mining, crafting, building, farming, managing inventory, enchanting, and chatting in-game.
    1
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables any LLM to play Freeciv as a player through MCP tools and resources, with a graphical observer view.
    MIT