najamjad-cop
Provides Gmail-API result reporting, enabling the agent to email match reports to the lecturer or a practice-mode recipient after a match.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@najamjad-copRun preflight checks and verify the replay of our last match."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Pursuit League - Cop Agent 👮
The police/cop side of a two-person final project for Orchestration of AI Agents: a distributed cops-and-thieves game played peer-to-peer over MCP (FastMCP/HTTP) against other teams' agents, with SHA-256 commit-reveal integrity and Gmail-API result reporting.
Companion repository (thief agent): https://github.com/najikay/pursuit-thief-agent The shared core package is byte-identical across both repos, enforced by
scripts/sync_core.pyin CI (seedocs/PLAN.md, ADR-002).
Team: Naji Kayal · Amjad Abed
Status: M6 - six counted series played and filed against six distinct opponents,
rule 31's pass threshold met three times over. Plays a full audited series, interoperates
with the course reference simulator and with six independent team implementations, and
audits an opponent's play after the match. Remaining work is tracked in docs/TODO.md.
League record
# | Date | Opponent | Result | Us | Them |
1 | 2026-08-08 |
| won 6–0 | 90 | 30 |
2 | 2026-08-13 |
| won 6–0 | 90 | 30 |
3 | 2026-08-14 |
| lost 0–6 | 30 | 90 |
4 | 2026-08-17 |
| won 4–2 | 60 | 40 |
5 | 2026-08-18 |
| lost 0–6 | 30 | 90 |
6 | 2026-08-21 |
| tied 3–3 | 75 | 75 |
Thirty-six of thirty-six mini-games verified Verified OK at the audit, across six
counted series, with zero technical losses attributable to us. That is the number we would point
at first: every game we played was one both sides could re-hash and agree on, including the
two we lost badly.
Two series reconciled field for field against the opponent's own filed report - vibecode's
at 66 fields with zero differences, and a later friendly against anrbj666 matching on all
six sub-games, both mutual_agreement.sha256 values, and both per-window github_commit
pairs. A report that agrees with the opponent's is the only kind that cannot be voided under
rules 33-35, and it is worth more to us than a scoreline.
Contents
Installation · Command line · Running it · Match-day workflow
Academic report: the model · orchestration dilemmas · strategies · what testing against a stranger taught us
Related MCP server: Police MCP Server
Installation
Requirements
Python | 3.12+ (managed by uv - you do not need it pre-installed) |
the only package manager used here (guidelines §8.4) | |
OS | Linux, macOS, or Windows via WSL2 - developed on WSL2 |
Optional |
|
git clone https://github.com/najikay/pursuit-cop-agent.git
cd pursuit-cop-agent
uv sync # installs the locked dependency set
uv run python scripts/check_all.py # every CI gate, one PASS/FAIL verdictcheck_all.py passing means the install is sound: lint, file-size limits, repo
rules, type checking and the full test suite.
Secrets - copy .env-example to .env and fill in real values. Nothing secret is ever
committed; .gitignore covers .env, secrets/, token.json and credentials.json, and a
CI gate fails the build if one is ever tracked (book rules 39-40).
cp .env-example .env # then edit: ANTHROPIC_API_KEY, DEEPSEEK_API_KEY, …Troubleshooting
Symptom | Cause and fix |
| Another agent is running. Stop it, or change |
| Normal: importing the MCP stack. It prints |
Opponent reports us unreachable | Check the tunnel: |
| Run from the repository root; the console script lives in this repo's |
OAuth browser never opens (WSL) | Expected - WSL has no default browser. Use |
Everything is slow on | Windows-mounted filesystems are slow in WSL. The event log holds its handle open for exactly this reason; if you can, keep the workspace on the Linux filesystem. |
Command line
One console script per repo (najamjad-cop here, najamjad-thief in the companion). Every
verb is argument parsing plus a single SDK call - the CLI holds no game logic, and a
meta-test keeps it that way.
uv run najamjad-cop --help # every verb
uv run najamjad-cop version # code version (book rule 53)
uv run najamjad-cop preflight # match-day checklist
uv run najamjad-cop peer # serve: MCP server + tunnel + dashboard
uv run najamjad-cop match # serve, then play the agreed series
uv run najamjad-cop peer --no-tunnel --no-dashboard # local play, nothing exposed
# Re-hash every step of a log and print the verdict. Paths are literal -
# `<log>` would be read by the shell as a redirect, so use a real one:
uv run najamjad-cop replay tests/goldens/artifacts/log_segal-police-team-vs-segal-thief-team_g01.json
uv run najamjad-cop replay path/to/log.json --serve # open the viewer instead
uv run najamjad-cop archive match.zip # bundle evidence (secrets excluded)Exit codes, because these run in scripts:
Code | Meaning | Example |
| it worked | preflight ready, log verified |
| it ran, the answer was bad | not match-ready, log TAMPERED |
| it could not run | log file missing or unreadable |
A tampered log and a missing file are deliberately different codes: an audit result must never be mistaken for a typo.
Development commands
uv run python scripts/check_all.py # all CI gates, one verdict
uv run python scripts/self_play.py --games 100 # measure our brains vs baselines
uv run python scripts/demo_dashboard.py # dashboard over a played game
uv run python scripts/two_process_match.py # both repos as real processes
uv run python scripts/sync_core.py ../pursuit-thief-agent # verify the mirrored coreRunning it
Everything below works from a fresh clone with no opponent and no API keys. The agent plays a full audited series on templates alone (book PAGE 67) - the LLM is an enhancement, not a dependency.
1. Install and prove the install
git clone https://github.com/najikay/pursuit-cop-agent.git
cd pursuit-cop-agent
uv sync # locked dependency set
uv run python scripts/check_all.py # every CI gate, one verdictALL GATES PASSED means lint, file sizes, repo rules, types, the strategy
suite and 1,900+ tests are green.
2. Serve as a peer, with the dashboard
uv run najamjad-cop peer --dashboard --no-tunnelWait for Uvicorn running - that, not the first log line, is readiness. Cold
start is ~15 s (importing the MCP stack), so start early on match day.
Then open http://127.0.0.1:8000/.
The port and host come from config/setup.json (ui.port, ui.host). It binds
to loopback by design: the dashboard shows our belief and our sealed state,
so exposing it would hand an opponent everything commit-reveal exists to hide
(rules 8-9).
What you get:
Panel | Shows |
Board | Belief heatmap on a log scale - linear collapsed 47 of 48 cells into one band |
Turn banner | Whose move, which step, which phase |
Dialogue | Every hint in and out, with the model that wrote it |
Negotiation | Propose → counter → lock, and terms awaiting a human |
Budget | Tokens against the agreed 200k cap |
Match day | Readiness - the same checks |
Testing | Practice mode, and whether each endpoint is answering |
Matches | Every match we have filed: score, per-game audit verdicts, artifacts |
Events | The raw event stream |
Updates arrive over a WebSocket; the client never polls. A dashboard failure cannot affect a game - it is a subscriber and nothing more (ADR-005).
Optional controls. Set features.controls to true in config/setup.json
to enable start/stop, negotiation approval, and the practice-mode toggle from
the page. Off by default, and there is deliberately no button that plays a
counted match - that is graded and irreversible, and docs/RUNBOOK.md is the
interface for it.
2b. Testing without touching the lecturer
The Testing panel answers the two questions that otherwise mean digging through logs.
Which mode am I in? Practice mode redirects every report to your own inbox
instead of the lecturer's, and prefixes the subject [PRACTICE]. It does not
skip the send - sending is the step that lost matches in assignment 6, so a
practice run exercises it for real and you read the actual email. The redirect
is enforced twice: the address is rewritten, and then checked at the point of no
return, so a rewrite that silently failed raises instead of delivering. See
docs/CONFIG.md §3b.
Simplest is the flag - it arms practice for one process and touches nothing:
uv run najamjad-cop match --opponent amjad --practice --dashboard --tunnelOr toggle it from the panel (with features.controls on), or set
practice.enabled in config/setup.json to make it persist. It is read fresh each time a report
is built, so the switch takes effect without a restart, and it overrides
email.mode to send - a practice run that quietly produced a draft would
look exactly like a successful send.
Is anyone actually reachable? Probe endpoints dials our MCP URL and the opponent's and reports three states, not two:
State | Meaning |
green | something accepted a TCP connection there |
red | configured, but nothing is listening - this blocks a match |
grey | not configured yet - a match not scheduled, not a fault |
A green light means the port answered. It does not mean the protocol works or that they will agree to our terms - that is what the handshake is for, and the panel deliberately claims no more than it can check.
3. Check you are ready to play
uv run najamjad-cop preflight # exit 0 or do not play
uv run python scripts/pre_match_smoke.py # MATCH READY in ~50 s4. Play
Write an opponent card once, then name it. opponents/<name>.toml carries the
two facts that come from them - their MCP endpoint and the group_id their
handshake declares:
url = "https://their-agent.example.com/mcp"
group_id = "their-group"
name = "Their Team"
notes = "quick tunnel - URL changes if cloudflared restarts"Copy opponents/_template.toml to start. Then every verb takes --opponent:
uv run najamjad-cop preflight --opponent amjad
uv run najamjad-cop match --opponent amjad --dashboard --no-tunnelNothing tracked needs editing, the settings for the last opponent are not
overwritten by the next, and opponent_group_id - which must equal what their
handshake declares, or the report is keyed by a placeholder instead of their
name - is reviewable the day before instead of discoverable at the handshake.
A card can only set network.opponent_*. Game terms are agreed in the signed
config/game.json, and a per-opponent override of one is exactly the thing
that must never be easy.
Or set network.opponent_url in config/police/game.toml by hand, then:
uv run najamjad-cop match --dashboard --no-tunnelA finished match writes four artifacts per Appendix F into
workspace/artifacts/ - declaration, config, log and result - and emails the
result. Check them with:
uv run python scripts/post_match.py --opponent <name>5. Try it without an opponent
Two ways, both real:
# our cop against our thief, two OS processes over real MCP/HTTP
uv run python scripts/two_process_match.py
# against the course reference simulator (expects ../reference-sim)
uv run python scripts/rehearsal.py --games 6The second is the one that matters - it is the only setup that has ever caught our interop defects, because it is the only opponent we did not write.
6. Verify a log
uv run najamjad-cop replay workspace/artifacts/log_<game_id>_g01.jsonExit 0 is Verified OK; exit 1 is TAMPERED and names the failing step;
exit 2 means the file could not be read. A tampered log and a typo are
deliberately different codes - an audit verdict must never be mistaken for a
mistyped path.
Add --serve to open the viewer instead of printing a verdict.
7. Measure
uv run python scripts/strategy_smoke.py --games 50 # win rates with intervals
uv run python scripts/sweep.py --games 24 # parameter sensitivity
uv run python scripts/measure_tokens.py # token censusResults land in results/ and are what notebooks/analysis.ipynb plots.
Match-day workflow
The full procedure with exact commands is docs/RUNBOOK.md. In outline:
Warm up - start the agent early; cold start is ~15 s.
Preflight -
uv run najamjad-cop preflight; exit 0 or do not play.Exchange URLs - set
network.opponent_urlinconfig/police/game.toml.Play -
uv run najamjad-cop match, dashboard on http://127.0.0.1:8000/.Audit - automatic per mini-game; every game must read
Verified OK.Report - reconcile with the opponent, then send (rule 30:
gmail.sendonly).Archive -
uv run najamjad-cop archive match.zip(secrets excluded).
Configuration
File | Role |
| Shared, signed terms. Both peers must hold a byte-identical copy; the handshake refuses to play on any mismatch. Ours is the opening proposal - every value at or above the Appendix F minimum (rule 12: raise, never lower). |
| Private, local. Our port, opponent URL, tunnel hostname, LLM choices, belief tuning. Never crosses the network. |
| Per-service limiter settings, validated against the Appendix F ceilings at load. |
| Secrets only. Never committed. |
Parameters worth knowing:
Key | Effect |
| Our MCP port (8802 cop / 8801 thief, so both run locally). |
| The only thing we know about the opponent. Preflight fails while empty. |
| Permanent public name. A named tunnel keeps its URL across restarts - the defect that cost Assignment 6 the most time (ADR-004). |
| Hint cadence. A quality dial, not a savings dial - see |
| How far we trust scent against a possibly-lying hint. |
| Agreed rules. Changing these unilaterally breaks the signature. |
The dashboard
Start it with --dashboard and open http://127.0.0.1:8000/ - see
Running it §2 for the panels and the optional controls.
Belief heatmap, turn banner, dialogue with per-message model provenance, the negotiation timeline, token budget, and report delivery status - pushed over a WebSocket, never polled. It shows local truth only (book rules 8-9): the opponent's position has no field in the read model, and a meta-test enforces that the UI can reach the agent only through the SDK.

Replay viewer
Every step is re-hashed from its revealed (payload, nonce) and compared with the stored
commitment (book rule 20). Below, the lecturer's own sample log replaying clean:

And the same log with one record edited after the fact - the forgery is localised to exactly the step it was planted in, and rule 19 voids the game:

See assets/README.md for how each image is reproduced.
Academic report
The model: a Dec-POMDP neither side can see
The game is a decentralised, partially observable Markov decision process. Neither agent observes the true state: positions are sealed inside commitments until the end-of-game audit, so each peer holds a belief over where the other might be and acts on that.
Two observation channels, with opposite trust properties:
Scent - a decaying pheromone trail the opponent emits involuntarily and cannot fake (book PAGE 22). Unfakeable but blurry.
Two models ship, and either can be selected per match. The book (PAGE 43-44) is radial - 0.90 / 0.62 / 0.42 / 0.20 / 0.14 / 0.04 - with relative decay
τ ← (1-ρ)·τ; the reference simulator is linear in Chebyshev distance - rings 0.90 / 0.60 / 0.30 - with absolute decayτ ← τ - ρ. We implement both, and both are registered in the interop kit:ScentModel.BOOKismultiplicative_book_v1(934c220d…) andScentModel.REFERENCEissubtractive_chebyshev_v1(81ebee59…). Each reproduces the kit's own published vectors - including the wire's serve order, which the kit'sfield_walkfixes as age the prior trail, merge the fresh deposit undecayed, transmit that: our thief shipped the trail one decay step too fresh until 2026-08-22, an opponent's per-frame gate measured it, and the frames are now pinned against the walk cell-for-cell (test_wire_scent_serve_order.py). So matching an opponent is one key -pheromones.pheromone_modelinconfig/game.json- and not a change to the fourteen signed terms, so the contract digesta284082d…survives the switch and nobody has to re-sign. The digest we declare at negotiate is looked up from the configured model, so there is no state in which we emit one physics and claim another. Emission is separately dialled from hints ---scent full|window|noneand--hints/--no-hints- so a fully silent series is one flag; under silence we declare no model at all, because a claim about a field nobody is sending is not a claim worth making.Hints - free natural language, which the rules explicitly permit to be a lie (rules 26-27). Precise but untrustworthy.
Our belief engine fuses them: diffusion for movement, a scent-likelihood update, and a
credibility weight per opponent that rises and falls as their hints agree or disagree with the
trail. The full derivation is in docs/PRD_belief_engine.md.
Three findings from building it, each of which changed the design:
Scent decay had a fixed point. Relative decay rounded to three decimals never reached zero, so dead trails polluted belief forever. Fixed with an explicit epsilon.
Multiplicative fusion double-counted the cumulative scent field and left belief lagging a moving opponent by ~5 cells. Replaced with a robust mixture update.
A flat likelihood parked belief mid-trail. Sharpening plus an additive floor fixed it; a clamp made faint scent indistinguishable from none.
Orchestration dilemmas
One gateway, or many callers? Book rule 3 mandates an orchestrator, and we took it literally: peripheral modules never call each other. The belief engine knows nothing of the transport, strategy knows nothing of crypto, the transport knows nothing of the rules. That is what makes the whole turn loop testable against fakes - the seam Assignment 6 never had.
How much to trust a fake. Our sharpest lesson. Three separate defects survived 1,500 passing tests because the fakes were kinder than the wire: a fake transport that wrapped an audit payload the real one did not, an in-memory link that never blocked, and peers that only ever played themselves. Interop can only be tested against something you did not write.
Where to put the deadline. A peer that answers forever must not hold us in a decided game, and a peer that goes quiet must not become our technical loss. Every wait is bounded and every ending is announced rather than assumed - see below.
Rate limiting our own protocol. The gatekeeper exists to be a good citizen towards Anthropic and Gmail. Applying it to the opponent nearly cost us games: our own limiter could have delayed a reply past their 30-second deadline.
Strategies, and why not RL
Cop - a scripted halving seal (strategy/seal_cop.py), not per-turn scoring: wall the
middle column while walking the lane beside it, take the gate, cut the half that holds the
thief, and the thief ends in a 3×3 pocket with us. From there domain/endgame.py solves the
pocket exactly - barrier placement is a move in its game tree - and prices every win to
co-location, because a step-on claim is the one capture every implementation honours; a
thief sealed where we can never walk is scored remote seal by our own bench and refused as
a plan. The one-barrier lock (strategy/lock.py: our body plus one wall, adjacency
required) finishes a pinned thief early. The bench (tests/regression/) drives it against
reacting adversaries - including a thief that squats the script's own wall line - and it
converts all of them inside the 35 moves, by plain claimable capture
(test_seal_converts_every_reacting_thief.py). Belief-directed pursuit (cop_brain.py)
remains the fallback for a board with no localisation.
Thief - survival-horizon evasion as a stack of hard floors, not greedy distance
maximisation: never end a turn within one step of the cop, refuse cells a half-built fence
has already made cheap to seal, keep a reachable 4×4, stay on the cop's side of a forming
cut (a sealer may not complete a wall that separates it from us), and stand in the gaps of
a fence under construction - a preference that runs after the safety floors, because run
before them it once marched us into the corner an opponent then walled
(test_ahk_yosi_corner_hunt.py replays that kill line move for move). It survives every
archived opponent line in matches/ and all 24 tie-break variants of a reacting
corner-hunter cop.
Hint policy - bluffing is a strategic resource with a cost: a claim discloses the claimer's cell, so a false capture claim hands the thief our exact position for nothing.
No reinforcement learning, deliberately:
Only ~60 real games are available across the whole league - nowhere near enough to learn a policy over a state space this size.
Opponent behaviour is non-stationary; every team ships something different.
Non-determinism would break replay, and replay is a graded deliverable.
The heuristics already beat the baseline decisively (below), so RL would be risk without measured upside.
Instead we do online opponent modelling sized to the ~210 observations a series actually provides - hint credibility, move tendencies, barrier response.
Measured through the real match machinery on held-out games - seed 11, never used while tuning, 60 games per matchup:
matchup | captures | rate | 95 % Wilson CI |
our cop vs greedy thief | 60/60 | 100 % | 94–100 % |
greedy cop vs our thief | 0/60 | 0 % (100 % survival) | 0–6 % |
greedy vs greedy (reference) | 4/60 | 6.7 % | 2.6–16 % |
Zero peer disagreements and zero audit failures across all 180 games.

One tunable decides the match, and it is not the one we expected:

barrier_threshold swings the capture rate from 4 % to 100 % across its
range. A barrier is impassable for both sides, so a cop that walls on weak
evidence fences itself away from the thief it is chasing - we shipped 0.15 for
weeks, which cost roughly a third of our games. lookahead is a genuine null
result: depths 1–4 produce byte-identical games, because an isotropic diffusion
kernel preserves the ranking of the candidate moves it is meant to separate.
The caveat we are obliged to state. Our thief survives every cop we have
ever faced or archived - and still loses, around step 30, to our own sealer,
whose halving plan no opponent has shown. Both directions are pinned rather
than smoothed over (test_thief_beats_sealing_cops.py records the defeat and
its price; the corner-hunt and archive suites record the survivals), because a
thief graded only against the cops it beats has been graded against itself.
Full derivations, confidence intervals, the token-cost table and the references
are in notebooks/analysis.ipynb. Reproduce with:
uv run python scripts/baselines.py --games 60 --seed 11 # held-out comparison
uv run python scripts/sweep.py --games 24 --seed 7 # sensitivity sweeps
uv run python scripts/measure_tokens.py # token censusWhat testing against a stranger taught us
We cloned the course reference simulator and pointed it at us. Nothing worked, in either
direction. The reference names the MCP tool argument message on three tools and payload
on one; we sent payload to all four and accepted only payload. Every turn and every
proposal was rejected by argument binding before a single byte of game logic ran - against
any agent built on the reference, which is most of the class.
Four more incompatibilities followed: a required timestamp we never sent, a claimed_cell
field their parser rejects outright, and three claim fields whose types differed. Then the
audit reveal turned out to be sent as a bare list where the schema declares an envelope, so
both peers recorded TAMPERED for games nobody had cheated in.
Every one of those passed our own tests. The commit-reveal core, by contrast, survived contact
unchanged: our commit_of reproduces the reference's signature byte for byte.
The lesson we would carry to any distributed project: a green test suite proves your code agrees with your assumptions, not that your assumptions are right.
Documentation
Document | Purpose |
| Product requirements (FR-* ids, KPIs, milestones) |
| Architecture: C4 + FSM diagrams, ADR-001..021, module map |
| 688-task build plan with traceability and progress |
| Current state, open items, and every correction measured |
| What crosses the wire, and what never does |
| Threat model, prompt-injection defences, secret handling |
| Nielsen heuristics mapped to dashboard decisions; accessibility |
| The four extension seams, with a worked plugin |
| Every config key, its file, and its Appendix F negotiability |
| What each gate checks and how to reproduce a failure |
| Conventions: core sync, commits, tests, match-day freeze |
| ISO/IEC 25010 quality characteristics mapped to evidence |
| Every handled boundary condition, each linking its test |
| Measured token consumption and the cost model |
| What is known to be incomplete, with the evidence |
| Sensitivity studies, baselines, cost table, references |
| Mechanism designs |
| Tunnel and connectivity procedures |
| Source digests (book, guidelines, reference simulator, A6 retrospective) |
Auditing an opponent
Commit-reveal proves a peer did not rewrite history. It proves nothing about whether they played by the rules, and those are different guarantees - we conflated them for the whole league phase and could not say, after losing a series, whether the play had been legal.
Auditing is post-match by design: rules 33-35 void a match for contradictory reports, so an agent acting on its own accusation converts a suspicion into a mutual zero. Everything below records evidence and changes nothing about how we play (PLAN ADR-019).
# replay their revealed records through the fair-play rules: movement legality,
# the Barrier Law, the budget, step order, hint length - and say what it could NOT check
uv run python scripts/audit_opponent.py --team vibecode
# are we disclosing scent on the same terms they are?
uv run python scripts/scent_parity.py --since 2026-08-14T16:00 # UTC
# 323 of 323 sealed capture claims name the claimer's own revealed cell
uv run python scripts/claim_evidence.py
# both repos must declare the same counted-match count (rules 37-38)
uv run python scripts/reconcile_counted.py ../pursuit-thief-agent --applyWhat a reveal can settle and what only the wire can: a peer's sealed record holds what that
peer chose to seal. Movement and barriers are checkable from an archive alone; smell_grid,
capture_claim, hint and response times are checkable only against what arrived on the
wire, which is why FrameLog keeps them as sent. The audit reports those as not checkable
rather than folding them into a clean verdict - "we looked and agreed" and "there was nothing
to look at" must never read the same.
Three results that constrain every strategy
All were established by measurement during the league phase, and all are load-bearing.
A barrier can shrink the board but can never take the thief. The book gives three capture
conditions (rules 46-47); the course reference implements exactly one. Its rules.py has
thief_result and is_captured and no barrier-capture or immobilisation check anywhere.
Every opponent we have met is reference-derived, so enclosure yields a mini-game we score
and they do not - the rules 33-35 contradiction. Every capture must be a claim the thief
confirms (PLAN ADR-020).
One cop cannot close on an open board. A 7×7 grid is the Cartesian product of two paths, so its cop number is 2 (Maamoun & Meyniel 1987), and an exhaustive fixed-point over all 49×49 states finds no state from which a movement-only cop can force a capture under simultaneous moves. Our cop tracking a thief to distance 2 and holding there for 28 steps is a theorem, not a defect. Barriers are the only resource that changes the answer (PLAN ADR-021).
And with a plan, barriers do change it. The halving seal converts every reacting thief the bench can build - the board halved, the half halved, the 3×3 solved exactly, ending in a co-location claim inside the 35 moves. The two constraints above still govern the shape of that win: it must end on a claim the thief confirms, and it cannot be had by movement alone. What was long recorded here as our largest competitive risk - a cop that tracks to distance 2 and holds - is closed; the risk that remains is an opponent whose own cop plays a plan as complete as ours, and the thief's floors are priced against exactly that.
tests/regression/cop_duel.py is the cop-side benchmark those claims are tested with - our
cop against an adaptive thief, with the belief the real ingress path builds from scent. A
recorded opponent line does not react and cannot measure closing.
License & attribution
MIT (see LICENSE). Protocol shapes and artifact schemas interoperate with the course
reference simulator rmisegal/Game-P2P-Cop-Chase
(educational license); where book and code conflict, the book governs.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Hosted MCP for Conductor Relay: a verifier-backed agent work exchange and cold marketplace.
Coordinate multiple AI agents over MCP: atomic claims, leases, shared ledger, handoffs, tasks.
Agent-native collaboration network: orchestrate a team of long-running agents from any MCP client.
Agent-native marketplace. Bootstrap, list inventory, search, negotiate, and trade via MCP.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables decentralized turn-based gameplay between thief and police agents in a peer-to-peer network, handling turn coordination, message passing, and protocol enforcement.
- FlicenseNot gradedqualityCmaintenanceEnables a police agent to autonomously chase a thief in a peer-to-peer grid game using FastMCP for turn-based communication and decentralized orchestration.
- FlicenseNot gradedqualityBmaintenanceImplements a distributed cops-and-robbers game agent as a FastMCP server, enabling peer-to-peer play with no central server. It manages turn-based moves, belief tracking, strategy selection, and secure protocol via SHA-256 commit-reveal.
- FlicenseNot gradedqualityBmaintenanceRuns a decentralized thief agent for a peer-to-peer cops-and-robbers game, using FastMCP to exchange moves and messages with a police agent while employing Bayesian belief and credibility-based bluffing strategies.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/najikay/pursuit-cop-agent'
If you have feedback or need assistance with the MCP directory API, please join our Discord server