adversary-gate
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| ADVERSARY_GATE_POLICY | No | Gate flags (`--sandbox bwrap`, `--triage jev`, floors). An agent that can pass `--coverage-floor 0` grades itself. Evidence flags here are refused. | |
| ADVERSARY_GATE_PYTHON | No | The interpreter's site-packages are harness too — a plugin installed there runs inside pytest. Use one the agent cannot write to. | |
| ADVERSARY_GATE_BASE_REF | No | The baseline is the oracle: an agent could commit a rewritten test and name that commit as the base. Unpinned, the answer says "baseline_chosen_by": "agent". |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| verify_repoA | Judge the working tree of |
| gate_policyB | What this server enforces. Read-only: the agent cannot change it. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 2 tools
verify_repo executes tests and judges the working tree, while gate_policy returns read-only enforcement information. These purposes are completely distinct with no overlap, so an agent can easily select the right tool.
Both names use snake_case, but verify_repo follows a clear verb_noun pattern while gate_policy is a bare noun phrase. This mixes action-oriented and resource-oriented conventions, though both remain readable.
With only two tools, the surface feels thin even for a focused verification server. The tools are well-chosen but leave no room for auxiliary operations like baseline introspection.
The core verification and policy-reading operations are present, covering the main lifecycle for an adversary gate. Minor gaps exist, such as retrieving baseline details or historical gate results, but agents can work around them.