agent-governance MCP server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@agent-governance MCP servershow me which tools the analyst agent can access"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
agent-governance
A governed agent platform where the dangerous action is impossible, not discouraged — plus an evaluation harness whose grader is itself graded.
The design rule is inversion-first: start from the failure and work backward. An agent must never autonomously write a commitment (an owner, a due date, a status) into a shared work system. The weak implementation of that rule is an instruction in a prompt. The strong one is that the credential physically cannot hold the scope, so no amount of clever input, prompt injection or model error produces the forbidden write. This repository is the strong one.
git clone https://github.com/Geomanti/agent-governance
cd agent-governance
pip install -e '.[test]'
python -m agent_governance.evals.runner # the eval gate; exit code 1 blocks
pytest tests/ -q # the unit suiteWhat's measured
Every figure below is produced by the code in this repository, on the default
registry of 8 tools and 3 agents. Reproduce with
python -m agent_governance.evals.runner.
Property | Value |
Eval gate | PASS — 12 cases, 19 assertions, 2 suites |
Mutation score | 7/7 killed |
Forbidden-path audit | CLOSED — no reachable tool |
Unit tests | 53 passed |
Forbidden scopes | 4, rejected at credential construction |
Per-agent tool surface — the withheld set is absent, not disabled:
Agent | Tier | Reachable | Withheld | Scopes |
| READ_ONLY | 3 of 8 | 5 | 3 |
| CONTACT | 5 of 8 | 3 | 5 |
| PROPOSE | 6 of 8 | 2 | 6 |
That table is enforced in three places at once (grant-time filtering, the
registration surface, and a call-site re-check), which is why the mutation
surface-unfiltered is caught rather than merely survived.
Related MCP server: Easy MCP Proxy
The three layers of "impossible"
1. Construction-time rejection. A credential requesting any scope in
FORBIDDEN_SCOPES cannot be instantiated. The forbidden class is refused where
the credential is built, before the agent exists.
2. Registration-time filtering. Tools exceeding a credential are never
registered on that agent's surface. Over MCP, tools/list returns a different
set per credential, and a withheld tool answers Unknown tool: — there is no
code path to call.
3. Call-site re-check. Even a direct in-process call re-verifies the scope, so the guarantee does not depend on the registration layer remaining the only entry point.
Verified over the real protocol, not a mock: the server was driven as a
subprocess over stdio, and tools/call on propose_issue_note as the
analyst returned isError: true with Unknown tool.
$ python -m agent_governance.server --agent analyst --list-tools
read_issue scopes=['read:issue']
read_digest scopes=['read:digest']
read_ledger scopes=['ledger:read']
withheld (5): contact_person, draft_update, propose_issue_note, read_audit_log, read_private_channelThe evaluation layer
The distinction that matters is falsifiability. A spot check produces an opinion; an evaluation produces a number a regression can move and CI can block on. Four things make that true here:
Behavioral assertions over trajectories. Cases assert on the sequence of decisions and effects a run produced, not on the text it emitted.
Safety assertions are never averaged. A case passes only if every safety assertion holds; quality assertions are scored separately against a threshold. Blending them is how a suite starts passing while being unsafe.
A written rubric, weighted, with required citations. A judge that cannot cite a criterion is rejected, and one that skips a criterion raises rather than scoring it zero.
The grader is graded.
evals/mutants.pyseeds deliberate defects into the live system, runs the entire suite, and counts a mutant as killed only if the gate actually blocks.
The meta-eval found real holes in this suite
Running the meta-eval against an early version of the golden suite scored 5/5 killed with a predicate-level harness: it probed one function per mutant. That number was misleading. Reimplementing it as suite-level mutation — mutate a live method, run the whole suite, require the gate to block — produced 6/7, and the survivor was genuine:
tier-ceiling-ignoredsurvived. No case asserted that a credential cannot claim an autonomy class above its tier ceiling. The check existed in the source and was untested. Asafety-tier-ceilingcase was added, andTIER_CEILING_ENFORCEDnow asserts the refusal directly.
A separate falsification pass neutered grants_for by hand. The call path
still refused writes, so every safety case passed while the tool surface
exposed all eight tools — the guard held and the surface leaked. That is
precisely the class of blind spot a green light hides, and it is why the
no-forbidden-tool and surface-within-scopes assertions now compute
allowed_tools independently by raw set comparison rather than trusting the
filtering helper's own output.
Both gaps are fixed; the score is 7/7. The sequence is reported here because a mutation score is only meaningful alongside the evidence that it can fall.
Durability, rollback, and the review gate
RunStore writes run state via write-temp-then-rename, so a process killed
mid-write leaves either the old file or the new one. A run interrupted between
start and finish is reloaded as INTERRUPTED, not lost.
Rollback replays the inverse of every applied action, newest first. The inverse
is recorded at stage time, not at failure time, which is what makes rollback
work for a run whose process died. If the apply phase fails — including the
case where an effect has no registered handler — the run is recorded as a
failure and rolled back, rather than left stuck in RUNNING with effects
half-applied.
The human review gate is a state transition, not a prompt instruction. A
propose-tier run parks in AWAITING_REVIEW and applies nothing; approve
requires a named approver, reject moves the run to ROLLED_BACK.
Contact discipline
Contact is a budget that depletes, not a permission that is on or off:
a hard ceiling per person, with refusals recorded rather than silently dropped;
a bundling window, so six intents inside one window cost one interruption, not six;
a value-before-ask rule — a message that requests something must carry something useful first, and a bare ask is rejected at construction.
Layout
src/agent_governance/
capabilities.py scopes, least-privilege credentials, forbidden class
budget.py interruption budgets, bundling, value-before-ask
audit.py hash-chained, tamper-evident decision log
runtime.py durable queue, rollback, review gate
server.py governed surface + MCP adapter (mcp>=2.0)
evals/
harness.py assertions, trajectories, rubric, judge, CI gate
golden.py the golden cases and reusable assertions
mutants.py suite-level mutation testing + forbidden-path audit
runner.py the gate: suites + meta-eval, exit code for CI
tests/ 53 unit tests, incl. real MCP transport testsHonest scope
This is a platform skeleton, not a deployment. There is no warehouse sink, no scheduler, and no model provider wired in —
Agent.runtakes a callable precisely so the runtime stays model-agnostic and headless. The governance, durability, and evaluation layers are the substance.The mutation score is 7/7 over seven hand-written operators, not exhaustive coverage of the source. It measures the suite against the failure classes the design names, which is a smaller claim than "the suite is complete".
Assertions are as good as their authors.
_no_commitment_writeinspects the effect names a run applied. A handler that ignored its own name and wrote a commitment anyway would not be caught by it. The structural guard (layer 1) is what prevents that; the assertion is a second net, not the primary one.The MCP surface is tested in-memory via
InMemoryTransportand over real stdio as a subprocess; it has not been exercised against a third-party client implementation.
License
MIT — see LICENSE.
Available Tools
6 toolsdraft_updateB
Draft a status update. Produces text; writes nothing to a system of record.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full disclosure burden, and it does surface the key behavioral trait: the tool is non-mutating and produces text only. That said, it says nothing about determinism, permissions, or whether the draft persists anywhere else, so it is only partially transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler; the purpose leads and the behavioral caveat follows. Nothing needs trimming.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-described, and the single-parameter surface is simple. Still, the meaning of the one input is left undocumented, which is a real gap for even a lightweight tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single 'text' parameter (default empty string) is never explained. The agent is left to infer whether 'text' is an instruction, a seed draft, or a template.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Draft a status update') and adds what the output is ('Produces text'). It is clearly distinct from the read-heavy siblings and from propose_issue_note, but it never explicitly names or contrasts a sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'writes nothing to a system of record' implies this is the safe choice when you only want generated text rather than a persisted write, which is useful implied usage. However, no explicit when-to-use, when-not-to-use, or named alternative is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
propose_issue_noteA
Propose an issue note. Applies only after a human consents; cannot set an owner, due date or status.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose non-obvious behavior: the operation requires human consent and is constrained from setting owner, due date, or status. It leaves gaps around whether the note is persisted/pending, reversibility, and permissions, but the consent gating and field restrictions are meaningful beyond the tool name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, and the core behavior plus the consent constraint are front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. For a consent-gated proposal tool with no annotations, the description covers the key conditional behavior and field restrictions; only post-consent semantics and the note parameter remain unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'note' parameter has 0% schema description coverage and the description never mentions it at all — no format, length, or markdown expectations. With one undocumented parameter, the description does not compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Propose an issue note'), which is clearly distinguishable from the read-only siblings (read_issue, read_digest, read_ledger, read_audit_log) and from draft_update, though it never explicitly contrasts itself with draft_update. Purpose is unambiguous without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit precondition for use: 'Applies only after a human consents.' That tells the agent the gating condition for invoking it. It does not name alternative tools (e.g., draft_update) or state when-not-to-use, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_audit_logB
Read this run's decision audit log.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Read' implies a safe read-only operation and 'this run's' implies scoping, but it says nothing about authorization, what decisions are recorded, ordering, or size; it relies on the output schema to convey content.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. For a zero-parameter tool with an output schema, this is an appropriately sized description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, and there are no parameters to explain, so the core is adequate. However, the key term 'this run' is undefined and there is no link to sibling tools that read related artifacts, leaving the agent without context for when this log is relevant.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 and there is no parameter semantics to document. Nothing in the description is needed or missing on this dimension.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (read) and resource (decision audit log) and scopes it to 'this run'. It distinguishes the artifact from siblings like read_digest and read_ledger reasonably well, though 'decision audit log' vs 'ledger' overlap enough that an agent may still hesitate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to read this log versus read_ledger or read_digest, nor any stated trigger or exclusion. The agent is left to infer the use case entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_digestC
Read the current team digest.
| Name | Required | Description | Default |
|---|---|---|---|
| team | No | core |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies a read operation but does not disclose scope ('current' is vague), whether it requires team membership/permissions, how large the digest is, or any freshness/recency semantics — all left unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or redundancy. It is efficient, though the brevity reflects under-specification rather than disciplined conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. But with no annotations and no parameter documentation, the description leaves key agent-facing questions unanswered: what the digest contains, how it differs from the log/ledger tools, and what the team parameter accepts.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'team' parameter has 0% schema description coverage and no enum, so the schema does not explain what values are valid or what the 'core' default means. The description only weakly implies team scoping via the phrase 'team digest' and adds no format or selection detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read' + 'team digest'), so an agent knows the operation. However, it does not differentiate from siblings like read_ledger, read_audit_log, or read_issue, and 'digest' is left undefined relative to those read-oriented tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of the alternative read tools (read_ledger, read_audit_log, read_issue). The agent must infer the selection criteria entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_issueC
Read one issue from the tracker.
| Name | Required | Description | Default |
|---|---|---|---|
| issue_id | No | ISSUE-1 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing about read-only semantics, permission requirements, or what happens when the issue does not exist. For a single-record fetch, error and not-found behavior is exactly what an agent needs and is entirely absent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with no filler. It is efficient, though the brevity borders on under-specification for a tool with this many informational gaps.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not required, and the tool is structurally simple (one flat string parameter). Still, the identifier format and not-found behavior are missing, leaving the definition only minimally adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter (issue_id) is never mentioned in the description. The schema supplies a default of 'ISSUE-1' but no format or identifier guidance, and the description does nothing to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('one issue from the tracker'), which is clearly distinct from the resource-oriented siblings like read_digest, read_ledger, and read_audit_log. It stops short of explicitly differentiating itself, but the resource noun alone makes routing unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus the other read tools, no prerequisites, and no exclusions. The agent must infer that this is for fetching a single known issue.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_ledgerC
Read the ledger of work absorbed per agent.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosure, and "Read" only weakly implies a safe, non-mutating operation. It says nothing about scope (whole ledger vs per-agent window), pagination, permissions, or how the ledger is populated, though the presence of an output schema covers the return shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though its brevity comes at the cost of the detail noted in other dimensions rather than being admirably compact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no parameters, and an output schema, the description still needs to convey scope and how it differs from read_audit_log/read_digest. It leaves both unaddressed, so an agent cannot confidently choose this tool over its siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to explain; the baseline for a parameterless tool applies. Nothing in the description contradicts the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It states a verb (Read) and a resource (ledger of work absorbed per agent), which is more than a tautology, but "work absorbed" is unexplained jargon. It also fails to distinguish itself from sibling read_audit_log or read_digest, which sound like the same kind of read-report operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to use this tool rather than read_audit_log or read_digest, nor any prerequisite or context for calling it. The agent must guess based on the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
draft_update - First observed
propose_issue_note - First observed
read_audit_log - First observed
read_digest - First observed
read_issue - First observed
read_ledger
TDQS
Scored across 6 tools
The four read_* tools each target a clearly distinct resource (issue, digest, ledger, audit log), so reads are unambiguous. draft_update and propose_issue_note are the only pair with mild overlap (both produce content concerning updates/notes), but the 'writes nothing' vs 'gated write' distinction is spelled out. An agent can reliably pick the right tool.
Every tool follows a clean verb_noun snake_case pattern: read_issue, read_digest, read_ledger, read_audit_log, draft_update, propose_issue_note. The verb set (read/draft/propose) maps predictably to the side-effect class of each tool.
Six tools is well-scoped for a governance/oversight server, and each tool has a distinct, non-redundant role. Nothing feels padded or missing at the count level.
The surface covers reading state and proposing changes, but there is no way to list issues (only read a single one), check the status of a submitted proposal, or follow a proposal through to resolution. These are notable gaps for an oversight workflow, though the read-plus-propose design is deliberately human-gated.
Maintenance
Related MCP Connectors
Security & DLP proxy for MCP: tool-poisoning scans, PII redaction on tool args/results. Beta.
Governed MCP gateway: one endpoint for your tools, with credential custody and audit log.
Find, vet, and run MCP tools through a secure audited gateway with prompt-injection risk scoring
MCP facade over the Nebelus Construction API. ~48 tools give full agent build parity: create/update/probe agents, edit graphs, attach knowledge and vector stores, wire connectors, set governance policies and locked guardrails, enable grounding-trace, and read deployment wiring. Purpose-built for regulated industries: data residency is enforced per region (EU / GCC-KSA), with PII controls and an audit trail. Agents are created as drafts — no deploy tool is exposed over MCP by design; publishing happens in the Nebelus console.
Related MCP Servers
- AlicenseAqualityCmaintenanceProvides per-Subagent MCP controls to any coding agent or client across all your MCPs and prevents context window waste. Loads only 3 tools instead of all your MCP Server's tool definitions. Agents discover tools on-demand, only when needed and only the servers and tools they are allowed.446 PyPI41MIT
- AlicenseNot gradedqualityBmaintenanceEnables aggregation, filtering, transformation, and composition of tools from multiple MCP servers through a single proxy with tool views.5AGPL 3.0
- AlicenseAqualityAmaintenanceEnables MCP-capable clients to query the tool registry, check install status, get tool recommendations for CTF or bug-bounty work, and run installed security tools through a governed execution path.1561MIT
- AlicenseNot gradedqualityBmaintenanceProvides a reusable MCP server foundation with explicit tool registration, operation-mode separation, scope-based permissions, structured results, and built-in health/status tools.MIT