persona-constitution
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@persona-constitutionscan this code for placeholder violations"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
CTO-MCP — persona-constitution MCP Server + Agentic PR Review
An MCP (Model Context Protocol) server that serves the Oluwaferanmi Oluwagbamila Agentic Engineering Persona — LLM Operational Constitution v3.0.0 to any MCP-capable coding agent, and a PR review gate that enforces the constitution's Zero-Framework-Tolerance rules on pull requests — as an MCP tool, a CLI, a reusable GitHub Action, and an opencode agent.
Grounded in SWEBOK v4.0 (18 Knowledge Areas), the NASA/JPL Power of 10, and Zero Framework Tolerance.
The constitution exists to counteract a specific, structural LLM failure mode: producing code that has the shape of a solution but none of the substance — skeletons, TODOs, stubs, and "you can extend this to…". This server makes the constitution queryable, ships a scanner that mechanically detects those violations in generated code, and turns that scanner into a diff-aware pull-request reviewer.
The Supreme Law — Every code output must be complete, executable, and correct. Not a scaffold. Not a pattern. Not a direction. Code that runs. Logic that is correct. Implementation that is done.
Requirements
Python 3.9+ — the test suite is run against CPython 3.9.6 and 3.14.6.
Zero runtime dependencies. CodebaseCSI (MIT), which backs the scanner, is vendored at
codebase_csi/— provenance, pinned upstream revision, and the local-modification ledger live incodebase_csi/VENDORED.md.Optional
[ast]extra: tree-sitter grammars that upgrade the scanner from regex to real AST analysis for JavaScript, TypeScript, Java, Go, Rust, Ruby, C and C++. Without it the scanner still runs and says so in itsenginesoutput. The pins are an ABI compatibility matrix verified on the 3.9 floor — see the comment block inpyproject.toml.
Related MCP server: AI Knowledge Center MCP
Layout
CTO-MCP/
├── persona_constitution/
│ ├── __init__.py Package API re-exports
│ ├── scanner.py Detection engine: CodebaseCSI + prose rules + Python AST
│ ├── ast_bridge.py constitution-xast: tree-sitter engine for brace languages
│ ├── logic_rules.py Deep logic rules: Po10 metrics, empty loops, identical
│ │ branches, constant conditions, unreachable code
│ ├── server.py MCP server: JSON-RPC 2.0 over stdio
│ ├── review/
│ │ ├── diff.py Unified diff parser (git-quoted paths, hunk edge cases)
│ │ ├── engine.py Diff-aware review: attribution, C-03 policy, exclusions
│ │ ├── config.py .persona-review.json: one policy for both gate factors
│ │ ├── report.py Renderers: text, Actions annotations, GitHub review payload
│ │ ├── github_client.py Stdlib GitHub REST client (bounded retries)
│ │ └── cli.py persona-pr-review: --staged/--diff/--git/--github/--install-hook
│ └── data/ Ships inside the package, so an installed copy works
│ ├── CONSTITUTION.md Full constitution (the served corpus)
│ └── DIRECTIVES.md Distilled directives for system-prompt injection
├── codebase_csi/ Vendored CodebaseCSI (MIT) — see VENDORED.md
├── .persona-review.json This repository's own review policy (dogfood)
├── .opencode/
│ ├── agent/pr-review.md Factor-2 agent: business-logic test research + G1-G5
│ └── command/review-staged.md /review-staged: both factors before committing
├── .github/workflows/
│ ├── ci.yml Lint, tests (3.9/3.14 x with/without [ast]), packaging
│ └── pr-review.yml Dogfood: this repo's PRs pass through its own gate
├── action.yml Reusable composite GitHub Action for any repository
├── tests/ Unit, engine, review, two-factor integration, smoke, e2e
├── tools/
│ └── benchmark_scanner.py Adversarial accuracy benchmark (regression gate)
├── LICENSE
├── pyproject.toml
└── README.mdSetup
From the repo root:
python3 -m venv .venv
.venv/bin/pip install -e ".[ast]" # or plain `.` to skip the tree-sitter tier.venv/bin/python is then the interpreter your MCP client must launch. Starting the
server with an interpreter that cannot import the vendored codebase_csi exits 1 with an
explanatory message on stderr — it will not fall back to a weaker scanner and report
misleadingly clean results.
Install into opencode
Two mechanisms, used together. The instructions file injects the constitution into the system prompt of every session, for every configured model; the MCP server provides on-demand structured lookup and mechanical verification.
Injection is not enforcement. instructions is system-prompt text, and whether a model
follows it is a property of that model, not of this repo — only the MCP scanner performs a
mechanical check. Two caveats worth knowing before you rely on it:
Small-context models can choke on the payload.
DIRECTIVES.mdis a substantial system prompt; on a 16k-context deployment (tested: Azure Phi-4) sessions hung rather than degrading gracefully. Prefer models with a large context window, or trimDIRECTIVES.mdfor small ones.Compliance is per-model and worth spot-checking. Verified by direct observation on OpenAI- and Anthropic-adapter models, which reproduced gate and law text verbatim on request. That is a sample, not a proof across every provider — re-verify on yours.
Add to ~/.config/opencode/opencode.json (or opencode.jsonc), replacing <REPO> with the absolute path to this clone:
{
"$schema": "https://opencode.ai/config.json",
"instructions": ["<REPO>/persona_constitution/data/DIRECTIVES.md"],
"mcp": {
"persona-constitution": {
"type": "local",
"command": ["<REPO>/.venv/bin/python", "<REPO>/persona_constitution/server.py"],
"enabled": true
}
}
}instructions is global opencode config, so the directives are injected into the system
prompt of every model and provider you have configured — there is no per-model setup.
Note that the directives consume context: models with small context windows may struggle.
Restart opencode afterwards — config is loaded once at startup and is not hot-reloaded.
Install into other MCP clients
Any client that speaks MCP over stdio works. Claude Desktop (claude_desktop_config.json):
{
"mcpServers": {
"persona-constitution": {
"command": "<REPO>/.venv/bin/python",
"args": ["<REPO>/persona_constitution/server.py"]
}
}
}Tools
Tool | Arguments | Returns |
|
| Table of contents + Supreme Law by default; any named section; or |
|
| One SWEBOK v4.0 Knowledge Area with its LLM operational discipline; omit |
|
| One Power of 10 rule with code / architecture / organisational applications, or all ten |
| none | The G1–G5 pre-emission gates and the prohibited-marker checklist |
|
| JSON verdict |
|
| Diff-aware review JSON: per-file findings attributed to changed lines, pre-existing debt counted separately, C-03 test-presence policy, verdict |
| exactly one of | G4 made mechanical: extracts imports (Python AST, JS/TS specifiers), classifies stdlib/built-ins/first-party locally, then verifies the rest exist on PyPI/npm. Hallucinated packages (slopsquatting surface) → |
section values for get_constitution
toc · preamble · identity · anti-deception · intelligence-architecture · t-shape · swebok · consensus-protocol · iteration-protocol · agentic-pathway · power-of-10 · operational-directives · knowledge-graph · invariants · references · full
(hive-mind is still accepted as a deprecated alias for consensus-protocol.)
The scanner
scan_code_for_violations is a union of five engines, because no one of them is adequate alone:
Engine | Contributes |
CodebaseCSI | Structural stubs, mock implementations, always-success functions, print-only bodies, fake data, pass-through functions, TODO markers. |
Constitution prose rules | Class 2 / Class 5 narrative deferral, and empty-body / unimplemented-stub detection for JavaScript, TypeScript, Java, Go and Rust |
Python AST analysis | Suppresses markers inside ordinary string literals; distinguishes genuine stubs from legitimate abstract declarations; classifies bare vs. typed |
constitution-logic (Python) | Deep logic shape, all warnings: Po10 Rule 1 (cyclomatic > 10) and Rule 4 (function > 50 lines), empty loop bodies, identical if/else arms, constant |
constitution-xast ( | Tree-sitter parse of JavaScript, TypeScript, Java, Go, Rust, Ruby, C and C++; judges hardcoded-return stubs (including |
Verdict philosophy: mechanical certainties (stubs, scaffold markers) are violations and FAIL; judgement calls (metrics, logic shape, test presence) are warnings and REVIEW - the agent layer adjudicates them, never silently.
Coverage by failure class:
Class 1 — Framework Generation:
TODO,FIXME,XXX, "your code here", "implement … here/later",raise NotImplementedError,todo!(),unimplemented!(),panic("not implemented"), empty function and method bodies, and bodies consisting only ofpassor...Class 2 — Scaffold Deception: "rest of the implementation", "follows the same pattern", "omitted for brevity", "and so on for the rest", "similar for the others"
Class 3 — Confidence Mismatch: empty
catch {}blocks, bareexcept: pass, always-success functionsClass 5 — Iteration Deferral: "left as an exercise", "you can extend this", "this is a starting point", "you would want to add", "the full implementation would", "in production you would"
Verdicts: FAIL if any violation fires, REVIEW if only warnings fire, PASS otherwise.
Measured accuracy
Measured against a 37-case adversarial corpus — 24 real violations across seven languages, plus
13 pieces of legitimate code specifically constructed to resemble violations (a linter that
matches on the string "TODO", a typing.Protocol whose methods are ..., a documented
except OSError: pass, a React placeholder= attribute, an anonymous no-op callback, an empty
Java constructor, a noop-default arrow binding, a busy-wait loop):
Configuration | Correct verdicts |
CodebaseCSI alone | 15/37 — 40% |
Constitution prose + structural rules alone | 22/37 — 59% |
Union without the | 30/37 — 81% |
Union with the | 37/37 |
The engines fail on largely disjoint inputs, which is why the union beats each: CodebaseCSI
misses every non-Python structural stub and every prose deferral; the prose rules miss
Python-semantic stubs such as always-true and print-only functions; and seven corpus cases
(function getUser(id) { return null; }, a bare UnsupportedOperationException, a
template-literal throw new Error(`not implemented`), a braceless empty Ruby method, a Go
return nil stub, and stubs bound through const name = (args) => {...} declarators) are
decidable only on a real syntax tree. Both baselines are enforced in CI across both
configurations.
Reproduce with:
.venv/bin/python tools/benchmark_scanner.pyRead these numbers with suspicion. The corpus is small and was written by the same author as the rules, which biases the result upward. It is a regression guard, not a general accuracy claim.
A PASS is necessary but not sufficient. Static analysis proves the absence of placeholder
markers — it cannot prove executability, correctness, or dependency honesty. The G1–G5 gates still apply.
Two-factor review
The scanner generalises to a two-factor review gate through
persona_constitution/review/: the same deterministic engine judges the change at two moments
— factor 1a when files are staged (pre-commit hook, contents read from the git index so
what is judged is exactly what would be committed) and factor 1b at PR time (Action / CLI /
MCP tool). Factor 2 is the agent layer, which at both moments researches the project's
business-logic tests before exercising judgement. Findings are attributed to the lines the
change introduces; pre-existing debt in touched files is counted and surfaced but never gates
the merge.
One policy file, .persona-review.json at the repository root, drives every surface (both
factors, the Action, and the agent), so a rule can never be enforced at one gate and forgotten
at the other:
{
"exclude": ["tests/*", "vendor/*"],
"require_tests": "warn",
"min_test_trigger_lines": 5,
"test_globs": ["qa/*"],
"business_logic": {
"description": "what this system's correctness actually means",
"critical_paths": ["billing/*"],
"business_logic_tests": ["tests/test_billing.py"],
"test_commands": ["python -m unittest discover -s tests"]
}
}require_tests is the C-03 policy (non-trivial code ships with tests): when a diff changes at
least min_test_trigger_lines of production logic and touches zero test files, every such
file is flagged — warn surfaces it for adjudication, fail gates the merge. The deterministic
policy is deliberately diff-global; mapping which tests cover which changed behaviour is the
agent's business-logic research, not a glob matcher's.
Delivery surfaces, one engine:
0. Staged gate (factor 1a) — install once per clone:
persona-pr-review --install-hook # writes .git/hooks/pre-commit (refuses to clobber
# a foreign hook without --force)
persona-pr-review --staged --json # what the hook runs: git diff --cached, contents
# from `git show :0:path` (the index, not the worktree)A commit with staged violations is blocked; bypassing with git commit --no-verify is loud and
still lands in front of factor 1b and the agent.
1. MCP tool — review_patch (table above). The server stays offline and deterministic: the
agent brings the diff (gh pr diff), the tool returns structured findings.
2. CLI — installed as persona-pr-review:
# local working tree against a base
persona-pr-review --git origin/main...HEAD --root . --json
# an existing diff file, annotated for GitHub Actions
persona-pr-review --diff change.diff --annotate
# a GitHub PR, posting REQUEST_CHANGES/COMMENT with inline comments
GITHUB_TOKEN=... persona-pr-review --github owner/repo#42 --post \
--require-tests fail --exclude 'vendor/*'Flags override .persona-review.json; --exclude appends to it. Exit codes: 0 PASS (or
REVIEW), 1 FAIL (or REVIEW with --fail-on-review), 3 operational error (including a
malformed policy file — a broken policy stops the gate rather than silently weakening it). The
GitHub client is stdlib urllib with bounded retries; the reviewer never emits APPROVE — a
scanner can prove the absence of markers, not the presence of correctness.
3. Reusable GitHub Action — the composite action at the repo root:
permissions:
contents: read
pull-requests: write
security-events: write # only needed when sarif-file is set
steps:
- uses: actions/checkout@v5
- uses: QuantmindSSI/CTO-MCP@main # pin a tag/sha in production
with:
exclude: "vendor/*" # appended to the repo's .persona-review.json
require-tests: "warn" # C-03; empty = use the repo's config
fail-on-review: "false"
sarif-file: "persona-review.sarif" # optional: findings in the Security tabViolations become ::error annotations on the changed lines and a posted review that
REQUEST_CHANGES; fork PRs are automatically downgraded to annotations-only so the token never
serves untrusted code. With sarif-file set, the review is also uploaded to GitHub code
scanning as SARIF 2.1.0 — findings appear in the repository's Security tab, tagged with
their MITRE CWE IDs (CWE-546 suspicious comments, CWE-1071 empty bodies, CWE-1069 empty
catches, CWE-684 contract-faking stubs, and the rest of the mapping documented in
scanner.py). This repository dogfoods the action on its own PRs
(.github/workflows/pr-review.yml).
4. opencode agent (factor 2) — .opencode/agent/pr-review.md defines the agentic layer.
Its protocol is layered and ordered: Layer 0 reads .persona-review.json and researches
the project's test landscape (which layers exist — unit, integration, e2e, smoke, regression —
and which tests reference the changed symbols, via git grep over the test tree); Layer 1
runs the deterministic gate through the MCP tool (its FAIL verdict cannot be overridden);
Layer 2 judges business-logic coverage — changed behaviour vs. covering tests found,
updated or not, with test-command runs as evidence; Layer 3 applies gates G1–G5 and the
nine review dimensions. Deterministic findings are supreme; agent judgement is additive only;
never APPROVE.
5. /review-staged command — .opencode/command/review-staged.md runs the full two-factor
flow on the staged index before a commit: the deterministic staged gate first, then the same
agent protocol, concluding "COMMIT" or "DO NOT COMMIT" with file:line-anchored required fixes.
The five verification gates
Run before emitting any code. All five must pass; if any fails, regenerate from the problem statement rather than patching.
Gate | Question |
G1 Executability | Copy-pasted into a blank file with the stated dependencies, does it run without modification? |
G2 Completeness | Does every function contain a real implementation? Any placeholder, TODO, or empty body? |
G3 Correctness | Execution traced for the happy path, the primary error paths, and the stated edge cases? |
G4 Dependency Honesty | Does every import, call, and referenced module exist in this output or a verified dependency? |
G5 Problem Fit | Does this solve the stated problem, at the stated scale, under the stated constraints — not a simpler adjacent one? |
Configuration
Variable | Effect |
| Absolute path to an alternative |
The server exits with status 1 and a message on stderr if the constitution file is missing or empty — a broken install fails loudly rather than silently serving nothing.
Tests
Run from the repo root, using the virtualenv interpreter:
.venv/bin/python -m unittest discover -s tests -v
.venv/bin/python tools/benchmark_scanner.py # regression gateThe suite is organised as a full test taxonomy:
Smoke (
test_smoke.py) — the critical path of every delivery surface in seconds: package import and version coherence, one stub/one clean scan, one review verdict, the console entry point, and a real MCP stdio handshake listing all six tools. A broken install fails here before anything else spends time.Unit (
test_server.py,test_ast_bridge.py) — constitution loading, markdown section extraction (all 14 sections, 18 KAs, 10 rules), scanner line-number accuracy and verdict boundaries, JSON-RPC dispatch, per-language xast stub detection, the deep-logic rules (Po10 metrics, empty loops, identical branches, constant conditions, unreachable code,while True:exemption), and the constant-drift guards that keep cross-engine deduplication sound.Regression (
test_server.pyfalse-positive classes,tools/benchmark_scanner.py) — every previously confirmed false positive stays fixed (string literals,Protocol,@abstractmethod, documentedexcept: pass,{}literals inside Python bodies, triple-quote parity), and the 37-case adversarial corpus enforces environment-aware accuracy baselines (37/37 with[ast], 30/37 without) in CI.Review engine (
test_review.py) — diff parsing (renames, binary, git-quoted unicode paths, submodule and mode-only changes, CRLF, lying hunk headers, adversarial garbage), changed-line attribution vs. pre-existing debt, exclusion globs, renderer contracts (neverAPPROVE, comment budgets, annotation escaping), and thereview_patchtool over the dispatcher including C-03 arguments.Two-factor integration (
test_two_factor.py) — real git repositories: the policy config contract (strict schema, loud failures), C-03 test-presence enforcement in warn/fail modes, staged-gate index authority (a fixed worktree must not mask a broken index), hook installation (foreign-hook refusal,--force), an installed hook blocking a genuinegit commitand then admitting a clean, tested change, and factor parity (staged gate and PR gate reach identical findings on the same change).End-to-end transport (
test_server.py) — a real subprocess driven over stdio:initializehandshake,tools/list, a full multi-tool session, malformed-input recovery, notification suppression, non-zero exit on a missing data file, and the invariant that stdout carries only protocol frames.
Protocol notes
Transport: newline-delimited JSON-RPC 2.0 over stdio, one message per line.
Methods:
initialize,tools/list,tools/call,ping.notifications/*are accepted and correctly produce no response frame.Protocol version:
2025-06-18; the client's requested version is echoed when it supplies one.Tool-level failures (bad arguments) return
isError: trueinside the result so the model can read and self-correct. Protocol-level failures return proper JSON-RPC error codes (-32700,-32600,-32601,-32602,-32603).Diagnostics go to stderr exclusively. stdout is never polluted with non-protocol bytes.
Credits
The structural detection layer of the scanner is provided by
CodebaseCSI, used under the MIT License and
vendored at codebase_csi/ (upstream revision and local modifications are recorded in
codebase_csi/VENDORED.md). This project adds the MCP interface, the Constitution corpus, the
Class 2 / Class 5 prose rules, the cross-language structural rules, the Python AST
false-positive suppression, the tree-sitter xast engine, and the diff-aware PR review stack.
License
MIT — see LICENSE.
CodebaseCSI is also MIT, and its license text is reproduced verbatim in the LICENSE file
under Third-Party Components, as its terms require.
References
IEEE Computer Society (2024). Guide to the Software Engineering Body of Knowledge (SWEBOK) v4.0. Ed. H. Washizaki. 18 Knowledge Areas.
Holzmann, G.J. (2006). The Power of 10: Rules for Developing Safety-Critical Code. IEEE Computer 39(6), 95–97.
Model Context Protocol specification — https://modelcontextprotocol.io
CodebaseCSI — forensic AI-generated-code detection. https://github.com/Thundastormgod/CodebaseCSI
Available Tools
7 toolsget_constitutionA
Retrieve the Agentic Engineering Persona constitution (v3.0.0) that governs all programming and software development work: SWEBOK v4.0, NASA Power of 10, Zero Framework Tolerance. Call with no arguments for the table of contents plus the Supreme Law; pass section for a specific part or 'full' for the whole document.
| Name | Required | Description | Default |
|---|---|---|---|
| section | No | Which section to return. Omit (or 'toc') for the table of contents plus the Preamble/Supreme Law. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose the return behavior by argument (TOC+Supreme Law vs. single section vs. whole document), which is useful, but says nothing about size/length of the full document, caching, or other operational traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the resource and its scope front-loaded before the invocation rules. The first sentence is dense with enumerated standards, which adds identity but is slightly heavier than strictly needed for tool selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-optional-parameter read tool with no output schema, the description covers the resource and the call behaviors well enough to invoke correctly. It stops short of clarifying retrieval granularity relative to the subset siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is fully enumerated in the schema, so the baseline is 3. The description largely restates the schema's own note ('omit or toc' → TOC + Preamble/Supreme Law) with only marginal added meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Retrieve') and resource ('Agentic Engineering Persona constitution v3.0.0') and enumerates the standards it codifies. However, the relationship to siblings like get_power_of_10 and get_knowledge_area, which appear to be subsets of this same document, is only implied rather than made explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent how to invoke in each case: no arguments yields the TOC plus Supreme Law, a `section` value returns a part, and 'full' returns everything. This is clear invocation guidance, though it offers no when-not-to-use or alternative-tool routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_knowledge_areaA
Retrieve one of the 18 SWEBOK v4.0 Knowledge Areas (KA-01 Requirements through KA-18 Engineering Foundations) including its LLM operational discipline. Pass ka as a number 1-18 or a name (e.g. 'Software Security'). Omit ka to list all 18.
| Name | Required | Description | Default |
|---|---|---|---|
| ka | No | Knowledge Area number (1-18) or name substring. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It makes clear this is a non-mutating lookup by using 'Retrieve' and by describing the returned content, and it discloses the no-argument default (list all 18). It does not mention pagination, size, or error behavior on an out-of-range `ka`, which is a minor remaining gap for a read-only lookup.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with what is retrieved and followed immediately by the input forms. Every clause earns its place; no filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema, the description covers what is returned, the valid input forms, and the default behavior. It could be slightly more complete by hinting at the shape of an individual KA entry, but there is no output schema to lean on and the core call semantics are fully specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and there is only one optional parameter, so the baseline is 3. The description goes beyond the schema by clarifying that `ka` accepts a range of numbers 1-18, a name substring, and that omission produces a full listing, and it supplies a concrete example ('Software Security').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Retrieve') plus a concrete resource (SWEBOK v4.0 Knowledge Areas, KA-01 through KA-18) with the exact content scope ('including its LLM operational discipline'). An agent can immediately distinguish this from the sibling retrieval tools (get_power_of_10, get_constitution, get_verification_gates), which map to different standards.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the two invocation modes: pass a number 1-18 or a name substring to fetch one area, or omit `ka` to list all 18. This is clear conditional guidance. It stops short of naming alternative sibling tools when, e.g., the user wants a different artifact, so it is not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_power_of_10A
Retrieve NASA/JPL Power of 10 rules (Holzmann 2006) with the persona's code-level, architecture-level, and organisational-level applications. Pass rule (1-10) for one rule, or omit for all ten.
| Name | Required | Description | Default |
|---|---|---|---|
| rule | No | Rule number 1-10. Omit for all rules. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the content shape (per-rule applications at three levels) but never states that the operation is a side-effect-free, static-corpus read, nor anything about return format or caching. Adequate but thin for an annotation-free tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the resource and its content front-loaded; nothing is bloated. The second sentence largely duplicates the schema's parameter description, so it is efficient rather than perfectly economical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter knowledge lookup with no output schema, the description conveys what is returned (the ten rules plus three levels of persona application) and how to scope the request, which is enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is already fully documented in the schema (1-10, omit for all). The description simply restates the same bounds and omission behavior, adding no syntax or semantic meaning beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Retrieve) and a precise resource (NASA/JPL Power of 10 rules, Holzmann 2006), and even names the payload structure (code-, architecture-, organisational-level applications). It does not, however, differentiate itself from adjacent knowledge-retrieval siblings such as get_knowledge_area or get_verification_gates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how to invoke it (pass `rule` 1-10 for one, omit for all ten) but gives no guidance on when to reach for this corpus versus the sibling knowledge/verification tools. Usage is implied by the resource name rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_verification_gatesA
Retrieve the five pre-emission verification gates (G1 Executability, G2 Completeness, G3 Correctness, G4 Dependency Honesty, G5 Problem Fit) and the prohibited-marker checklist. Run these gates before delivering any code output.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden, but for a zero-parameter static content retrieval the risk surface is minimal. The verb "Retrieve" implies a read-only, side-effect-free fetch, and the description discloses both the returned content (five gates plus the prohibited-marker checklist) and the required workflow behavior. It adds no auth, rate-limit, or pagination notes, but none are plausibly needed here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler: the first front-loads what is retrieved and enumerates the gates, the second states the imperative usage. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, annotation-free retrieval tool with no output schema, the description covers what an agent needs: the resource identity, the exact contents (five named gates plus the checklist), and when to invoke it. Nothing actionable is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate at the argument level; the baseline for a no-param tool is 4. The description adds no spurious parameter talk, which is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (Retrieve) and resource (five pre-emission verification gates) and even enumerates the gate names, so the agent knows exactly what payload to expect. It does not explicitly distinguish itself from similarly-named sibling retrievers like get_constitution or get_knowledge_area, which keeps it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Run these gates before delivering any code output" states a clear triggering condition and lifecycle placement, which is genuinely useful routing guidance. It stops short of naming alternatives or exclusions (e.g., when not to fetch gates), so it is a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_patchA
Review a unified diff (git diff / gh pr diff output) against the Zero-Framework-Tolerance rules. Every changed code file is scanned with the full engine union (CodebaseCSI, constitution prose rules, Python AST, tree-sitter xast) and findings are attributed to the lines the change introduces; pre-existing debt in touched files is counted separately and never fails the gate. Supply files (path -> full new-version content) for full-fidelity AST analysis; without it, added lines are scanned as fragments. Returns JSON: verdict (FAIL/REVIEW/PASS), per-file findings with new-file line numbers, totals, and the engines that ran. Deterministic and offline: fetch the diff yourself (e.g. gh pr diff) and post reviews yourself. A PASS gates nothing but markers - gates G1-G5 remain the reviewer's responsibility.
| Name | Required | Description | Default |
|---|---|---|---|
| diff | Yes | Unified diff text, e.g. from `git diff` or `gh pr diff`. | |
| files | No | Optional map of changed-file path (new version) to its complete file content, enabling full-file AST analysis instead of fragment scanning. | |
| exclude | No | Optional fnmatch globs for paths the gate must not judge: detector rule definitions and test fixture corpora contain hunted patterns as data. | |
| test_globs | No | Project-specific test path globs added to the built-in conventions (tests/*, test_*.py, *.spec.ts, *_test.go, ...). | |
| require_tests | No | C-03 enforcement: flag production-logic changes when the diff touches no test files. 'warn' surfaces them for judgement, 'fail' gates the verdict. Default off; the reviewing agent should normally pass 'warn'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden and does so: deterministic and offline, no fetching or posting side effects, pre-existing debt counted separately and never fails the gate, and an explicit caveat that a PASS gates nothing but markers while G1-G5 remain the reviewer's responsibility. This is unusually complete behavioral context for a mutation-free analysis tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and verdict semantics, then parameters, then caveats; every sentence carries information (engines, attribution rules, offline contract). It is dense and longer than average, but the length is justified by the tool's complexity rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the return payload: verdict (FAIL/REVIEW/PASS), per-file findings with new-file line numbers, totals, and engines that ran. Combined with the offline/gating caveats, an agent has everything needed to invoke and interpret the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so baseline is 3, but the description adds rationale beyond the schema: `files` as path -> full new-version content for AST analysis, `exclude` to keep detector definitions and fixture corpora out of judgement, and a recommendation on require_tests. Only `test_globs` gets no extra prose beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Review a unified diff') and immediately scopes the domain: Zero-Framework-Tolerance rules, engine union, line-level attribution. It is clearly distinguishable from siblings like scan_code_for_violations, which does not operate on a diff.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete operational conditions: supply `files` for full-fidelity AST analysis vs. fragment scanning without it, normally pass require_tests='warn', and fetch the diff / post reviews yourself. It does not explicitly compare against the sibling tools (e.g., when to prefer scan_code_for_violations over reviewing a diff), so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_code_for_violationsA
Statically scan code for Zero-Framework-Tolerance violations: TODO/FIXME markers, stub bodies (pass, ellipsis, NotImplementedError, unimplemented macros, panic stubs), empty function and catch bodies across Python, JavaScript, TypeScript, Java, Go and Rust, scaffold deception phrases ('rest of the implementation', 'omitted for brevity'), and iteration-deferral phrases ('left as an exercise', 'you can extend this'). Backed by the CodebaseCSI forensic detector plus Constitution prose rules; Python input additionally gets AST analysis so abstract stubs (Protocol/ABC/@abstractmethod) and markers inside string literals are not falsely flagged. Returns a JSON verdict (PASS/REVIEW/FAIL) with line-numbered findings. Use before delivering generated code. A PASS is necessary but not sufficient - still run gates G1-G5.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The complete code to scan. | |
| language | No | Optional language hint, e.g. 'python', 'javascript', 'typescript', 'java', 'go', 'rust'. Selects the string-masking strategy and enables Python AST analysis. Omit to infer automatically. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the underlying detector (CodebaseCSI) plus prose rules, the Python-specific AST behavior that prevents false positives on Protocol/ABC/@abstractmethod and markers in strings, and the exact return shape (PASS/REVIEW/FAIL with line numbers). This is unusually rich behavioral disclosure for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One long but densely packed sentence front-loads the enumerated violation classes, followed by the mechanism, return format, and usage note. Every clause earns its place, though the first sentence is heavy enough that a reader must parse carefully.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the absence of an output schema, the description explains the return value (PASS/REVIEW/FAIL verdict with line-numbered findings) and the follow-up obligation. For a two-parameter scanning tool this is complete enough for an agent to call and interpret correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both 'code' and 'language'. The description's mention of string-masking and Python AST ties to the language parameter's effect, but adds no syntax or format detail beyond the schema. Baseline 3 is appropriate when the schema does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (statically scan), resource (code), and enumerates the exact violation classes detected. An agent can immediately distinguish this from siblings like review_patch or verify_dependencies, which do not perform static violation scanning.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: 'Use before delivering generated code' and explicitly routes to a follow-up path by naming gates G1-G5 and warning that a PASS is not sufficient. It doesn't name a specific sibling tool as an alternative, but the timing trigger and hand-off are concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_dependenciesA
Verify that every external dependency introduced by code or a diff actually exists on its public registry - the mechanical half of gate G4 (Dependency Honesty). LLMs hallucinate package names and attackers register them (slopsquatting), so a missing registry entry is both an incompleteness defect and a supply-chain risk. Extracts imports (Python via AST incl. importlib/import literals; JS/TS via import/require specifiers), classifies stdlib/Node built-ins/first-party-in-diff/excluded locally, then checks the rest against PyPI (PEP 503) and the npm registry. NETWORK NOTICE: this is the only tool here that touches the network - package names and nothing else are sent over HTTPS, bounded (50 packages/call, 10s timeout, 3 attempts). Verdicts: FAIL = something does not exist (hallucinated/misspelled); REVIEW = unverifiable (offline/registry errors - never silently passed); PASS = everything resolves. Existence only: version pinning and integrity stay with the reviewer.
| Name | Required | Description | Default |
|---|---|---|---|
| code | No | Source text to extract imports from (requires 'language'). | |
| diff | No | Unified diff; imports are extracted from added lines of Python/JS/TS files with new-file line numbers. Exactly one of 'code' and 'diff'. | |
| exclude | No | fnmatch globs for package names that must never be sent to a registry (private/internal packages). | |
| language | No | Language of 'code'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses that it is the only networked tool, exactly what data leaves (package names only, over HTTPS), hard limits (50 packages/call, 10s timeout, 3 attempts), and the three verdict semantics including the safety-critical 'REVIEW = unverifiable... never silently passed'. It also states the deliberate scope boundary (existence only).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose then rationale, extraction mechanics, the network notice, the verdict legend, and the scope limit. Every clause is information-bearing, though the parenthetical on importlib/__import__ literals and the enumerated caps make it denser than strictly necessary for selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain returns and does so via the PASS/FAIL/REVIEW verdict legend, while also covering extraction sources, network exposure, and the explicit exclusion of version pinning and integrity. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 and the schema already documents 'code' requiring 'language', the diff/new-file line-number behavior, and the fnmatch glob semantics of 'exclude'. The description reinforces the security intent of 'exclude' (names that must never be sent to a registry) but adds no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (verify external dependencies exist on their public registry) and pins it to a named gate (G4, Dependency Honesty). It also explicitly distinguishes itself from every sibling by declaring it is 'the only tool here that touches the network', so an agent can route to it without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to reach for it (the mechanical half of gate G4, after code or a diff introduces imports) and scopes what it is not for (version pinning and integrity 'stay with the reviewer'). It does not name a specific alternative sibling or an explicit when-not-to-use branch, which keeps it just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v3.6.0- First observed
get_constitution - First observed
get_knowledge_area - First observed
get_power_of_10 - First observed
get_verification_gates - First observed
review_patch - First observed
scan_code_for_violations - First observed
verify_dependencies
TDQS
Scored across 7 tools
The three verification tools (verify_dependencies, scan_code_for_violations, review_patch) target clearly different inputs — package manifests, standalone code, and a diff — and the retrieval tools map to distinct documents. There is mild overlap since get_constitution can return everything that get_knowledge_area, get_power_of_10, and get_verification_gates offer separately, and review_patch subsumes scan_code_for_violations for changed lines, but the descriptions make the boundaries workable.
Names are readable and use consistent snake_case. Four retrieval tools form a clean get_* family, but the three action tools switch to verify_/scan_/review_ prefixes, which is a mild convention split though still sensible verb_noun naming.
Seven tools is well-scoped for a governance persona: four document retrievers plus three verification actions. Each tool clearly earns its place with no filler.
The surface covers the constitution's content areas (SWEBOK KAs, Power of 10, gates) and the mechanical verification half of gate G4 plus Zero-Framework-Tolerance scanning for both files and diffs. Gates G1, G2, G3, and G5 have no tool support and are explicitly deferred to the reviewer, a minor but acknowledged gap.
Maintenance
Related MCP Connectors
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
Repository knowledge graph MCP server for codebase understanding and debugging.
MCP server for secureFlows: token-free URL builders and integration-linting tools for AI agents.
Related MCP Servers
- AlicenseAqualityCmaintenanceArchitecture governance MCP server for AI-built codebases, enabling health checks, template management, and project scaffolding with migration support.57 npmMIT
- AlicenseNot gradedqualityCmaintenanceLocal-first MCP server that provides project context, verification gates, and structured tools for coding agents to discover knowledge, run diagnostics, and execute allowlisted commands within a repository.9 npmMIT
- AlicenseNot gradedqualityBmaintenanceMCP server that provides coding agents with structured repository context, including graph-based navigation, dependency analysis, runtime flow tracing, and configuration surface across supported stacks.57 npmMIT
- AlicenseNot gradedqualityAmaintenanceAn MCP server that enforces repository governance rules for AI coding agents, providing tools to validate plans, diffs, and scan for architectural and safety violations.1MIT