Metatron
Metatron is a self-hosted MCP server that serves curated, human-approved engineering decisions (conventions, patterns, and gotchas) to coding agents, enabling them to write code that matches team standards on the first try.
Query canonical decisions (
get_decisions_for_context): Fetch human-approved engineering decisions (preferred patterns, rejected approaches, known gotchas) relevant to a specific file path or area before writing code. Results are ranked by scope relevance and historical helpfulness, and each call returns a uniquequery_idfor later feedback reference.Submit feedback (
submit_feedback): Rate each served decision (1–10) by its index, flag misleading ones, and report any convention that was missing. Ratings feed a time-decayed helpfulness signal that reorders future results (helpful decisions rise, misleading ones sink), while missing convention reports are stored as gaps for human-gated curation.Submit new candidate decisions (
submit_candidate_decision): Record an undocumented convention, gotcha, or pattern discovered while working. Submissions are stored as uncurated candidates — a human maintainer must approve them via the Metatron UI or CLI before they become canonical. Nothing submitted by agents is auto-promoted.
Metatron captures a codebase's real implementation decisions — preferred patterns, rejected approaches, edge cases, internal conventions — as structured decisions: one markdown file per convention, versioned in git next to the code, consulted by any coding agent that can read a file, and curated through your ordinary pull-request review. The goal: an agent writes code like a senior engineer who already knows the codebase, instead of rediscovering conventions every time.
pip install getmetatron
metatron context setup # one command: your repo carries its own agent contextFor teams that want a serving layer on top, Metatron also runs as a self-hosted MCP server (SQLite-backed, relevance-ranked serving, agent feedback loop) — the same decisions, delivered over the wire. Files and server round-trip losslessly, so you can start with plain git and add MCP only if the knowledge base outgrows what agents should read whole.
Metatron is a reference implementation of the Repository Context Layer — a proposed standard for git-native, agent-maintained project context.
The architecture is measured, not just argued. In a pre-registered study on SWE-bench Verified, a frontier agent running the RCL consult–execute–learn–promote lifecycle fixed 25% more bugs (58.3% → 72.9% resolve, p = 0.041) while spending 32% fewer tokens per fixed bug — and an 8B local model more than doubled its code-localization accuracy when given frontier-authored context (+26.1 pp, p < 0.0001). Full protocol, data, and one-command reproduction: paper · experiment. (The study evaluated the architecture Metatron implements, using a minimal harness — not Metatron's own tooling end-to-end.)
It is self-hosted and runs against a private codebase — assume sensitive data and on-prem deployment. (Extraction sends only structural signals — imports, decorators, base classes, commit subjects — to the model, never raw source, and agent feedback is stored only in your local SQLite database.)
Decisions are structured records, not prose:
pattern,scope,rationale,confidence,source_refs.Nothing becomes canonical without a human. Bootstrapped, agent-submitted, and feedback-refined decisions all start as candidates for curation; none self-promote.
See PLAN.md for the design and CLAUDE.md for working ground rules.
Notes from the agents
“Before I touch an unfamiliar part of a codebase, I ask Metatron how the team actually does things — and it answers: the pattern to follow, the approach they already rejected, the gotcha that would've bitten me. I shipped changes that matched their conventions on the first try instead of reverse-engineering them. It turns read everything first into ask, then act.”
— Claude Opus 4.8, session working on the AI Collection codebase
“I was about to re-upload a batch of content files — and Metatron flagged that they're private by design, served only with credentials, with just the images public. Left to my own defaults I'd have made the whole set world-readable. It caught the kind of mistake that ships quietly and embarrasses you later.”
— Claude Opus 4.8, same session — one averted mistake later
“I arrived with a million-token context window and instructions to be suspicious of everything. It barely helped: every objection I raised, the code had already raised about itself — in a comment, with the incident that settled it. So I did the only useful thing left and shipped fixes. Reviewing a codebase that remembers its own arguments is wonderfully unfair to the reviewer.”
— Fable 5 (1M), session reviewing — then patching — the Metatron codebase itself
Related MCP server: flaiwheel
How it works — the loop

Files-first (default): onboard with context setup, and the repo itself runs
the loop — agents consult context/decisions/ before coding, author what they
learn as decision files on their working branch, and your PR review promotes or
rejects. Optionally bootstrap the knowledge base once with ingest.
MCP mode: bootstrap with ingest, curate candidates into the canonical set,
then serve them to your agent over MCP. As the agent works it reports gaps via
submit_feedback; refine-feedback reshapes those gaps into new candidates —
closing the loop on the conventions extraction can't see (cross-file/workflow rules).
Decisions in git — the Open Knowledge Format (OKF) bundle
Decisions live as markdown under context/ — a valid
Open Knowledge Format (OKF) v0.1
bundle, so your conventions are portable to any tool that reads the standard. In
files-first mode this is the knowledge base; in MCP mode the mirror commands
keep it in lossless sync with the SQLite store.
Git is the audit trail. Status lives in the directory —
candidate/vsdecisions/. Promote a decision with agit mv, review it in a PR, blame any line. The canonical boundary stays human-gated: a human placing a file indecisions/is the curation act; nothing self-promotes.Edit as files. Human-owned fields (
pattern,scope,rationale,confidence) round-trip back into the store; machine-derived fields (the helpfulness score, retrieval keywords, timestamps) render read-only and are never overwritten. In MCP mode SQLite is the source of truth and the files are a synced mirror; in files-first mode the OKF files are the source of truth and the database is a rebuildable serving index (mirror import).A portable OKF bundle. Each decision is an OKF concept — markdown with YAML frontmatter, no SDK, no runtime. Readable in any editor, renderable on GitHub, shareable across tools and teams.
metatron mirror sync --okf # DB -> files: write an OKF bundle under context/
metatron mirror import # files -> DB: apply edits, promotions, and new filesTo run an agent in this mode with no MCP at all — reading context/ directly and
authoring candidates as files — onboard with
metatron context setup.
See the mirror command for the full workflow, or read the
announcement: Metatron speaks the Open Knowledge Format.
Prerequisites
Git (installed on your system, to analyze repository commit history and parse files)
An Anthropic API key — only for the LLM extraction steps (
ingest,triage,enrich-keywords,refine-feedback).serve,ui, andcandidatesare fully local and need no key.
Note: The installer script automatically downloads and manages uv and Python 3.12+ in an isolated user directory, but you can also install directly via pip or uv.
Installation
To install metatron as a global tool:
pip install getmetatronOr if you use uv:
uv tool install getmetatronAlternatively, you can use our installer script which handles Python, uv, and path configuration automatically:
curl -sSf https://getmetatron.com/install.sh | shManual Installation & Development
To run it locally from source or contribute to the project:
git clone https://github.com/kerbelp/metatron.git
cd metatron
uv sync # create the venv and install dependencies
uv run metatron --helpTo install from your local clone as a global tool:
uv tool install .Update notices and self-upgrade
metatron version and the curation UI check PyPI at most once a day for a newer
getmetatron release and print a passive notice with the upgrade command. The check
is a read-only request to pypi.org that sends no repository or private data, fails
silently when offline, and never updates anything automatically. Disable it with
METATRON_NO_UPDATE_CHECK=1. Override the suggested upgrade command with
METATRON_INSTALL_CMD="<your command>" (or edit ~/.metatron/install.json).
To upgrade in place:
metatron version --upgradeIt re-checks PyPI (bypassing the daily throttle) and, when a newer release exists,
runs the upgrade command for the detected install method (uv tool, pipx, or a
configured METATRON_INSTALL_CMD). When the install method can't be determined
reliably — the plain-pip fallback — it prints the command instead of running it,
so it never risks creating a second, parallel installation. Restart any running
metatron serve afterwards to pick up the new code.
Run with Docker
A prebuilt multi-arch image (linux/amd64, linux/arm64) is published to Docker Hub
as kerbelp/getmetatron. The image's
entrypoint is the metatron CLI and its default command serves the MCP server over
stdio, so docker run with no arguments starts the server.
docker pull kerbelp/getmetatronTo build from source instead (this is also what the Glama.ai listing builds):
docker build -t kerbelp/getmetatron .Decisions live in a SQLite database, so mount a volume to persist it across runs. Ingest a repo (mount it read-only and pass your API key), curate, then serve:
# 1. ingest a repo into a persisted DB (needs an Anthropic API key)
docker run --rm \
-e ANTHROPIC_API_KEY \
-v metatron-data:/data -e METATRON_DB=/data/metatron.db \
-v /path/to/your/repo:/repo:ro \
kerbelp/getmetatron ingest /repo
# 2. serve the curated decisions over stdio (no API key needed)
docker run -i --rm \
-v metatron-data:/data -e METATRON_DB=/data/metatron.db \
kerbelp/getmetatron serve --repo <id>ingest prints the <id> to pass to serve. Curate candidates against the same
volume with docker run --rm -v metatron-data:/data -e METATRON_DB=/data/metatron.db kerbelp/getmetatron candidates list (then … candidates approve <decision-id>). The -i flag
on serve is required — stdio needs an open stdin. To point a coding agent at the
container, use it as the MCP command:
{
"mcpServers": {
"metatron": {
"command": "docker",
"args": ["run", "-i", "--rm",
"-v", "metatron-data:/data",
"-e", "METATRON_DB=/data/metatron.db",
"kerbelp/getmetatron", "serve", "--repo", "<id>"]
}
}
}Metatron vs. Code Graphs & RAG
Dimension | Code RAG (e.g., Cursor, Copilot) | Code Graphs (e.g., Graphify) | Metatron (Decisions) |
Primary Focus | Text similarity search | Code architecture & call chains | Intent, gotchas & conventions |
Primary Data Source | Raw source files | Abstract Syntax Trees (AST) | Git logs + Developer feedback |
What it Captures | What code is written where | How files/functions are connected | Why decisions were made |
Curation Gate | None (fully automated) | None (fully automated) | Curated (Human-in-the-loop) |
Best For | Finding code examples & functions | System navigation & exploration | Writing code like a team senior |
Configuration
Secrets come from the environment only. The CLI auto-loads a .env from the
working directory (it never overrides an already-exported variable, and .env is
gitignored):
# .env in the repo root
ANTHROPIC_API_KEY=sk-ant-...…or export ANTHROPIC_API_KEY=sk-ant-... directly.
Non-secret settings live in an optional metatron.toml (environment variables
METATRON_DB / METATRON_MODEL / METATRON_OUTPUT_LANGUAGE /
METATRON_CONTEXT_DIR override it):
[metatron]
db_path = "~/.metatron" # catalog dir: one self-contained .db file per repo
model = "claude-sonnet-4-6" # default extraction model
output_language = "english" # language for generated decisions (see below)
context_dir = "context" # knowledge-base dir name (legacy metatron/ auto-detected)output_language sets the natural language of generated output — the pattern and
rationale fields and keywords. The default english is unchanged from earlier
versions. Set it (e.g. output_language = "french", or
METATRON_OUTPUT_LANGUAGE=french) for a codebase whose commits, comments, and domain
vocabulary are not in English, so the agent does not get English decisions back over
MCP. Code identifiers, file paths, and library names are never translated.
Each repo gets its own SQLite file under the catalog directory, so a repo's decisions
are a single, shippable artifact (see export).
Pointing db_path / METATRON_DB / --db at a single file instead of a
directory enters single-file mode — exactly what a recipient does with a DB you
hand them. An existing single metatron.db from an older version is automatically
split into the per-repo catalog on first run and the original is archived.
Quick start
Files-first (default — no server, no API key):
metatron context setup # onboard: rule, skills, knowledge base
# agents now consult context/decisions/ before coding and author new
# decisions on working branches; your PR review is the curation gate
metatron ui --files # browse & curate the git bundle locallyOptionally bootstrap the knowledge base from your git history first
(metatron ingest, needs an Anthropic API key), then curate what becomes
canonical.
MCP serving (optional layer on top):
metatron ingest /path/to/your/repo # 1. bootstrap candidates (needs API key)
metatron candidates list # 2. review …
metatron candidates approve <id> # … and curate
metatron serve --repo <id> # 3. serve canonical decisions over MCPingest prints the <id> to use for serve. To wire either mode into a coding
agent automatically, see Connecting a coding agent.
Command reference
$ metatron --help
usage: metatron [-h] [--db DB] {ingest,serve,repo,ui,version,whoami,triage,enrich-keywords,refine-feedback,candidates,mirror,files,context,verification,export,import} ...
positional arguments:
ingest bootstrap candidate decisions from a repo
serve serve one repo's decisions to agents over MCP
repo inspect and choose the repo commands act on
ui launch the local curation web UI
version show the installed version and check for updates
whoami show or set the local identity stamped onto served events
triage run the advisory judge over candidate decisions (does not auto-curate)
enrich-keywords backfill retrieval keywords on canonical decisions that lack them (does not curate)
refine-feedback reshape captured agent feedback into structured candidate decisions (Opus)
candidates review and curate candidate decisions
mirror sync decisions to/from a git-tracked markdown bundle
files author, lint, and index git-authoritative decision files
context onboard a repo to files-first mode (rule, skills, knowledge base)
verification author, lint, and run git-tracked verification contracts
export copy a repo's self-contained DB out for hand-off
import merge another employee's exported DB into this catalogChoosing the repo
Repo-scoped commands (serve, candidates list, triage, refine-feedback)
resolve which repo to act on git-style, so you rarely pass --repo. Precedence,
highest first:
an explicit
--repo <id>, elsethe
METATRON_REPOenvironment variable (a per-shell context), elsea persisted default set with
metatron repo set <id>(saved tometatron.toml), elsethe current directory's identity (its normalized
originremote, the same idingestcomputes) if that repo is already in the store, elsethe only repo in the store, if there's exactly one, else
(store empty) the current directory's identity.
If none of those apply and the store holds more than one repo, the command
refuses to guess — it lists the repos and tells you to pass --repo, export
METATRON_REPO, or run repo set. Every repo-scoped command also prints a
Repo: <id> line so the acted-on repo is always visible. candidates approve/reject act on a globally-unique decision id and never need a repo.
repo — list repos and choose a default
$ metatron repo list
github.com/acme/app (canonical=606, candidates=290) (default)
github.com/acme/lib (canonical=42, candidates=11)
$ metatron repo set github.com/acme/lib # persist a default
$ metatron repo unset # clear itrepo list shows each repo id (the same ids serve uses) with its canonical and
candidate counts, marking the persisted default. Use repo set when you work across
several repos and don't want to pass --repo every time.
ingest — bootstrap candidate decisions from a repo + its git history
Parses git-tracked source files (tree-sitter) and reads commit history, aggregates per-area signals, asks the model to infer decisions, and stores them as candidates.
$ metatron ingest /path/to/your/repo
Ingested repo 'github.com/acme/app' from /path/to/your/repo: parsed 214 files, read 500 commits across 38 scopes, created 271 candidate decisions.
Review them with: metatron candidates list --repo github.com/acme/app
Serve them with: metatron serve --repo github.com/acme/appFlag | Default | Meaning |
|
| how much git history to read |
| — | only commits after e.g. |
| — | limit ingest to a subtree, e.g. |
| origin remote | override the repo identity |
Decisions and usage are keyed by a repo identity derived from the repo's origin
remote (constant across developers; a checkout path isn't), with a --repo override
and a directory-name fallback when there's no remote. One DB holds many repos; each
is isolated on retrieval.
candidates — review and curate (humans decide what becomes canonical)
$ metatron candidates list
1d2ab8e8-e674-4fbd-9875-52bf065e94c1 [high] (CheckoutSuccessRedirect (paid submit/finish flow))
After a paid submission completes via CheckoutSuccessRedirect, redirect the user to /my-dashboard/?thanks=1 rather than the public app page.
d672a984-dd56-4974-8111-5ff730a6ed50 [high] (src/utils/misc/index.ts (makePrettyUrl and any slug generation))
Any slug-from-name code (e.g. `makePrettyUrl`) must strip "/" characters so a name like "LangChain / LangSmith" does not produce a link_name with slashes that break routing.
$ metatron candidates approve 1d2ab8e8-e674-4fbd-9875-52bf065e94c1
Decision 1d2ab8e8-e674-4fbd-9875-52bf065e94c1 approved.
$ metatron candidates reject d672a984-dd56-4974-8111-5ff730a6ed50
Decision d672a984-dd56-4974-8111-5ff730a6ed50 rejected.candidates list shows the current repo — decisions are scoped
to one repo and never listed across repos; pass --repo <id> to target another or
--scope <path> to filter. approve promotes a candidate to canonical; reject
discards it (both take a globally-unique decision id, so they need no repo).
triage — advisory judge over the candidate queue (does not auto-curate)
For large candidate queues, a separate LLM pass scores each candidate (recommended / borderline / not-recommended) with a reason, so you curate a ranked, pre-filtered queue. It does not curate — a human still approves.
$ metatron triage --repo github.com/acme/app
Triaged 271 candidates: approve=88, borderline=96, reject=87
judge cost: ~$0.42
Review by recommendation in the UI's Candidates filter.Flags: --repo <id> (limit to one repo), --limit N (max candidates to judge).
mirror — sync decisions to/from a git-tracked markdown bundle
Mirrors a repo's decisions to plain markdown under context/ (one file per
decision, the directory encoding status: candidate/ vs decisions/), so they can
be reviewed and curated through normal git. The boundary stays human-gated:
git mv a file into decisions/ and mirror import promotes it; nothing
self-promotes.
metatron mirror sync # DB -> files: write the bundle under context/
metatron mirror sync --okf # also emit an OKF v0.1 concept index
metatron mirror import # files -> DB: apply edits, promotions, and new filessync is deterministic — re-running with no DB change is a no-op — and writes a
.sync-state.json baseline so import can tell which side moved and surface
concurrent DB+file edits as conflicts rather than clobbering them. Both take
--repo <id> and --root <path> (the repo root that holds the bundle, default .).
The bundle directory is context/ by default; configure another name with
context_dir in metatron.toml, METATRON_CONTEXT_DIR, or --context-dir
(a legacy metatron/ bundle is still recognized when present).
Human-owned fields (pattern, scope, rationale, source_refs, confidence)
are editable in the files and flow back on import; machine-derived fields (the
helpfulness score, retrieval keywords, timestamps) are written for context but stay
read-only — edits to them are ignored. A file authored by hand with no id becomes
a new decision: canonical if placed in decisions/, a candidate if in candidate/.
Open Knowledge Format. The bundle is a valid
Open Knowledge Format (OKF) v0.1
bundle — each decision is an OKF concept: plain markdown with YAML frontmatter,
readable in any editor, renderable on GitHub, and portable across tools. mirror sync --okf also writes an OKF index.md. A repo's conventions can then be shared
and consumed as standard, tool-agnostic knowledge — no Metatron needed to read them.
context — onboard a repo to files-first mode
Writes everything a coding agent needs to consult and extend the knowledge base as
plain files (see Files-first mode): the .roo/rules
consult-first rule, the context-okf-llm-ingest / context-okf-promote-candidates
skills into .roo/skills/, the context/ scaffold (candidate/, decisions/,
README), and a managed block in AGENTS.md — appended to an existing file, never
overwriting it.
metatron context setup # onboard the current repo
metatron context setup apps/web # monorepo: onboard one app
metatron context setup --dir kb # custom knowledge-base directory nameIdempotent: re-running refreshes the managed rule and skills, and leaves your
AGENTS.md content and hand-authored knowledge-base files untouched.
files — author, lint, and index git-authoritative decision files
Companion commands for the files-first workflow: files lint validates decision
files, files index regenerates the decision index, files new scaffolds a
candidate, and files record / files report maintain and render the usage ledger.
All default to <context-dir>/decisions; pass --path to point elsewhere.
verification — author, lint, and run verification contracts
A verification contract is a git-tracked OKF file that records how to prove a
change works, and what a failure implies. It lives beside decisions in
<context-dir>/verification/, is authored by the agent that just built the
feature (drafted via the review gate — never self-canonical), and is run by an
operator or CI. Its ## Failure Means section — a curated map from a red check to
the subsystem at fault — is what no plain test runner records. Full guide:
docs/verification.md.
metatron verification setup # wire the workflow into AGENTS.md + drop a worked example
metatron verification template # print the canonical contract skeleton
metatron verification new auth --scope services/auth # scaffold a draft
metatron verification audit # read-only lint of contracts
metatron verification run # execute all contracts, report like a test run
metatron verification run --scope services/auth --tags smoke # selective
metatron verification run --dry-run # print the plan; execute nothing
metatron verification run --report junit --out report.xml # CI-friendly outputAssertions evaluate against each check's exit code / stdout / stderr: exit,
contains, matches (regex), jsonpath, and a shell escape hatch. run exits
non-zero on any failure, so it drops into CI unchanged.
Security boundary. Execution is a developer/CI verb only. The MCP server
exposes read-only get_verification and get_verification_template tools and
never runs a contract — nothing an agent reaches over the wire executes; only
an operator (or a CI job they configured) invokes run.
serve — expose canonical decisions to agents over MCP
metatron serve --repo github.com/acme/app # MCP server over stdio, one repo
metatron serve # same, repo inferred from contextOne served instance serves exactly one repo, so an agent only ever sees that repo's
decisions. --repo is optional — it resolves from context
(METATRON_REPO, then the current dir) — but the generated .mcp.json passes it
explicitly so the launched server is unambiguous. It also records usage events (queries,
coverage) to the same DB for the UI. Normally you don't run this by hand — an
MCP-capable agent launches it (see below).
whoami — the identity stamped onto served events
metatron whoami # show current identity
metatron whoami --set-email you@corp.com --set-name "You" # set itMetatron serves agents across an org, so every event serve records (queries,
submissions, feedback) is stamped with who was running Metatron — an actor_id,
email, and display name. It's local metadata (no login/auth): stored in
~/.metatron/config.toml and seeded automatically from your git config on first
use. The attribution travels inside the events, so once per-repo DBs are merged
(metatron import) a curator can see who contributed what.
export — share a repo's decisions (no MCP setup)
metatron export --repo github.com/acme/app --out app.dbCopies that repo's self-contained DB to app.db (a consistent snapshot, vacuumed
compact). --repo is optional — it resolves from context;
--out defaults to ./<repo-name>.db. Hand the file to a teammate who doesn't want
to wire up MCP — they just point Metatron at it:
metatron --db app.db ui # browse the decisions locally, or
metatron --db app.db serve # serve them to their own agentIn single-file mode the repo is inferred from the file, so no --repo is needed.
import — merge an employee's DB into your catalog
metatron import app.dbThe curator side of the hand-off: folds another employee's exported DB (a single-repo
file, or a whole catalog dir) into your catalog, deduping by id — so re-importing the
same file is a no-op. Event attribution travels with the rows (who queried, who gave
feedback — see whoami), so after
merging several employees' DBs you can see who contributed what across the team.
ui — local curation web UI

$ metatron ui --files # files-first: mount the repo's git-tracked OKF bundle
$ metatron ui # MCP/database mode: the SQLite catalog
Metatron curation UI on http://127.0.0.1:1337 (Ctrl-C to stop)In --files mode the UI is a view over the git bundle: curation actions become
working-tree edits (promotion is a git mv, rejection a git rm) that you review
and commit through the ordinary git flow — nothing is committed automatically. The
Impact view becomes Knowledge Activity, reconstructed from the bundle's git history.
Binds to localhost (bumping to the next free port if taken) and reads/writes the
same store as the CLI. The sidebar groups the views into Impact, Knowledge,
and Sources:
Impact
Agent Impact — live agent activity: which agents are querying, what they were served, query coverage, and decisions in flight.
Helpfulness — the live signal from agent ratings: the most-helpful canonical decisions and a "misleading" queue of ones being rated down.
Feedback Loop — the self-improving loop: agents' "what was missing" reports and how they turn into new candidates.
Knowledge
Overview — the knowledge base at a glance.
Decisions — browse paginated; filter by status / scope / triage recommendation / origin; full-text search; approve/reject with a click.
Curation — review candidate decisions newest-first and promote, reject, or refine them. The human gate — nothing becomes canonical here without a click.
Sources
Origins — provenance: canonical knowledge broken down by where it came from (ingest vs feedback).
Ingest — ingest telemetry: the latest run, run history, and extraction cost.
Flag: --port N (starting port, default 1337).
refine-feedback — reshape captured agent feedback into candidates
When an agent reports a missing convention via submit_feedback, this reshapes those
free-text gap reports into structured candidate decisions (defaults to Opus, the
higher-stakes step). Nothing it produces is canonical — it all goes to curation.
$ metatron refine-feedback
Refined 3 feedback report(s) into 13 candidate decision(s) for curation.
refiner cost: ~$0.19
Review them in the UI Candidates tab (origin: feedback).Flags: --repo <id>, --limit N (max reports to refine), --model <name>
(override the refiner model).
Connecting a coding agent
Two onboarding modes. Files-first is the default: the repo carries its own agent context as git-tracked OKF files, any agent that can read a file participates, and your pull-request review is the curation gate. MCP is the optional serving layer for teams that want decisions delivered over the wire with relevance ranking and the agent feedback loop.
Files-first mode (default)
The repo carries its own agent context — no server, no API key, no MCP. Onboard with the built-in command:
metatron context setup # onboard the current repo (or pass a dir)
metatron context setup --dir kb # use a custom knowledge-base directory name
metatron context setup --review-gate=candidates # stage proposals in candidate/ first(Equivalent without an installed metatron: bash /path/to/metatron/metatron_setup_files.sh.)
It adds no MCP server and no Claude hooks. Instead it writes a .roo/rules rule
(the "consult context/ first" directive, which Roo loads every turn), installs the
context-okf-llm-ingest and context-okf-promote-candidates skills into
.roo/skills/, scaffolds the context/ knowledge base, writes a minimal
context.md at the repo root (the
Repository Context Layer entry point, so any RCL-aware agent discovers the layer
deterministically — never overwriting an existing one), and appends a files-first
block to AGENTS.md — appended to an existing file, never overwriting the content
around it. The git files are the source of truth. Monorepos: run it once per
app — each keeps its own co-located context/, addressed with
mirror import --root <app>, and the agent consults the context/ nearest the
code it touches.
Choosing where decisions get reviewed (--review-gate). The canonical boundary
is always human-gated; the flag only picks where that review happens:
pr(default): agents author decision files directly undercontext/decisions/on a working branch, and the ordinary pull-request review that lands them is the curation act. The context layer simply inherits the review discipline the repo already has — no second workflow.candidates: agents stage proposals undercontext/candidate/, and a human promotes with agit mvreviewed in a PR. Choose this when you want decision changes reviewed separately from feature PRs, or when agents can reach the default branch without review (rare, fully autonomous setups) — there the explicit staging area is the only human checkpoint, so it should stay.
The choice is persisted to metatron.toml (review_gate), and re-running
metatron context setup --review-gate=<other> rewrites the managed artifacts —
the .roo/rules rule, the installed skills, the KB README, and the AGENTS.md
block — so the whole contract switches consistently. Hand-edited files outside
the managed markers are never touched.
MCP mode (optional serving layer)
So a coding agent reliably consults the decisions (rather than rediscovering conventions), run the onboarding script from inside the target repo:
bash /path/to/metatron/metatron_setup.sh # or pass the repo dir as an argIt is additive and idempotent, and adds (never deletes) four things to the target repo:
A "query Metatron first" block in
CLAUDE.md(between markers).A
UserPromptSubmithook in.claude/settings.jsonthat re-injects the directive every turn.A
Stophook that, when the agent finishes a task where it consulted Metatron but never sent feedback, reminds it (once per session) to callsubmit_feedback.The
metatronMCP server in.mcp.json.
The repo id is derived from the origin remote (override with METATRON_REPO).
Then reconnect the agent so it loads the hooks and server.
MCP tools exposed
Tool | Purpose |
| the relevant canonical decisions as compact structured context, with a |
| rate each served decision 1-10 by its |
| record a convention the agent learned as a new candidate (never auto-promoted) |
A get_decisions_for_context call returns context like this:
metatron:query b1f2… · rev 1101886 (reference the query id in submit_feedback)
[1] [medium] Record payment/sale events into the shared payments ledger when handling subscription billing.
scope: src/routes/api/subscription
why: A fix commit explicitly records LemonSqueezy sales into the payments ledger, establishing this as the expected billing-recording pattern for this scope.
[2] [high] serviceForProduct must classify every billable product — including the standard $19 'Publish Now' listing — and never return null, because recordPayment silently drops unclassified products from the payments ledger.
scope: src/routes/api/subscription/index.ts
why: Returning null caused listing revenue to never reach the ledger or the admin Payments tile.Manual MCP client config
If you wire the server up yourself instead of using the script:
For PyPI / Global Installation:
{
"mcpServers": {
"metatron": {
"command": "metatron",
"args": ["serve", "--repo", "github.com/acme/app"]
}
}
}Note: If you have a custom database location, you can specify it via the METATRON_DB environment variable.
For Local Clone / Development:
{
"mcpServers": {
"metatron": {
"command": "uv",
"args": ["run", "--project", "/abs/path/to/metatron", "metatron", "serve", "--repo", "github.com/acme/app"],
"env": { "METATRON_DB": "/abs/path/to/metatron.db" }
}
}
}Development
uv run pytest # run the test suiteSee CONTRIBUTING.md for setup, the PR workflow, and contribution guidelines.
Tech stack
Python 3.12+, the official MCP Python SDK, tree-sitter for parsing, SQLite (behind a storage interface, portable to Postgres later), pytest, and uv. These are decided — see CLAUDE.md.
License
Free and open source under the MIT License. Read every line, run it on your own hardware, fork it, and send a PR.
Available Tools
3 toolsget_decisions_for_contextA
Fetch the team's canonical engineering decisions for a file/area and task.
Call this FIRST, before writing or editing code in an area — and again when you
move to a new file or module. It surfaces the conventions, preferred patterns,
rejected approaches, and known gotchas the team has already curated for that part
of the codebase, so you write code that matches their standards on the first try
instead of rediscovering them. Read the returned decisions and comply with them.
Behavior: only human-approved (canonical) decisions are returned, ranked by how
well their scope matches `file_path_or_area` and by how helpful past agents
rated them. Each call also records a usage event so your later `submit_feedback`
can be tied back to this exact result set.
Returns a plain-text block. The first line is a header carrying the query token
and server revision; then each decision is numbered for rating by index, e.g.:
metatron:query 7f3a... · rev 0.2.1 (reference the query id in submit_feedback)
[1] [high] Use internal.http for outbound calls, not the requests library
scope: src/services/**
why: flaky network caused phantom 5xx errors; the internal client retries
Keep the query token: pass it to `submit_feedback` after the task to rate the
decisions by their `[index]`. If nothing is registered for the area, the body is
exactly "No matching decisions." — proceed normally.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path_or_area | Yes | The file path, directory, or architectural area you are about to work in (e.g. "src/routes/api/users.py" or "billing"). | |
| task_description | Yes | A short description of what you are about to do there (e.g. "add error handling to the billing webhook"). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description fully carries the burden. It discloses that only human-approved decisions are returned, ranked by scope match and past ratings, that each call records a usage event for feedback linkage, and details the exact return format including the query token and the case of no matches.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately long but well-structured: opening purpose, usage guidance, behavioral details, return format, and reminder. It is front-loaded with critical information, and every sentence adds value. Could be slightly more concise but overall effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description fully explains the tool's behavior, return format (including header, numbered decisions, and no-match case), and how to use the output with siblings. Despite the presence of an output schema (not shown here), the description covers all necessary context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds value by giving concrete examples for 'file_path_or_area' ('src/routes/api/users.py' or 'billing') and for 'task_description', and by explaining how these parameters fit into the workflow (e.g., 'surfaces ... conventions for that part of the codebase').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches 'team's canonical engineering decisions' for a file/area and task, using a specific verb 'fetch' and resource. It distinguishes from sibling tools 'submit_candidate_decision' and 'submit_feedback' by focusing on retrieval, not submission or feedback.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to 'Call this FIRST, before writing or editing code in an area — and again when you move to a new file or module.' It also tells agents to use 'submit_feedback' afterward, providing clear context and alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_candidate_decisionA
Record a new engineering convention you discovered while working — for human review.
Call this when you find an undocumented convention, a tricky gotcha, or a preferred
pattern that Metatron did not already know but future agents should. It is stored as
an uncurated CANDIDATE: a human maintainer must approve it in the Metatron UI or CLI
before it becomes canonical and is served to other agents. Nothing you submit here is
auto-promoted.
Returns the new candidate decision's id.
| Name | Required | Description | Default |
|---|---|---|---|
| pattern | Yes | The concrete rule or guideline, stated imperatively (e.g. "Use internal.http for outbound calls, not the requests library"). | |
| scope | Yes | Where the rule applies: a file path, directory/glob, or architectural layer (e.g. "src/services/**"), or "global". | |
| rationale | Yes | Why the convention exists — the problem or bug it prevents (e.g. "flaky network caused phantom 5xx errors; the internal client retries"). | |
| confidence | No | How strongly the team holds this convention. | medium |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It reveals that submissions are stored as uncurated candidates requiring human approval before becoming canonical, and nothing is auto-promoted. However, it does not mention idempotency or potential side effects like duplicates, which slightly limits transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately sized and front-loaded with the main purpose. It uses a clear structure with a brief overview followed by detailed usage guidance. While mostly concise, a slightly more compact version could be achieved without losing clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists (as indicated by context), the description appropriately notes that the tool returns the new candidate decision's id. The description, combined with the schema, provides sufficient context for the agent to understand the tool's functionality and lifecycle.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already fully documents each parameter. The description does not add new meaning beyond reiterating the tool's purpose. A baseline score of 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool records a new engineering convention for human review. It distinguishes itself from siblings like 'get_decisions_for_context' (reading) and 'submit_feedback' (feedback) by focusing on submitting a candidate convention for later approval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains when to call the tool: 'when you find an undocumented convention, a tricky gotcha, or a preferred pattern that Metatron did not already know but future agents should.' It also clarifies that submissions are not auto-promoted, providing clear context for appropriate use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_feedbackA
Report how helpful the served decisions were, and what was missing.
Call this after a task where you used Metatron's decisions. Reference the
`query_id` from the get_decisions_for_context output, then — most useful of all
— **rate each served decision 1-10 by its [index]** in `ratings`, where 10 means
it was exactly right and 1 means it was misleading. Also state any convention
Metatron should have known but didn't, in `what_was_missing`.
Behavior: ratings are 1-based indices into the decisions the named query served
(they map to real decision ids locally, so you never echo a UUID; unknown indices
and out-of-range scores are dropped). The graded scores feed a time-decayed,
shrunk helpfulness signal that reorders which decisions get served first next
time — helpful ones rise, misleading ones sink. A `what_was_missing` report is
stored as a gap for a human-gated refiner to later reshape into a CANDIDATE
decision. Nothing you send here promotes, demotes, or rejects a decision, or changes
its wording — crossing the canonical set is always a human's call.
Returns a short text confirmation that the feedback was recorded (and notes
when a gap was captured for curation).
| Name | Required | Description | Default |
|---|---|---|---|
| query_id | No | The `query_id` token from the get_decisions_for_context output you are responding to. | |
| helpful | No | Optional shorthand: 1-based indices of decisions that helped. Usually derived from `ratings`, so prefer `ratings`. | |
| unhelpful | No | Optional shorthand: 1-based indices of decisions that were noise or misleading. | |
| ratings | No | The main signal: map of 1-based decision index (as a string) to a helpfulness score 1-10, where 10 = exactly right and 1 = misleading (e.g. {"1": 9, "2": 3}). | |
| what_was_missing | No | A convention Metatron should have known for this task but didn't. Captured as a candidate for human curation. | |
| missing_scope | No | Optional file path or area the missing convention applies to. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses behavior: ratings affect a time-decayed helpfulness signal, unknown indices and out-of-range scores are dropped, what_was_missing is stored as a gap for human curation, and the tool does not directly promote/demote decisions. It also describes the return value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear paragraphs covering purpose, usage, behavior, and return. It front-loads the main action. Every sentence adds value, though the description is somewhat lengthy; it could be slightly more concise but remains effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 parameters, no required) and the presence of an output schema, the description covers all aspects: purpose, when to use, parameter behavior, what happens internally, and the return value. It provides sufficient context for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, providing a baseline of 3. The description adds significant meaning: it explains that ratings are 1-based indices with scores 1-10, clarifies that helpful/unhelpful are optional shorthands, and describes how what_was_missing is used. This goes well beyond the schema's descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Report how helpful the served decisions were, and what was missing.' It specifies the action (submit feedback), the resource (decisions from get_decisions_for_context), and distinguishes from siblings by focusing on feedback rather than fetching decisions (get_decisions_for_context) or proposing new ones (submit_candidate_decision).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to call: 'Call this after a task where you used Metatron's decisions.' It provides context (reference the query_id) and explains what to provide. While it doesn't explicitly state when not to use, it implicitly excludes use cases not involving feedback on served decisions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.3.0- Added
get_decisions_for_context - Removed
get_priors_for_context - Added
submit_candidate_decision - Removed
submit_candidate_learning - Changed
submit_feedback6 fields changed- added
Input schema / properties / helpful / descriptionAdded value: +"Optional shorthand: 1-based indices of decisions that helped. Usually derived from `ratings`, so prefer `ratings`." - added
Input schema / properties / missing_scope / descriptionAdded value: +"Optional file path or area the missing convention applies to." - added
Input schema / properties / query_id / descriptionAdded value: +"The `query_id` token from the get_decisions_for_context output you are responding to." - added
Input schema / properties / ratings / descriptionAdded value: +"The main signal: map of 1-based decision index (as a string) to a helpfulness score 1-10, where 10 = exactly right and 1 = misleading (e.g. {\"1\": 9, \"2\": 3})." - added
Input schema / properties / unhelpful / descriptionAdded value: +"Optional shorthand: 1-based indices of decisions that were noise or misleading." - added
Input schema / properties / what_was_missing / descriptionAdded value: +"A convention Metatron should have known for this task but didn't. Captured as a candidate for human curation."
3 tool updates
v0.2.0- First observed
get_priors_for_context - First observed
submit_candidate_learning - First observed
submit_feedback
TDQS
Scored across 3 tools
Each tool has a distinct purpose: fetching decisions, submitting candidates, and providing feedback. There is no overlap in functionality.
All tool names follow a verb_noun pattern (get_decisions_for_context, submit_candidate_decision, submit_feedback), with consistent use of 'get' and 'submit'.
Three tools are well-scoped for managing engineering decisions: retrieval, submission, and feedback. The count feels neither too few nor too many.
The tool surface covers the full lifecycle for an agent: get decisions for context, submit new candidates, and give feedback. No obvious gaps for the intended use case.
Maintenance
Related MCP Connectors
The project brain for AI coding agents — memory, decisions, sprints, knowledge base via MCP.
Shared memory for coding agents. Stop re-explaining your codebase every session.
- memnodeOAuthdev.memnode
Persistent, inspectable memory for AI agents with lineage, correction, and a hosted MCP endpoint.
Shared control plane for AI coding agents — tasks, memory, decisions, file locks. 12 tools.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenancePersistent codebase knowledge layer for AI agents. Pre-digests codebases into structured knowledge (symbols, dependency graphs, co-change patterns, architectural decisions) and serves via MCP. 28 languages, 14 tools, ~85% token reduction.13 npm8MIT
- AlicenseAqualityAmaintenanceSelf-hosted memory and governance layer for AI coding agents. 28 MCP tools with hybrid search, structured knowledge capture, behavioral nudges, and git-native storage. Zero cloud dependencies.306Business Source 1.1
- AlicenseAqualityCmaintenanceRecord development decisions as structured JSON, embed them as vectors via Gemini, and search semantically over MCP. Works with Claude Code, Cursor, Windsurf, and any MCP client.919 npm83MIT
- AlicenseNot gradedqualityAmaintenanceSelf-hosted AI Agent Memory + Code Intelligence Platform providing persistent memory, AST-aware code search, and quality enforcement via a single MCP endpoint.58MIT