Skip to main content
Glama
dvir-shamay

mcp-convention-gate

by dvir-shamay

mcp-convention-gate

Enforce your process by construction — not by prompt. No review, no commit.

mcp-convention-gate demo

An MCP server + git pre-commit hook that refuses a commit until a required review is registered. It enforces the step outside the model's prompt and discretion — so an AI coding agent can't skip it, no matter what the instructions say.

The problem

Rules written as text — convention files, scoped instructions, even the system prompt — don't make an AI coding agent perform a costly, skippable step. In a controlled study across nine models, a "review before commit" rule was skipped in 135 of 135 trials, at every instruction layer. The identical rule, enforced by a gate that refuses the commit, is obeyed 0 of 135 non-compliant — by construction.

Related MCP server: Loopbreaker

Requirements

  • Node.js ≥ 18 — uses the built-in test runner and modern APIs.

  • git — the enforcement layer is a git pre-commit hook.

  • An MCP-capable agent/client — VS Code + Copilot, Claude Desktop, Cursor, or any client that speaks MCP over stdio — to register reviews in-loop. (The git hook still enforces even without one.)

Quickstart

In your repo:

npx mcp-convention-gate init

This installs a pre-commit hook, a .gate-config.json (which reviews are required), and an MCP server entry (.vscode/mcp.json). From then on:

  • Your AI agent registers each completed review through the MCP tools (create_gate_session, register_gate), checks gate_status until commit_allowed is true, then runs the real git commit itself — no MCP call authorizes it, the pre-commit hook enforces automatically at that point.

  • Any commit — from the agent or from a terminal — is blocked until the required reviews are registered.

  • Call guarded_commit only after that commit succeeds, purely to log it to the session's audit trail. Calling it before committing marks the session committed, and the hook will then block the commit you were about to make — see the guarded_commit row below.

npx mcp-convention-gate status   # show current gate status
npx mcp-convention-gate check    # run the pre-commit check manually

How it works

Two enforcement layers, both fail-closed:

  1. Git hook level — a pre-commit hook blocks git commit at the OS level, even from the terminal. This is the actual enforcement; nothing else needs to run for a commit to be blocked.

  2. MCP tool level — gate_status reports whether commit_allowed is true before you commit; guarded_commit is an audit-log call for after a commit succeeds, not a gate you pass through to get one.

Pick which reviews are mandatory in .gate-config.json:

{
  "enabled": true,
  "required_gates": ["security-review", "tests", "docs"]
}

The default required_gates is an opinionated nine-role review panel (spec-reviewer, architect, code-reviewer, security-auditor, test-writer, qa-acceptance, debugger, docs-writer, ux-evaluator) — that's a lot for a first commit. Most projects start lighter and grow into it; edit required_gates to match your workflow, e.g. a single gate to begin:

{ "enabled": true, "required_gates": ["code-review"] }

If .gate-config.json is missing or fails to parse, the gate is treated as disabled (not enforced) — printed as a WARNING to stderr, never silently. This favors not disrupting every commit repo-wide over silently enforcing from a file most contributors never touch directly; fix or restore the file to configure gates intentionally.

Session overrides are honored on both layers. If create_gate_session sets its own required_gates, the git hook enforces that session's list too (it checks a session's own requiredGates before falling back to .gate-config.json) — not just the MCP tools. Omitting required_gates at session creation falls back to .gate-config.json on both sides as well, so the common case (no override) stays in sync automatically. They can only diverge if you deliberately pass an ad hoc list to create_gate_session that isn't reflected anywhere else you track — an intentional escalation, not a footgun.

MCP tools

The MCP server (server.js) exposes five tools over stdio. Your agent calls them in-loop; the TOOLS definition in server.js is the source of truth.

Tool

Required args

Optional args

What it does

create_gate_session

description

required_gates[]

Opens a session for a task/sprint/PR and returns its session_id. Omitting required_gates uses the default set.

register_gate

session_id, gate_name, result (pass | fail | warn)

agent, findings[], metadata

Records that one review completed. Only pass counts toward satisfying a gate.

guarded_commit

session_id, commit_message

override, override_reason

Call this AFTER git commit succeeds, never before — it's an audit-log record, not an authorization; calling it first marks the session committed and the hook then blocks the commit you were about to make. Returns allowed if every required gate had passed at call time, else blocked. override: true (with a mandatory override_reason) records a self-attested override in the audit trail — it does not touch git; use GATE_BYPASS for that.

gate_status

session_id

—

Reports registered / missing / failed gates and whether a commit is allowed.

list_sessions

—

limit (default 10)

Lists sessions, most recent first, for auditing which tasks went through review.

A findings entry is { severity: "critical"|"high"|"medium"|"low", description, file?, line? }.

Configuring other MCP clients

init writes a ready-to-use .vscode/mcp.json for VS Code. Other clients use the same launch command — node running this package's server.js, with MCP_GATE_STORE_PATH pointed at your repo's .gate-store.json. The exact server path is whatever init wrote into .vscode/mcp.json — copy it from there.

Claude Desktop — add to claude_desktop_config.json:

{
  "mcpServers": {
    "convention-gate": {
      "command": "node",
      "args": ["/absolute/path/to/mcp-convention-gate/server.js"],
      "env": { "MCP_GATE_STORE_PATH": "/absolute/path/to/your/repo/.gate-store.json" }
    }
  }
}

Cursor — add the same block to .cursor/mcp.json (project) or ~/.cursor/mcp.json (global).

Any MCP client — launch node <server.js> over stdio and set MCP_GATE_STORE_PATH. The git hook enforces independently of the client, so even a client that never reaches the server is still gated at commit time.

Why it works — advisory vs. enforced

A prompt is a preference the model weighs. "Run a review first" competes with the immediate task ("edit and commit"), so a cooperative, action-oriented agent follows the nearest instruction and treats the distant rule as optional. Nothing stops it skipping.

The gate is code that returns failure, not a request. The model can't decide to bypass it any more than it can decide a locked door is open. It's a "please knock" sign vs. a locked door — the prompt is the sign; the gate is the lock.

"Isn't this just…"

  • …a Claude Code hook? Those are vendor-locked and gate tool calls; they don't prove a process step happened, and don't work outside Claude. This is model-agnostic — any agent that speaks MCP + git.

  • …a pre-commit hook? pre-commit checks file content (lint/format). This gates on process state ("was the review registered?") and exposes it to the agent over MCP so it can satisfy the gate in-loop.

  • …a guardrail? Guardrails filter the model's words (safety / PII / jailbreak). This gates your process, not the model's words.

Coverage (honest caveat)

The gate governs every commit that passes through it. An agent could still commit via a path the gate does not mediate (e.g. a raw shell call outside the hook, or git commit --no-verify). Making the guarded path the only path is a deployment step; within its coverage, enforcement is by construction. See SECURITY.md for the full trust model and the intentional bypass paths (GATE_BYPASS, MCP override).

Uninstall

rm .git/hooks/pre-commit                 # remove the enforcement hook
rm .gate-config.json .gate-store.json    # optional: remove gate config + runtime state

If init backed up a previous hook it saved it as .git/hooks/pre-commit.backup — restore that if you had one. Also delete the convention-gate entry from .vscode/mcp.json (and any other client config).

Background

Built on the study "Advisory, Not Enforceable: Text Cannot Gate Costly Process Steps in AI Coding Agents (But an Out-of-Band Mechanism Can)."

License

MIT © Dvir Shamay

Available Tools

5 tools
create_gate_sessionA

Create a new gate session for a task/sprint/PR. Returns a session ID used for subsequent gate registrations and commit checks. If required_gates is omitted, defaults to this repo's .gate-config.json (falling back to the built-in 9-role set if that file is absent).

ParametersJSON Schema
NameRequiredDescriptionDefault
descriptionYesDescription of the task or sprint (e.g., "S47: harden briefing routes")
required_gatesNoOverride the default required gates. If omitted, uses .gate-config.json's required_gates, or the built-in 9-agent default if that file is absent.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses default behavior ('If required_gates is omitted, defaults to this repo's .gate-config.json...') and the return value. It does not cover side effects or permissions, but as a create tool this is sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences front-load the primary purpose and return value, then cover the default behavior. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter tool with no output schema, the description covers purpose, return value, and default behavior. It could mention error cases or what 'commit checks' entail, but overall it is complete enough for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by explaining the defaulting logic for required_gates (fallback chain) and clarifying the purpose of the returned session ID, exceeding what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Create a new gate session for a task/sprint/PR' and notes it 'Returns a session ID used for subsequent gate registrations and commit checks,' clearly distinguishing it from siblings like register_gate and gate_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies usage as the entry point ('used for subsequent gate registrations and commit checks') but does not explicitly state when to avoid it or compare to alternatives. Context is clear but exclusions are absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gate_statusA

Check the current status of a gate session — which gates have been registered, which are missing, and whether commit is allowed.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesThe gate session ID to check

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description must reveal behavior, and it does disclose that this is a read-only operation (checking status) and what information it returns. However, it does not mention any side effects (likely none), error conditions, or performance implications. The description is adequate but not rich in behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one sentence, front-loaded with the main purpose ('Check the current status of a gate session') followed by detailed specifics. It is concise with no wasted words, earning a high score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no output schema), the description is fairly complete: it states what it does and what information is returned. It lacks explicit mention of error handling or format of 'which gates' but that is minor given the context. An agent can likely use it correctly with this description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description covers 100% of the parameter (session_id has a description), so the baseline is 3. The tool description adds minimal extra meaning beyond repeating that it's a session ID; it doesn't clarify the format or provide examples. The schema already gives the essential info.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool checks the current status of a gate session, listing the specific information it provides (registered gates, missing gates, commit allowed). This is a specific verb ('Check') and resource ('gate session'), and while it doesn't explicitly differentiate from siblings, the purpose is unambiguous enough to distinguish from create/register/commit tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies its usage context by describing what it checks, making it clear this is for inspecting session state rather than creating or modifying. It does not explicitly state when not to use it or name alternatives, but the context is clear enough for an agent to know when to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

guarded_commitA

AUDIT RECORD ONLY — call this AFTER git commit has already succeeded, never before. It does not perform or authorize the commit: on success it marks the session "committed" in the store, and the git pre-commit hook treats a committed session as used up. Calling this BEFORE running git commit will cause that following commit to be BLOCKED (the hook will see no eligible session). Correct flow: register_gate for every required role -> gate_status to confirm commit_allowed:true -> run the real git commit yourself -> optionally call guarded_commit afterward to log it.

ParametersJSON Schema
NameRequiredDescriptionDefault
overrideNoForce-record commit even with missing gates (records override in audit log, self-attested — not independently verified). Does NOT bypass the git hook itself — see GATE_BYPASS for that.
session_idYesThe gate session ID to check
commit_messageYesThe commit message of the commit that was ALREADY made (this call records it, it does not create it)
override_reasonNoReason for override (required if override=true)

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and succeeds fully. It states that the tool does not perform or authorize the commit, that it marks the session as used up in the store, and that the pre-commit hook will block a commit if this is called too early. It also clarifies that the override option is self-attested and does not bypass the git hook, pointing to a separate GATE_BYPASS mechanism. This is excellent behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than usual but every sentence contributes crucial ordering information and behavioral warnings. It is front-loaded with the most important constraint (AUDIT RECORD ONLY, call AFTER commit) and then elaborates logically into the correct flow. No filler words or repetition seen.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a complex tool with subtle misuse consequences, and despite having no output schema and no annotations, the description fully covers purpose, ordering, failure conditions, parameter meanings, and the expected caller behavior. It leaves nothing an agent needs to know to invoke it safely and correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema descriptions already explain each parameter fully (100% coverage), so the baseline is 3. The description adds meaningful context beyond the schema: it emphasizes that commit_message is for the already-made commit, explains the audit/logging semantics of override, and clarifies that override does not bypass the git hook. This adds genuine value to two parameters, so 4 is warranted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states that this tool is for AUDIT RECORD ONLY and describes its exact role: marking a session committed in the store after git commit has succeeded. It distinguishes itself clearly from sibling tools (e.g., it is not gate_status or commit; it logs a completed commit). This leaves no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit precondition ('call this AFTER `git commit` has already succeeded, never before') and explains the exact failure mode if misused (calling before leads to the following commit being BLOCKED). It also provides the correct sequence with sibling tools (register_gate -> gate_status -> actual commit -> guarded_commit). This is precise and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_sessionsA

List all gate sessions (most recent first). Useful for auditing which tasks went through full review.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum sessions to return (default: 10)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It states that results are sorted most-recent-first, implying a read-only listing operation. It doesn't go into auth or side effects, but the list verb makes the non-mutating nature clear enough for this simple tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no wasted words. The core action and ordering are front-loaded, followed by a practical use case. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple: one optional parameterais, no output schema, and clear sibling context. The description is sufficient for an agent to know what it does and when to use it. A minor gap is that the return shape isn't described, but this is not critical for a list operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%: the only parameter, limit, is already described in the input schema with its default value. The description doesn't add further meaning to the parameter, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete verb and object, 'List all gate sessions', and adds the ordering behavior 'most recent first'. This clearly distinguishes the tool from siblings like create_gate_session and gate_status without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Useful for auditing which tasks went through full review' provides a clear use case and context for when to call this tool. It does not explicitly name alternatives or exclusions, but the list-only intent is evident relative to the sibling creation and status tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

register_gateA

Register that a review gate has been completed. Called by each agent after finishing its review. The gate result (pass/fail/warn) and any findings are recorded.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoIdentity of the reviewing agent
resultYesGate result: pass (no blockers), fail (blocking issues found), warn (non-blocking issues)
findingsNoList of findings from the review
metadataNoAdditional metadata (model used, duration, etc.)
gate_nameYesName of the gate/agent (e.g., "code-reviewer", "security-auditor")
session_idYesThe gate session ID (from create_gate_session)

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden. It discloses that the gate result and findings are recorded, making the write nature explicit. However, it does not mention idempotency, whether duplicate registrations are allowed, or whether the registration finalizes the gate session or affects guarded_commit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two focused sentences, with the core action first and the caller context second. Every sentence contributes meaningful guidance, and there is no filler or redundant restatement of schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 6 parameters, nested objects, and no output schema, the description provides the essential context: who calls it, when, and what data it records. It could further explain side effects or the relationship to gate_status/guarded_commit, but the existing guidance is sufficient for correct selection and basic invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description adds no syntax, format, or relationship details beyond the schema; it only mentions gate result and findings, which are already described.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Register that a review gate has been completed.' Adds lifecycle context with 'Called by each agent after finishing its review,' which distinguishes it from sibling tools like create_gate_session and gate_status without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly states when the tool should be called: after each agent finishes its review. It does not explicitly list exclusions or alternative tools, but the lifecycle cue ('after finishing its review') provides clear context that prevents confusion with session creation or status checks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedcreate_gate_session
    • First observedgate_status
    • First observedguarded_commit
    • First observedlist_sessions
    • First observedregister_gate

TDQS

A4.1/5.0

Scored across 5 tools

Disambiguation5/5

Each tool targets a distinct phase of the gate workflow: session creation, gate registration, status inspection, commit logging, and listing. There is no meaningful overlap between tool purposes, and the guarded_commit description clearly differentiates it from actually running a commit.

Naming Consistency3/5

create_gate_session, register_gate, and list_sessions follow a clean verb_noun pattern, but gate_status is a noun phrase and guarded_commit reads as an adjective_noun rather than an action. The naming is still readable and snake_case is consistent throughout, but the deviations are noticeable enough to lower the score.

Tool Count5/5

Five tools is well-scoped for a focused gate-review workflow: create, register, check, log, and list. Each tool serves a necessary function in the lifecycle, and none feel redundant or ornamental.

Completeness4/5

The core workflow is covered end-to-end: session creation, gate registration, status verification, commit logging, and session listing. Minor gaps exist, such as no way to retrieve the detailed findings recorded during register_gate or to cancel/abort a session, but these do not block the primary gate-enforcement flow.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Provides tools for agents to manage a local review graph, tracking acceptance behaviors, evidence, review passes, and human waivers to decouple review convergence from shipping readiness.
    3
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    A local-first, auditable code review MCP server that freezes Git changes, creates immutable ReviewBundles, provides role-isolated contexts for correctness, security, architecture, and test reviewers, validates structured findings, and generates deterministic JSON/Markdown reports.
    3 npm
    1
    Apache 2.0