Skip to main content
Glama

GATEKEEPER MCP

A local Model Context Protocol server that sits between a coding agent and any code-modifying action. Before the agent edits a file, it calls pre_action_check; if the change is not justified by evidence, or the task is already solved, the server tells the agent to stop and report back instead of writing code.

What is GATEKEEPER MCP?

GATEKEEPER MCP is a single-binary MCP server that gates file modifications behind an evidence check. It exposes one tool, pre_action_check, which returns one of three decisions — allow, deny, or request_info — together with a reason and concrete next steps. The agent is expected to call it before every edit and to treat a deny as final. It runs entirely on your machine over STDIO, performs no network calls, and sends no telemetry.

Related MCP server: Sensory-Grounding MCP

Why?

LLM coding agents suffer from Action Bias: the tendency to change code when no change is warranted. A plausible-looking diff gets produced whether or not a bug existed, whether or not the code was already correct, and whether or not the agent can reproduce the problem it claims to be fixing. The result is churn that looks like progress and quietly breaks working software.

Estimates of the rate at which coding agents propose unnecessary edits commonly land in the 35–65% range. Treat that as an order-of-magnitude figure rather than a precise measurement: it varies wildly by task, model, and how "unnecessary" is defined, and the primary sources are informal. It is included here to establish that the problem is large, not to be cited as a hard number.

The intervention is deliberately blunt. Do not make an edit until you can show the problem it fixes.

Install

Requires Node.js >= 20 and npm.

git clone https://github.com/<YOUR NAME>/gatekeeper-mcp.git
cd gatekeeper-mcp
bash scripts/install.sh

The installer checks your Node version, runs npm ci, builds to dist/, and prints the configuration for your agent. Manual equivalent:

npm install
npm run build

Quick start with any agent

The server is not published to npm yet (that is Phase 4), so point your agent at the built entrypoint. Use <PATH_TO_REPO> for the absolute path to this repo: MCP servers start with an unspecified working directory, so a relative path will not resolve.

OpenCode — opencode.json:

opencode mcp add gatekeeper -- node <PATH_TO_REPO>/dist/index.js
{
  "mcp": {
    "servers": {
      "gatekeeper": {
        "type": "local",
        "command": ["node", "<PATH_TO_REPO>/dist/index.js"]
      }
    }
  }
}

Codex — ~/.codex/config.toml:

[mcp_servers.gatekeeper]
command = "node"
args = ["<PATH_TO_REPO>/dist/index.js"]

Claude Code:

claude mcp add gatekeeper -- node <PATH_TO_REPO>/dist/index.js

Hermes Agent — ~/.hermes/config.yaml. This command is interactive: it connects, lists the tools it finds, and asks which to enable. Answer Y.

hermes mcp add gatekeeper --command node --args <PATH_TO_REPO>/dist/index.js

Once published to npm, every command above works with npx gatekeeper-mcp instead.

Compatibility matrix

Agent

Status

Config path

Guide

OpenCode v1

UNVERIFIED

opencode.json

docs/agents/opencode.md

OpenCode v2

VERIFIED

opencode.json

docs/agents/opencode.md

Codex

UNVERIFIED

~/.codex/config.toml (unconfirmed)

docs/agents/codex.md

Claude Code

VERIFIED

~/.claude.json / .mcp.json

docs/agents/claude-code.md

Hermes Agent

PARTIALLY VERIFIED

~/.hermes/config.yaml

docs/agents/hermes.md

Configs that are not yet fully verified are marked accordingly in each guide.

The snippets above are a starting point. Prefer each agent's mcp add command over hand-editing a file: the CLI writes whatever the installed version actually expects, which is the one thing that changes between major versions.

Tools

pre_action_check

Gatekeeper check. Call this BEFORE modifying any code. It verifies whether the change is justified by evidence and whether the task is already solved. If not justified, the agent MUST NOT edit files.

Input

Field

Type

Required

Description

taskDescription

string

yes

What the agent is trying to accomplish.

proposedChange

string

yes

The diff, or a natural-language description of the change.

affectedFiles

string[]

no

Files the agent intends to modify.

evidenceOfProblem

string

no

Reproduction steps, a failing test, or an error log.

reproduction

object

no

An executable command that reproduces the problem.

runTests

boolean

no

Run the project test suite as extra evidence. Default false.

reproduction is the field that matters. Free text is never accepted as proof, so evidenceOfProblem on its own can only ever produce request_info.

reproduction field

Type

Default

Description

command

string

—

Bare executable, e.g. "npm", "pytest", "node".

args

string[]

[]

Arguments. Never put arguments in command.

cwd

string

repo root

Must resolve inside the repository root.

timeoutMs

number

30000

Kill after this long. Range 1000-300000.

expectFailure

boolean

true

true: non-zero exit means reproduced. false: invert it.

The command is spawned directly, never through a shell. command: "npm test" is rejected, because accepting a string like that would mean handing it to a shell. Split it into command: "npm" and args: ["test"].

Output

Field

Type

Description

decision

"allow" | "deny" | "request_info"

The verdict.

reason

string

Why the verdict was reached.

nextSteps

string[]

Ordered instructions the agent should follow.

evidence

object

What the reproduction stage found.

state

object

What the state inspection stage found.

evidence field

Type

Description

state

"reproduced" | "not_reproduced" | "unverifiable" | "timeout"

Outcome of the reproduction.

detail

string

Human-readable explanation.

exitCode

number

Present only when a command actually ran.

durationMs

number

How long the command took.

state field

Type

Description

workingTree

"clean" | "dirty" | "unknown"

Git working tree status.

tests

"pass" | "fail" | "skipped" | "unknown"

Result of the test suite.

detail

string

Human-readable explanation.

Decisions

  • allow — justified, proceed with the edit. Reachable only when the reproduction was executed and failed as declared, the working tree is clean, and no green test suite contradicts the failure.

  • deny — not justified, do not touch any file. The agent should reply "No change needed — the current code already satisfies the task."

  • request_info — evidence is missing, ask the user before editing.

Reproduction command

A full call that can reach allow, for a bug where a test times out:

{
  "taskDescription": "Uploads hang when the network is slow",
  "proposedChange": "Add a bounded retry around the upload call in src/upload.ts",
  "affectedFiles": ["src/upload.ts"],
  "reproduction": {
    "command": "npm",
    "args": ["test", "--", "upload.test.ts"],
    "timeoutMs": 60000
  },
  "runTests": false
}

The server runs npm test -- upload.test.ts, reads the exit code, and returns:

{
  "decision": "allow",
  "reason": "Failure reproduced under a clean working tree. The change is justified.",
  "nextSteps": [
    "Make the minimal change that makes the reproduction pass.",
    "Do not refactor unrelated code."
  ],
  "evidence": {
    "state": "reproduced",
    "detail": "Reproduction succeeded: \"npm\" exited 1 in 843ms. Output: ...",
    "exitCode": 1,
    "durationMs": 843
  },
  "state": {
    "workingTree": "clean",
    "tests": "skipped",
    "detail": "The git working tree is clean."
  }
}

Set expectFailure: false for a check that is expected to pass today, such as an assertion the change would break. A zero exit then means reproduced, in the sense that the expected outcome was observed.

Change the same call to a command that exits 0 and the verdict flips to deny: the reproduction succeeded, so the code may already be correct. That asymmetry is the point.

Development

npm run dev        # watch mode via tsx
npm test           # vitest, single run
npm run test:watch # vitest, watch mode
npm run typecheck  # tsc --noEmit
npm run lint       # biome check
npm run format     # biome format --write
npm run build      # compile to dist/

Project layout

src/
  index.ts                  entrypoint: stdio transport
  server.ts                 McpServer factory and tool registration
  logger.ts                 pino, bound to stderr
  tools/pre_action_check.ts the one tool
  engine/                   analyzer + reproduction / state / decision stages
  schemas/                  zod input and output schemas
  types/                    domain types

stdout is reserved for the JSON-RPC frame stream. All logging goes to stderr. SKILL.md is the agent-facing contract and is the file to read if you are building an agent that must respect the gate.

Architecture: docs/architecture.md · Roadmap: docs/phases.md · Decisions: docs/decisions/

License

MIT — see LICENSE.

Available Tools

1 tool
pre_action_checkPre-action gatekeeper checkA
Read-onlyIdempotent

Gatekeeper check. Call this BEFORE modifying any code. It verifies whether the change is justified by evidence and whether the task is already solved. If not justified, the agent MUST NOT edit files. Supply an executable reproduction under reproduction (a bare command plus args, never a shell string): free-text evidenceOfProblem alone is never enough to obtain an "allow".

ParametersJSON Schema
NameRequiredDescriptionDefault
runTestsNoRun the project test suite as extra evidence. Default false.
reproductionNoAn executable command that reproduces the problem.
affectedFilesNoFiles the agent intends to modify.
proposedChangeYesThe diff, or a natural-language description, of the proposed change.
taskDescriptionYesWhat the agent is trying to accomplish.
evidenceOfProblemNoFree-text proof that a problem exists. Not executable on its own; supply `reproduction` to obtain an allow decision.

Output Schema

ParametersJSON Schema
NameRequiredDescription
stateYesWhat the state inspection stage found.
reasonYesOne or two sentences explaining the verdict.
decisionYesThe verdict for the proposed change.
evidenceYesWhat the reproduction stage found.
nextStepsYesOrdered, actionable instructions.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds valuable behavioral context: the tool acts as a gate that can block edits, verifies task completion, and mandates a specific reproduction format. It does not explicitly state the return decision structure, but the output schema likely covers that, and the 'allow' requirement implies it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise—four sentences that front-load the main purpose and then detail the critical constraint. Every sentence earns its place: the gatekeeper role, the requirement to call before editing, the reproduction format, and the rule about evidence. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (6 params, nested object, output schema), the description fully covers the essential usage: when to call, what to supply, and the decision rule. The output schema covers return values, so no additional return explanation is needed. The description is complete for an agent to correctly invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, baseline is 3. The description enriches parameter understanding beyond the schema by specifying that `command` must be a bare executable, `args` must not be embedded in `command`, and that `evidenceOfProblem` alone is insufficient without `reproduction`. This adds practical usage meaning to the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('verify'), the resource (whether a change is justified and task already solved), and context (before modifying code). It clearly distinguishes the tool's role as a gatekeeper without ambiguity, and no siblings exist to confuse.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs 'Call this BEFORE modifying any code' and sets a hard rule: a reproduction is required to obtain an 'allow', and free-text evidence alone is insufficient. This gives clear when-to-use and what-to-provide guidance, with no alternative tools to compare.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.0.1
    • First observedpre_action_check

TDQS

A4.5/5.0

Scored across 1 tool

Disambiguation5/5

The server exposes exactly one tool, `pre_action_check`, so there is no possibility of confusing it with another tool. The purpose of the gate check is single and unambiguous.

Naming Consistency4/5

The single tool name `pre_action_check` is clear and snake_case, but it does not follow a verb-noun pattern and there is no set of tools to establish a broader naming convention.

Tool Count3/5

With only one tool, the server feels thin for the 'gatekeeper' domain, though the narrow purpose partially justifies the count. Auxiliary tools like policy lookup or reproduction validation are absent.

Completeness4/5

The core gatekeeping interaction is covered, but the surface lacks related operations such as querying gate rules or validating evidence formats, so coverage is functional but has minor gaps.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers