Skip to main content
Glama

# SentinelMCP

Sprint 6 · Cybersecurity · Client: Ironclad Cyber Defense

A production-style MCP server exposing security-triage tools to any MCP client, paired with an AutoGen developer/reviewer agent loop that drafts human-approved remediation pull requests — and an AutoGen-free deterministic engine so the whole path is testable with no API key.

The system scans a codebase, persists findings against severity-based SLAs, and attempts to fix them. It can propose, review, and draft. It cannot merge. That is enforced five ways, one of which is a test that AST-walks the codebase to prove no new code can add a merge path.

Claude Desktop ──▶  MCP server ─┐
                                ├──▶  gRPC ScannerService ──▶  Semgrep / fixtures
human analyst ───▶  FastAPI ────┘                                  │
                                                                   ▼
                                                          PostgreSQL

What this cannot do — read this first

The remediation engine is six regex rules over a single line of source. That is the whole of it, and these are the measured numbers on real code.

measured

Real vulnerable files where a rule could act at all

3 of 30 (10%)

Patches that stopped a real exploit, of those proposed

5 of 5

Rewrites that need a human step before the code will run

2 of 6

Fixture CWE classes with no rewrite, which escalate to a human

7 of 13

Real merged fixes our reviewer would reject

9 of 30 — 8 of 30 with the enclosing block the loop is now shown

All of this is measured, not asserted, by three tiers of evidence against real exploits, an independent static analyser, and 19 published CVEs. Full write-up, including the eight defects the harness found in the engine and in itself: docs/evidence.md.

So the honest pitch is triage, not autonomous remediation. The queue ordering, the SLA clock, the policy gates and the audit trail work on any finding a scanner produces. The drafting is narrow: mechanical, high-confidence, single-line fixes, with everything else escalated to a human. That is a real product — a managed security provider's volume problem is triage, and an agent that drafts the 10% it can do correctly while escalating the rest beats an agent that guesses at the 90% it cannot.

The number that is not a limitation: the merge block is real. A test AST-walks backend/ to prove no future commit can add a code path that merges. That is the part of this project I would defend without qualification.


Related MCP server: saglitzsecure-mcp

Why this exists

A managed security provider's analysts get hundreds of SAST/DAST/dependency findings a week and spend most of their time triaging rather than remediating. The two expensive things are deciding what matters, and turning a finding into a change — most of the queue can do the first and not the second.

So: order the queue by time remaining rather than severity, and let a machine attempt the change on the narrow class where a mechanical fix is unambiguous. The second one is only safe with a hard human gate, which is the part this project is actually about.

Resume line

Built a production-style MCP server exposing security-triage tools to any MCP client, paired with an AutoGen developer/reviewer agent loop that drafts human-approved remediation pull requests — with a three-tier evidence harness that found eight defects in the engine, including one asserted as correct by a green test.


What it does

4 MCP tools

list_findings, get_cve_details, propose_fix, create_ticket — over stdio (what Claude Desktop launches) or streamable-http

36 HTTP endpoints

scans, findings, remediation, pull requests, SLA dashboard, tickets, CVE lookup, read-only audit

gRPC scanning service

protobuf-typed Scan and Health; the fingerprint is computed on the wire side

2 agent engines

AutoGen (opt-in, LLM) and a deterministic 6-rewrite one (default) behind one protocol

7 policy gates

severity, confidence, status, path allow-list, snippet, attempt budget — re-evaluated fresh, never trusted from a stored flag

16 audit actions

every finding, every agent turn, every human decision

3 evidence tiers

patch correctness measured by execution, by an independent static analyser, and against 19 real CVEs

414 tests

ruff + mypy + pytest, all green; 159 of them against real Postgres, real sockets, and real upstream commits

Try it

git clone <this repo> && cd SentinelMCP
cp .env.example .env

# Option A: everything in Docker (a real gRPC boundary)
docker compose up -d --build
docker compose exec api python -m backend.scripts.seed
open http://localhost:8000/docs

# Option B: no Docker beyond Postgres
docker compose up -d postgres
pip install -e ".[dev]"
alembic upgrade head
make seed
make api          # http://localhost:8000/docs

Then, to see the whole thing work end to end across two real processes:

python scripts/smoke_e2e.py
  [PASS] the API reports itself healthy
  [PASS] and that the scanning service behind it is reachable
  [PASS] auto-merge is blocked and human approval is required
  [PASS] fixture targets are discoverable over gRPC
  [PASS] the scan succeeded                          new=8 total=8
  [PASS] findings were persisted
  [PASS] the outcome is a legitimate agent result     outcome=approved rounds=1
  [PASS] the draft is blocked from auto-merge         status=draft auto_merge_blocked=True
  [PASS] an ANALYST cannot approve                   HTTP 403
  [PASS] an APPROVER can approve
  [PASS] and approving it STILL does not merge it    status=open blocked=True
  [PASS] the auto-merge block was recorded by the policy engine, not the approver
SMOKE PASSED -- 21 checks

Note what this does and does not show. It scans data/fixtures/, which is hand-written to match the rule table, so "the outcome is a legitimate agent result" means the plumbing works end to end — not that the fix is correct. On real code the engine engages on 10% of vulnerable files. make evidence is the version that measures the fix.

Connect it to Claude Desktop

{
  "mcpServers": {
    "sentinel": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/SentinelMCP",
               "python", "-m", "backend.scripts.run_mcp_server"],
      "env": {
        "SENTINEL_DATABASE_URL": "postgresql+asyncpg://sentinel:sentinel@localhost:5432/sentinel",
        "SENTINEL_SCANNER_TRANSPORT": "in_process"
      }
    }
  }
}

in_process because a stdio server is a subprocess the client spawned — there is no second process to dial. See ADR 003.


The guarantee

The system cannot merge its own code. Not "it doesn't by default" — enforced, redundantly, because a guarantee that depends on nobody ever finding the exception is a convention, not a guarantee.

  1. No merge code exists. publisher.merge_pull_request() exists and its body raises.

  2. The column default is True. auto_merge_blocked is Boolean, default=True in the schema and hardcoded at every insert. The database would have to be altered to store an unblocked pull request.

  3. SENTINEL_ALLOW_AUTO_MERGE is read only to raise. Set it and the process refuses to start. It is a tripwire, not a switch.

  4. A test AST-walks the source. No assignment anywhere sets auto_merge_blocked to anything but True; exactly one function writes PullRequestStatus.MERGED; no route path ends in /merge. Check 4 catches a new merge path, which the other three would not.

  5. The smoke test proves it on two real processes: blocked → analyst 403 → approver OK → still not merged.

Full reasoning, including the alternatives that were refused: ADR 001.

This is the load-bearing claim, and it is the one thing here I would defend without qualification. The remediation engine is narrow and I have said so; this is not narrow and it is not a matter of degree.


Is the patch actually a fix?

The system can prove it cannot merge its own code. Nothing in it could say whether a drafted patch works — until this.

make evidence        # three tiers, six real advisories
make evidence-all    # all nineteen

tier

question

whose answer

needs a network

1 execution

does the patch stop a real exploit?

ours, by running it

no

2 static

does an independent detector stop reporting it?

Semgrep's

first run only

3 corpus

what happens on real CVEs and real upstream fixes?

upstream's

yes

Tier 1 writes a real vulnerable module, runs a real attack against it, applies the agent's patch, and runs the attack again. Tier 2 asks Semgrep whether the finding is still reported. Tier 3 fetches the real fix commit for 19 GitHub Security Advisories and compares.

There is deliberately no pass rate: every denominator available is flattering by construction. The report leads with what each tier could not check, and the counts of patches that are incomplete, that regress the function, and that no tier covers.

What it found. Eight findings — seven code defects and one measurement defect. Four of the code defects were invisible to a 300-test suite, and one of those was being actively asserted as correct by a green test:

  • The SQL rewrite emitted a statement that psycopg2 rejects and sqlite3 executes while silently returning the wrong rows — a LIKE wildcard colliding with the %s placeholder. The existing test asserted the text contained %s; the broken form does too.

  • The reviewer requested changes on a correct one-line signature change, because it reconstructed the patched source from the diff's added lines and got a block header with no body. Fixing that was not enough: the same check was also rejecting any diff that was not a like-for-like line replacement, which was 15 of 18 rejections on real upstream fixes.

  • Two rules emitted Python into Go files, and the test for one of them asserted that was correct, on a .go path.

  • eval( matched the declaration of a method named eval; task_keyprefix was treated as a secret.

  • Two bugs in the harness itself, found by its own mutation control.

  • And a measurement defect: the reviewer-rejection rate was two-thirds an artifact of feeding the reviewer a one-line snippet and judging 67-line fixes against it.

All fixed, all pinned by regression tests. The unflattering numbers are still reported, and one of them turned out to be the harness's fault rather than the engine's:

  • The rule table engages on 3 of 30 real vulnerable files (10%) — an upper bound, since the candidate line is found by matching the table's own patterns. A six-regex table is the wrong shape for most real vulnerabilities, and no amount of work changes that without a different product.

  • 9 of 30 real, merged, human-reviewed fixes would be rejected by our own reviewer — 8 of 30 once the loop is shown the finding's enclosing block instead of the single line a scanner reports. A ±10 line window takes it to 3 of 30, and is not what is shipped: it satisfies the check by sweeping in neighbouring lines by accident rather than by being the right context. The residual is fixes that also edit an import far from the finding, which the reviewer is arguably right to flag.

Full write-up, including why a bad --config URL exiting 0 makes a static-analysis gate silently meaningless, and why one of these numbers was a measurement artifact before it was a product finding: docs/evidence.md.


Architecture

Three processes, one database.

                    ┌──────────────────────────────────┐
 Claude Desktop ───▶│  MCP server  (stdio / http)      │
 or any MCP client  │  4 tools · errors returned,       │
                    │  never raised                    │
                    └────────────────┬─────────────────┘
                                     │
 human analyst ────▶┌────────────────┴─────────────────┐
                    │  FastAPI                          │
                    │  32 endpoints · JWT · 3 roles      │
                    └────────────────┬─────────────────┘
                                     │  ScannerBackend (ABC)
                                     ▼
                    ┌──────────────────────────────────┐
                    │  gRPC ScannerService  :50051      │
                    │  Scan() → typed findings          │
                    │  Health() → real capability       │
                    └────────────────┬─────────────────┘
                                     ▼
                    FixtureScanner · SemgrepScanner → PostgreSQL

Why the boundary is real. Scanning is CPU-bound and shells out to Semgrep; serving is I/O-bound and must stay responsive during an incident. In one process, one slow scan makes /health time out, and a health check that lies is worse than no health check.

SENTINEL_SCANNER_TRANSPORT picks the implementation — grpc, in_process, or auto (default: probe with a 2s deadline, fall back with a WARNING). Both implementations satisfy the same ABC, so ScanningService depends on the interface and the test suite needs no socket for 280+ of its 414 tests.

Full walkthrough: docs/architecture.md.

The data model, briefly

  • A finding is identified by a fingerprint, a SHA-256 over (scanner, rule_id, normalised path, target). It deliberately excludes the snippet and line numbers, so a scanner whose output shifts by two lines updates the finding instead of arriving as a new one — and the SLA clock does not restart. Line number is not identity.

  • Enums are VARCHAR + CHECK, not PostgreSQL ENUMs. Adding a value to a native enum takes an ACCESS EXCLUSIVE lock; here it is a one-line change.

  • audit_logs has no updated_at and nothing updates it. An event log with a mutable timestamp is a contradiction.

  • The engine is keyed by event loop. This process runs two long-lived loops, and SQLAlchemy's async pool holds connections bound to whichever loop first checked them out.

The agent loop

for round in range(3):                          # and at most 8 LLM calls
    patch   = developer.propose(snippet, feedback)   # may ESCALATE
    verdict = reviewer.critique(patch)               # APPROVE / REQUEST_CHANGES / REJECT
    if approved: return patch
return escalate_to_human()

Two engines implement one protocol, so the service cannot tell which ran and the audit trail is identical:

  • Deterministic (default). Six regex rules, applied to the one line the scanner reported (eval → ast.literal_eval, SQL f-string → bound parameter, shell=True → argv, disabled TLS verification, hardcoded secret → env, proxy-auth-on-redirect), and a reviewer that checks the patch parses, that the vulnerable construct is gone, and that it is not oversized. It escalates on a class it has no rewrite for, on a language its replacement does not compile in, and on a finding with no snippet — rather than guessing. Free, no network, deterministic — so it is the test oracle for the loop's control flow, and it makes the whole remediation path runnable in CI. This is the 10% in the table at the top, and the reason is structural: a real vulnerability is usually not one line.

  • AutoGen (opt-in). Same protocol, LLM behind it, with per-call token and cost attribution. The two roles run as separate single-agent chats, never one team: a reviewer that has read the developer's justification anchors on it, and not anchoring is the entire value of the review step.

Each rewrite documents its own limits, and a test asserts they are documented — a fix that half-solves a problem is the most dangerous thing an agent can output. And because documentation is not evidence, the three evidence tiers measure those limits by execution. That harness removed a seventh rewrite (math.random → secrets): its pattern was Go's and its replacement was Python's, and no regex can bridge the two.


Design decisions

ADR

decision

the short version

001

Humans merge, always

five enforcement mechanisms; the AST test is the one that catches a new merge path

002

AutoGen, not LangGraph or CrewAI

the reviewer must not be anchored; AutoGen makes separate agents and model clients the default. LangGraph is a state graph and this is a counter

003

Typed gRPC boundary

scanning is subprocess-heavy and must be separable; and an error must never be representable as an empty success

004

Support both mcp majors

semgrep pins mcp<2; pinning either way makes the scanner and the MCP surface mutually exclusive. Verified: 10 identical checks pass on both v1 (local) and v2 (container)

Both SDK majors, actually verified

scripts/verify_mcp_sdk.py runs the same ten assertions against whichever mcp major is installed. The two environments genuinely differ, and that falls out of the dependency graph rather than being arranged:

environment

mcp resolves

why

local venv

v1

something in the tree pins mcp<2

runtime image

v2

the image deliberately does not install semgrep

make verify-mcp                                     # whichever you have
docker run --rm sentinelmcp:ci python scripts/verify_mcp_sdk.py   # the other one

CI runs both (mcp-sdk and docker jobs), so a change that works on only one major fails CI rather than a reviewer's environment.


Verification

make check          # format-check, lint, typecheck, test  (what CI runs)
make test           # pytest; integration tests skip themselves
make test-integration
python scripts/smoke_e2e.py
make evidence       # measured patch correctness -- see docs/evidence.md

gate

tool

result

format

ruff format --check

126 files

lint

ruff check

clean

types

mypy (disallow_untyped_defs)

114 files, clean

tests

pytest

414 passed

integration

pytest -m integration

159 passed against real Postgres, real sockets, and real upstream commits

end to end

scripts/smoke_e2e.py

21 checks across 2 processes

evidence

scripts/evidence_report.py --all

3 tiers, 19 real advisories, 0 failures

CI (.github/workflows/ci.yml) runs the same six gates plus a container build, a real /health/policy check on the running image, and tier 1 of the evidence harness. Only tier 1: tiers 2 and 3 need semgrep.dev and seven upstream clones, so a network blip would train contributors to re-run red builds. Tier 1 needs no network, no database and no key, runs in about two seconds, and is what notices when a rewrite stops working.

The suite's actual shape

group

count

what only it can prove

tests/unit/policy/test_safety_rails.py

21

the AST assertions — no new code can set the block or write MERGED

tests/integration/test_db_session_dependency.py

8

the real session dependency, which every other test overrides

tests/integration/test_mcp_tools.py

37

the four tools through a real MCPServer

tests/integration/test_grpc_scanning.py

22

typed schema and status codes over a real socket

tests/integration/test_pipeline_persistence.py

20

append-only audit log; a duplicate fingerprint rejected by the database

tests/integration/test_scanner_transport.py

12

which backend implementation each mode actually produces

tests/unit/agents/test_deterministic_agents.py

26

each rewrite's diff, and that it documents its own limits

tests/unit/policy/test_source_roots_setting.py

16

that the path allow-list is operator-configurable, traversal-safe, and fails closed

tests/unit/telemetry/test_telemetry_contract.py

16

the frozen cross-fleet event shape Sprint 11 depends on


Two bugs this project found in itself — and eight the evidence harness found

Included because a portfolio that only shows the happy path is a brochure, and because these are exactly the kind of defect a green suite hides.

These two were found by scripts/smoke_e2e.py, starting two processes and looking.

get_session() never committed. Its docstring promised "commits on clean exit, rolls back on any exception"; the body rolled back and did not commit. Closing the session therefore discarded everything a route had written: POST /scans answered {"new_findings": 8, ...} and left the findings table empty. Every write endpoint in the HTTP API was a no-op reporting success.

get_db_session() tore its generator down with asend(None). Resuming an async generator that has already yielded once makes it run to completion and raise StopAsyncIteration, which PEP 479 turns into RuntimeError on the way out of a coroutine. It sat in a finally, so it replaced the response of every authenticated request with a 500.

Both were invisible to a fully green suite, because every route test overrides get_db_session with a session it manages itself — so neither function was executed by a single test.

The naive fix for the second — swapping asend(None) for aclose() — removes the 500 and introduces a quieter bug, because aclose throws GeneratorExit at the yield and the commit never runs. The real fix was to stop hand-driving a generator: backend/database/session.py now owns a single transaction() context manager, and get_session, session_scope, and get_db_session are all thin wrappers over it.

The rule that came out of it, now written into the test suite:

If a test overrides a dependency, some other test has to run the real one. Otherwise the override is not a test double, it is a hole.

The eight the evidence harness found later

make evidence was written because of the two above, on the reasoning that if starting two processes found two session bugs, asserting on diff strings would find more. It found eight: a SQL rewrite that silently returned the wrong rows, two rules that emitted Python into Go files — one of which a green test was asserting was correct, on a .go path — a reviewer that rejected correct patches for two separate reasons, eval( matching the declaration of a method named eval, task_keyprefix treated as a secret, two bugs in the harness itself, and one number that turned out to be an artifact of the harness rather than a fact about the engine.

The rule that came out of that, also written into the suite:

A verification harness that cannot fail is worse than none. The exploit must land on the unpatched code first, the mutation control must prove it can tell a fix from a no-op, and a rule set that reports nothing must be reported as vacuous rather than as a clean scan.

Full account, with before-and-after numbers: docs/evidence.md.


Layout

backend/
  agents/        the developer/reviewer pair: autogen_engine, deterministic, and the loop
  api/           FastAPI routers, dependencies, correlation-id middleware
  auth/          JWT, Argon2 hashing, roles, the session dependency
  config/        pydantic-settings; every env var with its default
  core/          clock, ids and fingerprints, logging, asyncio bridging
  database/      models, repositories, enums, the loop-scoped engine
  evidence/      3 tiers of measured patch-correctness evidence; no runtime path
  grpc_service/  the ScannerService server, the two client backends, conversion
  mcp_server/    the 4 MCP tools and their registration
  policy/        the 7 remediation gates and the SLA state machine
  scanners/      base interface, registry, fixture and Semgrep adapters, snippet enrichment
  services/      scanning, remediation, approvals, publisher, sla, cve
  telemetry/     the frozen event contract
proto/           sentinel/v1/scanner.proto
alembic/         one initial migration, verified by `alembic check` in tests
data/evidence/   the checked-in advisory manifest; the git cache is gitignored
docs/            architecture.md + evidence.md + 4 ADRs
scripts/         smoke_e2e.py, evidence_report.py
tests/           unit + integration, 414 tests

Configuration

Every variable is documented in .env.example with its default and why it exists. The ones that change behaviour:

variable

default

effect

SENTINEL_ALLOW_AUTO_MERGE

false

Set it and startup raises. A tripwire, not a switch.

SENTINEL_SCANNER_TRANSPORT

auto

grpc | in_process | auto. grpc makes a dead target a startup failure.

SENTINEL_AUTOGEN_ENABLED

false

true + SENTINEL_LLM_API_KEY switches to the LLM engine.

SENTINEL_REMEDIATION_SOURCE_ROOTS

["app/","src/","services/","lib/","config/"]

the only paths the agent may draft patches under. Empty list fails closed.

SENTINEL_REMEDIATION_SNIPPET_MAX_LINES

120

ceiling on how much source the loop is shown, around the finding. 0 disables the ceiling.

SENTINEL_ENABLED_SCANNERS

["fixture"]

add "semgrep" for real SAST

SENTINEL_APP_ENV

development

production makes startup refuse a default JWT secret

SENTINEL_TELEMETRY_SINK

file

file | http | null

Note list-valued settings need JSON, not comma-separated: fixture,semgrep is an error at import.

Known limits

Stated here rather than left to be discovered. The first is the one that matters most, and it is measured rather than estimated.

  • The remediation engine covers about 10% of real vulnerabilities. Six regex rules over one line of source, against 30 real vulnerable files pulled from 19 published CVEs: a rule could act on 3. A real vulnerability is usually not one line, so this is a shape problem and more rules would not fix it. Seven of the thirteen CWE classes in data/fixtures/ have no rewrite and escalate to a human, which is the intended behaviour. Measured, with the reasoning and the two attempts to improve it: docs/evidence.md.

  • The reviewer is starved of context — was. It was handed the one line the scanner reported, so it rejected any fix touching neighbouring lines. The loop is now shown the finding's enclosing block, bounded by SENTINEL_REMEDIATION_SNIPPET_MAX_LINES, and the strategy used is recorded in the audit row. See backend/scanners/snippet.py; the residual rate and what it actually means are in docs/evidence.md.

  • CWE-295 has no verification. The TLS rewrite is Go-only, there is no Go toolchain here, and no Semgrep rule was found that fires on InsecureSkipVerify: true. Tier 1 and Tier 2 independently record it as unverified rather than passing.

  • CWE-338 has no rewrite. The rule that used to claim it emitted a Python name into a Go file, and math/rand.Int and crypto/rand.Int have incompatible signatures, so it was removed rather than patched.

  • Semgrep is not in the container image and the Dockerfile does not install it, deliberately — it is a heavy dependency and the fixture scanner keeps the demo deterministic. Point it at data/samples/vulnerable_app/ and it works; the scanner adapter is real (fixed argv, subprocess with a timeout, absolute paths resolved to repo-relative so fingerprints are machine-independent).

  • ScannerKind.ZAP and DEPENDENCY are declared but unimplemented. The schema can carry them; no adapter exists.

  • The MCP server has no authentication. No token check, no role check, no transport-level auth. Defensible on stdio (a pipe the client created, so no network peer can reach it), not defensible on streamable-http, which is therefore behind a compose profile that is off by default.

  • 28 of the 36 HTTP endpoints are unauthenticated — the read paths, by design. The write paths and all three approval endpoints are not.

  • The publisher is simulated. It never contacts a git host: no token, no remote, no HTTP. A portfolio project must not be able to push to a real repository by accident.

  • stdout is the MCP protocol channel, and backend/core/logging.py writes logs to stdout. There is a startup guard against a dirty stdout, but a runtime INFO line still lands on the wire. Routing logs to stderr is the real fix and is not done.

  • data/samples/vulnerable_app/ is deliberately vulnerable. Excluded from ruff, from mypy, and from import. Do not "fix" it.

More, including the unresolved tensions: docs/architecture.md §11.


Sprint coverage

Roadmap technologies for Sprint 6: MCP, AutoGen, gRPC, and SAST/DAST (stretch).

  • MCP — the whole mcp_server/ package, both transports, both SDK majors

  • AutoGen — agents/autogen_engine.py, three packages not the meta-package

  • gRPC — proto/, the server, two client backends, 22 tests over a real socket

  • SAST — a real Semgrep adapter. DAST (ZAP) is a schema placeholder only.

Deliberately not built here: Redis. An earlier revision of this project declared a Redis service, a SENTINEL_REDIS_URL, a SENTINEL_SCAN_CACHE_TTL_SECONDS, and a fakeredis dev dependency "for rate limiting and a scan-result cache" — and implemented none of it. Redis is not in this sprint's scope; the roadmap builds it properly in Sprint 4 and reuses it in Sprints 7 and 10. A cache that is declared but absent is worse than no cache, because a reader assumes it is protecting something. It has been removed, and docker-compose.yml says why.


Development

make install      # uv sync --extra dev + pre-commit install
make help         # all targets
make proto        # regenerate gRPC stubs after editing the .proto
make check        # the full quality gate
make evidence     # the patch-correctness harness -- see docs/evidence.md

make targets are Unix-only (Git Bash / WSL / macOS / Linux). On plain Windows PowerShell, run the underlying commands — .venv\Scripts\python -m ruff check ., -m mypy ., -m pytest.

Pre-commit runs the same three tools as CI: ruff-format, ruff, mypy. Not pytest, because it needs a database for the integration half and is too slow for a commit hook.

Further reading

MIT licensed.

Related MCP Connectors

Related MCP Servers