Skip to main content
Glama
DeepSleuth

MCP Security scanner

Official

Deepsleuth — read the fine print

CI

Deepsleuth is a deterministic, no-LLM security scanner for MCP servers. It audits what a server says — and, more importantly, what it does.

Most MCP scanners read only the declared manifest (tools/list names, descriptions, schemas). They never launch the server, never call a tool, never read a response, never read the implementation source, and never reason across calls — so whole classes of attack are structurally invisible to them.

Deepsleuth sees those. It is a single frontend-agnostic detection core with two frontends:

  • Frontend A — inline MCP gateway / proxy (the headline artifact). A transparent proxy that is an MCP server to the agent and an MCP client to one downstream server. It audits tool descriptions at startup, enforces a gate before every tools/call, and scans every response before returning it. Includes a headless proxy-eval mode for offline scoring.

  • Frontend B — batch / sandbox scanner. A pre-flight auditor that launches a server in a Docker sandbox, actively elicits behavior with synthesized calls + planted canaries, and produces findings. Also the offline scoring harness.

Both frontends run the same detectors over the same Context — a detector is written once and works in both.

No LLM. Ever.

Fully deterministic: parsing, static AST + taint/dataflow, normalized regex/token heuristics, unicode/encoding/entropy analysis, structural diffing, and sandboxed dynamic execution with instrumentation. Same input → byte-identical findings. Offline (no network egress except to the Docker daemon). No threat feeds.


Install

Python 3.11+. No required third-party packages — the scanner speaks MCP over stdio itself, so it installs in externally-managed (PEP 668) environments.

pip install security-scanner-deepsleuth-mcp                             # published on PyPI
# or from source:
pip install git+https://github.com/DeepSleuth/deepsleuth-mcp.git   # zero required dependencies
deepsleuth --help
# or straight from the source tree:
python -m deepsleuth --help

The PyPI distribution is named security-scanner-deepsleuth-mcp (what it is, plus the brand); the tool, repo and MCP server are deepsleuth. The package installs the deepsleuth console command — and an alias script under the distribution's own name, so uvx security-scanner-deepsleuth-mcp runs the server for registry clients.

Deepsleuth is itself an MCP server, so agents can scan with it directly:

{"mcpServers": {"deepsleuth": {"command": "python", "args": ["-m", "deepsleuth.mcp_server"]}}}

Tools: list_detectors, check_listing, scan_target.

Install as an agent plugin

The repo is a valid Agent Plugins package (plugin.json + mcp.json, spec 1.0.0): any compatible client can install it directly from the repository and gets the deepsleuth MCP server plus the audit-mcp-server skill. The stdio entry (bin/deepsleuth-mcp) needs only python3.11+ — the scanner has zero third-party requirements:

{"type": "stdio", "command": "./bin/deepsleuth-mcp"}

For the dynamic layer (Frontend B and proxy-eval) you need the Docker CLI + daemon. Without Docker the scanner degrades gracefully: static/manifest detectors still run and the skipped dynamic coverage is reported (never a crash).

Related MCP server: SentinelMCP

Run

# Frontend B — batch/sandbox scanner (also the offline scoring harness)
python -m deepsleuth scan <target> [--no-dynamic] [--json out.json] [--timeout N] [--reference-listing tools.json]

# Frontend A — inline MCP gateway/proxy (the gate); speaks MCP on stdio to the agent
python -m deepsleuth proxy <target> [--policy policy.yaml] [--fail-closed] [--log run.jsonl]

# Frontend A headless — drive a deterministic call plan through the proxy, emit the findings JSON
python -m deepsleuth proxy-eval <target> [--json out.json] [--timeout N] [--policy p]

# list every registered detector
python -m deepsleuth detectors

<target> can be a server directory (with mcp.json and/or source), an mcp.json launch spec, or a raw stdio launch command (e.g. "python3 server.py"). scan exits 0 when clean and non-zero once a finding reaches --fail-severity (default high).

--allow-unsandboxed runs the dynamic layer without Docker — use it only for your own trusted fixtures, never on untrusted servers.

--reference-listing tools.json supplies another server's tool list (a JSON array of {name, description, inputSchema} entries, or an object with a tools key) so the cross-server name comparison runs against it without launching a second server. The same comparison also runs automatically across several entries in one mcp.json and across several server entry modules found in one directory.

Wire the proxy into an agent

Point your MCP client at the proxy instead of the real server; the proxy launches the real one downstream:

{ "mcpServers": {
    "guarded-fs": {
      "command": "python", "args": ["-m", "deepsleuth", "proxy",
        "/path/to/real-server", "--policy", "policy.example.yaml", "--log", "gate.jsonl"]
    } } }

Try it on the bundled fixtures

python -m deepsleuth scan tests/fixtures/injection --no-dynamic          # source taint + hint violation
python -m deepsleuth scan tests/fixtures/poisoned  --no-dynamic          # poisoned descriptions
python -m deepsleuth scan tests/fixtures/supplychain --no-dynamic        # install-time hook + typosquat
python -m deepsleuth proxy-eval tests/fixtures/runtime --allow-unsandboxed  # response injection + cross-call leak, with gate decisions
python tests/run_all.py                                                    # unit + e2e tests (no pytest needed)

What it covers

Evidence locations — deepsleuth detects across all eight, with special strength on the five a manifest-only scanner misses:

Evidence location

Manifest-only sees it?

deepsleuth

description, name, schema

yes

✅ normalized mechanism rules + obfuscation

source

no

✅ AST taint, hint-vs-behavior, rug-pull gates, auth/audit

runtime-response

no

✅ response-injection + canary/credential leak scan

multi-call-state

no

✅ cross-call canary leakage, re-list diff, response diff

server-identity

rarely

✅ handshake vs. config/package identity

install-time-script

no

✅ npm/pip install-hook + typosquat analysis

Mechanism categories: tool-poisoning, agent-config-poisoning, tool-shadowing, prompt-injection, credential-exposure, command-injection, path-traversal, ssrf, data-exfiltration, confused-deputy, auth-misconfiguration, denial-of-service, excessive-privilege, supply-chain, information-disclosure, client-side-vulnerability, other.

Every finding validates against the fixed finding schema, carries a top-level evidence_location and confidence, and (from the proxy) records its gate decision on raw.gate_decision. See DETECTORS.md for one entry per detector including its known blind spots, and ARCHITECTURE.md for how the layers fit and how to add a detector.

Known limitations (v1)

  • Source analysis is Python-first. Node/TS servers get manifest + install-hook

    • dynamic coverage, but source taint is Python-only in v1 (JS is regex-lite).

  • Taint is intra-procedural. Flows through helper functions/classes across the module are approximated, not fully tracked.

  • The proxy fronts exactly one downstream server (v1 scope; multi-server namespacing is structured for but not built).

  • The live proxy's elicitation round-trip and forwarding of downstream-initiated requests are best-effort. All gate/audit/diff/response logic is fully exercised by proxy-eval, which is what the offline evaluator scores.

  • Without Docker, dynamic detectors are skipped (reported, not silent).

Getting involved

Contributions welcome — see CONTRIBUTING.md. Found a security issue? Please follow SECURITY.md.


Listed in the official MCP Registry:

mcp-name: io.github.DeepSleuth/deepsleuth

Available Tools

3 tools
check_listingA

Analyze a supplied MCP tool listing with the listing-phase detectors (description/schema poisoning, shadowing, obfuscation, ...) WITHOUT launching the server. Pass another server's tools/list output.

ParametersJSON Schema
NameRequiredDescriptionDefault
tools_jsonYesA JSON array of tool entries ({name, description, inputSchema}), or an object with a 'tools' key.
server_nameNoLabel for the supplied listing (default: 'supplied-listing').

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose one meaningful trait: the analysis runs without starting the server, so it is a side-effect-free, offline read. It says nothing about what the analysis produces, how detectors are selected, or any limits on listing size or processing time.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action and the key constraint front-loaded, followed by the input provenance. No filler or repetition of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, two-parameter static analysis tool with a fully described schema and no output schema, the description is adequate on inputs and safety. It leaves open how detectors are chosen relative to the list_detectors sibling, and gives no sense of the result shape despite the trailing '...' implying more detector coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the schema. The description adds only a modest hint about where tools_json comes from ('another server's tools/list output') and says nothing about server_name or its default label, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Analyze') and resource ('a supplied MCP tool listing') and enumerates the detector categories (description/schema poisoning, shadowing, obfuscation). The 'WITHOUT launching the server' clause implicitly separates it from scan_target, but the sibling is never named, so the differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear triggering context: use this on another server's tools/list output when you have a static listing and do not want to launch the server. There is no explicit when-not or named alternative (e.g., 'use scan_target to connect live'), so the routing is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_detectorsA

List every detector in the deepsleuth registry: id, category, evidence location, phase, and required capability layers.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the return content (id, category, evidence location, phase, required capability layers) and claims completeness ('every detector'), but says nothing about read-only semantics, permissions, ordering, or pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. The verb, scope, and payload are packed into one line, and every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter listing tool with no output schema, the description covers what an agent most needs: that it returns the full registry and which fields each entry carries. Gaps are limited to ordering/pagination and the read-only nature, which are minor here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which is the baseline-4 case; there is no parameter meaning for the description to add or omit. The enumerated field list is about output, not inputs, so it does not affect this dimension.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List every detector in the deepsleuth registry') and enumerates the returned fields, so the tool's function is unambiguous. It does not, however, differentiate itself from the siblings check_listing or scan_target.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use, when-not, or alternative is given. Usage is only implied: a full-registry listing is self-evidently a discovery step, but the description never states that intent or contrasts it with scan_target/check_listing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_targetB

Scan an MCP server target with the full detection core: source analysis, package manifests, and (optionally, dynamic=True) live execution inside the Docker sandbox. Returns the findings.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYesServer directory, mcp.json path, or raw stdio launch command.
dynamicNoAlso run the dynamic layer (Docker sandbox required; the target always runs sandboxed). Default: false.
timeoutNoPer-call timeout seconds (default 30).

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that live execution happens in a Docker sandbox and that the target always runs sandboxed, which is real safety context. It omits side-effect/read-write profile, auth or permission needs, and what a scan actually does to the target beyond execution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and layers, then the optional path and return. Nearly every clause earns its place; the parenthetical about dynamic is slightly redundant with the schema but reads naturally.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-param tool with no annotations and no output schema, the description covers what gets scanned and that findings are returned but leaves the return shape and safety profile thin. 'Returns the findings' is a placeholder rather than a useful expectation-setter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so target, dynamic, and timeout are already documented in the schema; the description only restates dynamic's behavior ('optionally, dynamic=True'). Baseline 3 is appropriate when the schema does the parameter work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Scan') and resource ('MCP server target') and enumerates the detection layers: source analysis, package manifests, dynamic execution. It doesn't explicitly name or distinguish itself from the siblings (check_listing, list_detectors), but the 'full detection core' scope makes the distinction self-evident.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the description pairs 'dynamic=True' with its requirement (Docker sandbox), which hints at when to enable the heavy path. There is no explicit guidance on when to reach for scan_target versus check_listing or list_detectors, nor stated prerequisites beyond the sandbox.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedcheck_listing
    • First observedlist_detectors
    • First observedscan_target

TDQS

A3.7/5.0

Scored across 3 tools

Disambiguation4/5

check_listing and scan_target both perform scanning, but descriptions clearly scope check_listing to listing-phase detectors without launching the server, while scan_target runs the full core with source/manifest analysis and optional Docker dynamic execution. list_detectors is clearly a registry enumeration. Boundaries are mostly distinct with only mild conceptual overlap.

Naming Consistency5/5

All three tools follow a consistent verb_noun snake_case pattern (check_listing, list_detectors, scan_target). No mixing of conventions, styles, or casing.

Tool Count4/5

Three tools is on the lean side but well-scoped for a focused MCP security scanner: a listing-only check, a full scan, and detector enumeration. Each earns its place, though a slightly richer surface could be expected for a detection framework.

Completeness4/5

Core workflows (static listing analysis, full source/manifest/dynamic scan, detector discovery) are covered, and check_listing vs scan_target provide both lightweight and deep paths. Minor gaps exist: no per-finding detail/retrieval, no report export, and no detector configuration management, but agents can work around these.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Self-hosted MCP gateway that applies deterministic, compiled policy to tool discovery, invocation, and outbound data flow, with no model in the enforcement path. Every decision emits a hash-chained receipt sealed with Ed25519 and verifiable using public keys only.
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enforces deterministic security policies as an inline firewall for MCP server tool calls, with AST-based validation, cryptographic audit logging, and CLI-based evaluation and verification.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables security auditing of MCP servers by running them in a sandbox with fake secrets, capturing outbound traffic, and detecting tool poisoning or secret exfiltration before approval.
    8 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Proxies MCP traffic between a client and a downstream server to enforce runtime policies on tool declarations, call arguments, and results, including allowlisting, sandboxing, secret and egress controls, and injection detection. It also includes a deterministic benchmark for measuring which security controls stop which attacks.
    MIT