mcp-hardened
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-hardenedsearch for .env files in the project root"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP-HARDENED
A deliberately small MCP server built to demonstrate security hardening against the OWASP MCP Top 10, with a test suite that proves the defenses hold.
Warning
This repository contains intentionally vulnerable code. Commits at and before the
naive-baselinetag implement path traversal, SQL injection, and SSRF on purpose — they are the "before" half of a security demonstration. The build order is deliberate: write the naive version, write attack tests that prove the vulnerability is reachable, then harden until the tests pass. Do not use any code from this repository as a reference implementation, and do not deploy it. This is a teaching artifact, not a library.
Ongoing project. Things may change.
Related MCP server: Vulnerable MCP Server
What this demonstrates
Three read-only tools, each hosting one vulnerability class (plus ping, a
smoke test):
Tool | Vulnerability class | Defense |
| Path traversal | Resolve path, assert inside sandbox root |
| SQL injection | Parameterized queries only |
| SSRF | Host allowlist, revalidated after redirects |
Full breakdown of each tool in docs/TOOLS.md.
Audit logging
Every tool invocation is written to logs/audit.jsonl as one JSON object per
line:
{"timestamp": "...", "tool": "search_files", "arguments": {"query": "../OUTSIDE_SANDBOX.txt"}, "outcome": "rejected", "error": "path escapes the sandbox root"}Fields: timestamp (ISO 8601 UTC), tool, arguments as received, outcome
(ok or rejected), error on rejection, result_len on success.
Return values are never logged — only their length. Logging content would reintroduce the over-sharing problem MCP10 addresses.
JSON Lines is not just convenient. json.dumps escapes newlines, so an argument
containing \n cannot forge a second log entry. A plaintext format would have
been forgeable. tests/test_audit.py enforces this.
The log contains attacker-controlled text by design — that is the point of an audit log — so anything consuming it must treat it as untrusted input.
Writes fail open: if the log cannot be written the tool call still succeeds, with the failure reported on stderr. A compliance context would invert this. See residual risk in the coverage doc.
The log is gitignored — it is generated data, like data/*.db.
Before and after
The same test suite, run against the naive implementation and against the hardened one:
docs/attack-suite-before.txt— attack tests failing, functional tests passingdocs/attack-suite-after.txt— everything passing
The hardening itself, as a diff:
naive-baseline...master
Explicitly out of scope
Prompt injection — unsolved industry-wide, not attempting
Tool poisoning — a consuming-side threat; a server cannot prevent a client from connecting to a poisoned server
DNS rebinding — the host allowlist checks the name, not the resolved IP
Multi-user authentication, rate limiting
OWASP MCP Top 10 coverage
ID | Item | Status |
MCP01 | Token mismanagement / secret exposure | Addressed |
MCP02 | Privilege escalation via scope creep | Addressed |
MCP03 | Tool poisoning | Not addressable at the server layer |
MCP04 | Supply chain | Partially addressed |
MCP05 | Command injection | Addressed — the focus of this project |
MCP06 | Intent flow subversion | Not addressable at the server layer |
MCP07 | Insufficient authentication / authorization | Out of scope for this architecture |
MCP08 | Lack of audit and telemetry | Addressed |
MCP09 | Shadow servers | Not addressable at the server layer |
MCP10 | Context injection / over-sharing | Addressed |
Full reasoning per item in docs/OWASP-MCP-COVERAGE.md
Stack
Python 3.12 · FastMCP 3.3.1 · pytest · SQLite · stdio transport
Running it
uv sync
uv run scripts/seed_db.py # creates data/records.db, needed by query_records
uv run pytest -vTo poke at the tools by hand:
npx @modelcontextprotocol/inspector uv run src/server.pyClaude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"hardened-demo": {
"command": "wsl.exe",
"args": [
"-d", "Ubuntu",
"--",
"/home/nullstep/.local/bin/uv",
"run",
"--directory", "/home/nullstep/mcp-hardened",
"src/server.py"
]
}
}
}Config location: %APPDATA%\Claude\claude_desktop_config.json on Windows,
~/Library/Application Support/Claude/ on macOS. Microsoft Store installs
redirect this into the package container under
%LOCALAPPDATA%\Packages\Claude_*\LocalCache\Roaming\Claude\.
Two non-obvious requirements when the server runs in WSL:
uvneeds an absolute path.wsl.exe -- <cmd>does not run a login shell, so~/.local/binis never added to PATH. A bareuvfails with exit 127.--directoryis required.wsl.exeinherits and translates the Windows working directory, so the process starts in/mnt/c/Windows/System32and relative paths do not resolve.
Restart Claude Desktop fully after editing — quit from the system tray, not just closing the window.
Note: ROOT = Path(__file__).resolve().parent.parent in src/server.py means
the sandbox, database, and log paths resolve correctly regardless of working
directory. --directory and ROOT are independent layers — the config could be
wrong and containment would still hold.
Attribution
Built with Claude Code. The implementation, threat model, and documentation in this repository were AI-generated.
My contribution was direction and review: setting the scope, deciding what stays out of scope, approving or rejecting each step, and verifying behavior in the MCP Inspector at each stage.
Available Tools
4 toolsfetch_docARead-onlyIdempotent
Fetch the text of a document at a URL.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds the behavioral detail that the output is 'text', which is useful context. However, it does not disclose error handling, URL format expectations, or other runtime behaviors beyond this.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loaded with the verb, and contains no unnecessary words. Every element contributes to understanding the tool's core function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter) and the existence of an output schema, the description is minimally adequate for basic invocation. However, it lacks usage guidance and behavioral edge cases, making it incomplete for a fully informed agent decision.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one parameter 'url' with no description. The description clarifies that the URL points to a document, providing essential meaning. However, it does not specify URL schemes, encoding, or other constraints, so compensating for 0% schema coverage is only partially achieved.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'fetch' and clearly identifies the resource (a document at a URL) and the output (text). This distinguishes it from siblings like search_files or query_records, which are search-oriented rather than direct fetches.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. There is no mention of scenarios, exclusions, or preconditions, leaving the AI without context for selecting this tool over its siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pingARead-onlyIdempotent
Echo a message back. Smoke test only.
| Name | Required | Description | Default |
|---|---|---|---|
| message | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safe read-only behavior is covered. The description adds the context that this is a smoke test, confirming it is a lightweight utility with no side effects. It does not contradict any annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only two short sentences, front-loaded with the primary action and followed by the usage caveat. It contains no fluff and every word contributes to understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a trivial tool with a single parameter, strong annotations, and an output schema present. The description fully covers its purpose and usage. The 'smoke test only' caveat is sufficient guidance for when to use it. No additional details are needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has only one parameter ('message') with no description, so the tool description must clarify its meaning. 'Echo a message back' implies the message parameter is the text to be echoed, but it does not provide additional details such as format, length limits, or echo behavior. The description offers minimal compensation for the 0% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Echo a message back.' The phrase 'Smoke test only' clarifies its limited scope, distinguishing it from siblings like search_files and query_records which perform substantive operations. This is a specific verb+resource with clear scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Smoke test only,' which tells the agent this tool is intended solely for connectivity checks and not for real-world data retrieval. While it does not name alternative tools, the context is clear enough that the agent should not use this for substantive tasks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_recordsARead-onlyIdempotent
Look up records by category.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds minimal behavioral context ('by category'), but does not disclose pagination, matching behavior, or limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no redundant words, front-loading the purpose effectively.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple single-parameter read-only tool with an output schema present, the description is largely complete; however, it lacks usage guidance and any detail about filtering behavior, so it is adequate but not rich.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has only one required parameter 'filter' with 0% description coverage. The description compensates by indicating that the filter is a category, giving meaning to the otherwise generic string parameter, though it does not specify matching semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('look up') and resource ('records') with a scoping criterion ('by category'), clearly distinguishing it from sibling tools like search_files (files) and fetch_doc (documents).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives such as search_files or fetch_doc. The description implies it is for record lookup but does not mention exclusions or alternative tool usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_filesDRead-onlyIdempotent
Read a note from the sandbox directory.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the operation as read-only and non-destructive, so the description adds little beyond noting the 'sandbox directory' scope. It fails to explain how the query is used or what behavior results, especially given the mismatch between the name and description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short (one sentence), but this is under-specification rather than conciseness. It lacks essential information about the query parameter and the tool's actual purpose, making it ineffective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema and only one parameter, the description is incomplete due to the name/description contradiction, lack of parameter semantics, and no usage context. It does not sufficiently leverage existing structural information or provide additional necessary context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description entirely omits the meaning of the 'query' parameter. The phrase 'Read a note' does not clarify how a query string factors into the operation, leaving the only parameter completely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Read a note') and a location ('sandbox directory'), but this conflicts with the tool name 'search_files', which suggests a search operation. It does not distinguish from siblings like fetch_doc or query_records, and the actual purpose is unclear given the query parameter.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus the sibling tools (ping, query_records, fetch_doc). It only states a generic action without context or exclusions, leaving the agent without direction for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
fetch_doc - First observed
ping - First observed
query_records - First observed
search_files
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: ping for smoke testing, search_files for reading a note, query_records for searching records, and fetch_doc for retrieving a URL. There is no overlap or ambiguity between them.
Three tools follow a verb_noun pattern (search_files, query_records, fetch_doc), but 'ping' is a standalone command and 'search_files' is misleading since it actually reads a note rather than searching files. This mixing of conventions reduces consistency.
The four tools form a compact and well-scoped set for a lightweight server, covering a smoke test, file access, record lookup, and URL fetching without unnecessary bloat.
The toolset offers read operations for different resources but lacks discovery methods (e.g., listing notes, categories, or available URLs) and any write capabilities. This limits the ability to perform full workflows without prior external knowledge.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MEOK MCP Hardening MCP — automated security red-team for any MCP server. Maps OWASP LLM Top 10
Security scanner for MCP servers. Detect vulnerabilities, prompt injection, and tool poisoning.
Experimental MCP server for current empirical verification of explicit public HTTPS endpoint claims.
Trust-scored search engine for MCP servers. 1,900+ sources indexed. IETF draft published. Referenced by OWASP MCP Security Cheat Sheet. L0-L4 trust levels based on cryptographic verification.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceA deliberately vulnerable MCP server that allows clients to interact with a database for educational purposes, demonstrating security vulnerabilities including SQL injection, arbitrary code execution, and sensitive data exposure.4-
- FlicenseNot gradedqualityDmaintenanceAn educational MCP server demonstrating common security vulnerabilities like command injection, path traversal, SQL injection, and XXE attacks. Designed for security training purposes only, not for production use.-
- FlicenseNot gradedqualityDmaintenanceA deliberately insecure MCP server designed as a pentest lab to demonstrate common vulnerabilities in MCP deployments.-
- FlicenseNot gradedqualityCmaintenanceAn intentionally vulnerable MCP server designed as a live demo target for the MCP Trust security scanner. It contains deliberate insecure patterns to demonstrate scanning capabilities.-