Skip to main content
Glama

mcp-secure-reports

An MCP server that lets an AI agent export reports from a system that has no API, without handing the agent more access than the person using it.

I built the first version for a hospital whose information system only has a web interface. This repository is a cleaned up public version: same design, synthetic data, no hospital names or real credentials.

Architecture

What it does

The agent (Claude Desktop, Claude Code or any MCP client) can list reports, start an export, poll until the file is ready, and look up insurance claims. Behind the server, a backend logs in to the legacy system and downloads the file.

Three things make it safe to point at sensitive data:

Tools depend on the role. One server process serves one role. Tools that role may not use are never registered, so the model does not see them in tools/list at all. Report types are filtered the same way. Permissions live in config/roles.json, so a change is a reviewed pull request, not a code change.

Every call is audited. Each tool call, allowed or denied, goes to an append only JSONL log with the role, user, parameters and duration. Fields such as patient_id are masked before they are written, including inside nested objects.

Credentials never reach the model. The login for the legacy system is stored in a Fernet encrypted vault. Only the backend reads it. No tool returns it, and each read is itself an audit event.

Related MCP server: legacy-mcp

Roles in the sample config

Tool

reporting_staff

unit_manager

finance

list_available_reports

✓

✓

✓

download_report

✓

✓

✓

get_report_job_status

✓

✓

✓

check_claim_status

✓

✓

download_billing_statement

✓

Try it

Requires Python 3.10 or newer.

python -m venv .venv
.venv/bin/pip install -e ".[dev]"        # Windows: .venv\Scripts\pip
.venv/bin/python scripts/demo.py

The demo starts the server twice over stdio, once as reporting_staff and once as finance, and prints what each role sees:

=== Connected as reporting_staff ===
  tools/list: download_report, get_report_job_status, list_available_reports
  > download_report({"report_type": "monthly_visits", ...})
    ok: {"job_id": "f1cf76f6...", "status": "queued", "next": "poll get_report_job_status"}
  > download_report({"report_type": "insurance_claims", ...})
    error: Error executing tool download_report: Report type 'insurance_claims' is not allowed for this role

=== Connected as finance ===
  tools/list: check_claim_status, download_billing_statement, download_report, get_report_job_status, list_available_reports

=== audit.jsonl ===
  reporting_staff  download_report     failed   {"report_type": "insurance_claims", ...}
  finance          check_claim_status  success  {"claim_id": null, "patient_id": "00******78"}

Run the tests:

.venv/bin/python -m pytest

Use it from Claude Desktop

Generate a vault key once:

python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())"

Then add the server to claude_desktop_config.json:

{
  "mcpServers": {
    "secure-reports": {
      "command": "/path/to/mcp-secure-reports/.venv/bin/secure-reports-mcp",
      "env": {
        "MCP_ROLE": "reporting_staff",
        "VAULT_MASTER_KEY": "paste the key here",
        "DATA_DIR": "/path/to/mcp-secure-reports/data"
      }
    }
  }
}

Ask Claude for "the monthly visits report for September 2026" and it will start the export, poll the job and tell you where the file is. Open data/audit.jsonl to see what was logged.

Configuration

Variable

Default

Purpose

MCP_ROLE

reporting_staff

Role served by this process

MCP_USER

demo_user

User id written to the audit log

VAULT_MASTER_KEY

none, required

Fernet key for the credential vault

BACKEND

mock

mock for synthetic data, playwright for a real system

MOCK_DELAY_SECONDS

2

How long a mock export takes

DATA_DIR

./data

Audit log, vault file and exports

CONFIG_PATH

config/roles.json

Role and tool definitions

Project layout

src/secure_reports_mcp/
  server.py    registers tools for the active role, audits every call
  tools.py     tool logic, no MCP or logging concerns
  backend.py   MockBackend (synthetic data) and the PlaywrightBackend interface
  vault.py     Fernet encrypted credential store
  audit.py     append only JSONL log with field masking
  config.py    loads and validates config/roles.json
scripts/demo.py   end to end run over the MCP protocol
tests/            35 tests: masking, vault, config, role filtering, job lifecycle

What is mocked

MockBackend runs the full flow, including the vault read and the async job, but writes a small CSV of made up numbers instead of logging in anywhere. PlaywrightBackend defines where real browser automation goes and raises NotImplementedError, because selectors depend on the target system. In the hospital deployment the next step is a read only database backend with a fixed set of queries per domain, rather than free form SQL.

Role selection by environment variable is fine for a single user on a desktop. A shared deployment should take the role from an authenticated session instead.

License

MIT

Available Tools

3 tools
download_reportB

Start exporting a report. date_range is {"from": "YYYY-MM-DD", "to": "YYYY-MM-DD"}. Returns a job_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
date_rangeYes
report_typeYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the important trait that this is an async export returning a job_id rather than the report content, but says nothing about permissions, whether the export is durable, or how/when to poll for completion.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action; nothing is bloated. Slight redundancy in stating the return value as its own sentence, but it is high-value information given no output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because there is no output schema, naming job_id is appropriately included, and the nested date_range format is covered. Still missing the source of report_type values, any async/polling workflow, and any safety or auth context, which matters for an unannotated tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and date_range is an unconstrained object, so the description's explicit {'from': 'YYYY-MM-DD', 'to': 'YYYY-MM-DD'} format is genuinely additive. However, report_type is undocumented in both schema and description, leaving the agent with no hint about valid values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Start exporting a report') and, via 'Returns a job_id', signals it is an async kickoff rather than a synchronous file fetch. It does not name or contrast with the siblings, so an agent must infer that get_report_job_status is the follow-up.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when/when-not guidance and no mention of the sibling tools. The agent is left to infer that report_type comes from list_available_reports and that job_id is polled via get_report_job_status, neither of which is stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_report_job_statusB

Check whether an export job has finished and where the file is.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the return content (finished? + file location), but says nothing about permission requirements, polling frequency, status values, or whether the file expires.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the two key outputs front-loaded. No padding, though it is arguably under-specified rather than optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Without an output schema or annotations, the description must carry more. It conveys the core return contract but omits status enumeration, error/not-found behavior, and how the resulting file location pairs with download_report.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions the job_id parameter. While 'job_id' is fairly self-explanatory as a name, the description adds no format, sourcing, or usage detail to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Check') and resource ('export job'), and names the two pieces of information returned: completion state and file location. It is clearly distinguishable from list_available_reports and download_report, though it does not explicitly reference them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given. An agent can infer this is a polling tool to run after starting an export, but the description never states that, nor does it relate the tool to download_report (which presumably consumes the file location this tool returns).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_available_reportsA

List the report types this role may download. domain is 'management' or 'billing'.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses that results are filtered by the caller's role permissions ('this role may download'), which is real behavioral context. It does not state that the operation is read-only, whether the list is paginated, or what is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the primary purpose and followed immediately by the one detail an agent cannot get from the schema. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, read-only listing tool with no output schema, the description covers purpose and the critical enum values. The remaining gap is minor: no statement of read-only nature or return shape, but nothing essential is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema declares 'domain' as a bare string with no enum, so the description is the only source of parameter meaning. It supplies the allowed values ('management' or 'billing'), which is essential to calling the tool correctly, though it does not explain what each domain covers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (list) and resource (report types) with a clear scope qualifier ('this role may download'). It implicitly distinguishes the tool from download_report and get_report_job_status by being the discovery step, but it never names them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: an agent would call this to discover what is downloadable before invoking download_report. However, there is no explicit when-to-use statement, no mention of prerequisites, and no reference to the sibling tools as alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observeddownload_report
    • First observedget_report_job_status
    • First observedlist_available_reports

TDQS

A3.7/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: listing available reports, initiating an export job, and polling job status. There is no overlap in resource or action, so an agent can easily choose the right tool.

Naming Consistency5/5

All three tools use consistent snake_case with verb-led names (list_..., download_..., get_...). The pattern is predictable and readable across the set.

Tool Count5/5

Three tools are well-scoped for a focused report-export server. Each earns its place in the core workflow without being excessive.

Completeness4/5

The set covers listing, starting an export, and checking job status, which forms a complete minimal lifecycle. Minor gaps exist (no cancel job or list jobs operation), but agents can work around them for the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    A
    maintenance
    A governance proxy for AI tools — every MCP/agent tool call is policy-gated, secret-redacted, and written to a hash-chained, offline-verifiable audit trail.
    13
    MIT
  • F
    license
    C
    quality
    C
    maintenance
    Wraps a legacy SOAP + stored-procedure backend with governed MCP tools, including a semantic data dictionary, compensation for transactionless writes, and load protection, enabling AI agents to safely operate on enterprise systems.
    10
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables controlled AI-agent access to enterprise-shaped tools with a deny-by-default gated write path, human approval, dry-run execution, and append-only audit logging.
    1
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enforces authenticated identity on every tool call and SSE frame, rotates vaulted credentials in place, and restricts tools via allowlists.
    71 npm
    MIT