Skip to main content
Glama

Local MCP Toolbox

Read-only inspection with explicit access controls.

A local Model Context Protocol server for checking files, Git changes, application logs, container health, and Python environments. You choose what it can inspect. Every request is authorized, sensitive output is redacted, results are bounded, and activity is audited.

Release Python Quality Security License Scope

Current release: v1.5.2. The v1.5 scope is complete. See the release contract and acceptance record.

Quick Start · Tools · Security · How It Works · Connect a Client · Architecture · Demo · Roadmap · Contributing


Limits in force

The server enforces the YAML file passed to --config. The shipped default, config/restricted.yml (the same policy values as config/default.yml), approves nothing:

  • No filesystem roots. File tools deny every path.

  • Docker, Git, GitHub, logs, scanners, infrastructure, incidents, Python environments, and external network are off.

  • File reads stop at 240,000 bytes. Directory scans stop at 500 entries. Responses stop at 100 records and 262,144 bytes. Subprocess and GitHub calls time out after 10 seconds.

  • Readable extensions are .md, .txt, .json, .yaml, .yml, .toml, .py, and .ts. Blocked names are .env, .env.*, id_rsa, id_ed25519, *.pem, *.key, and credentials*.

  • Home-directory paths are redacted. Email addresses and IP addresses are not, unless that policy option is turned on.

  • Audit events are capped at 8,192 bytes and segments at 8,388,608 bytes. Closed segments are kept for 30 days. A relative audit.path is resolved from the configuration file's directory, not the process working directory.

  • HTTP is disabled. profile: advanced is rejected.

Those are the running defaults, not the schema ceilings. An operator can raise several numbers inside a private profile; the ceilings, and the defaults that apply only after an integration is enabled, are in the security model. Pattern redaction is not a guarantee that every secret is recognized.

Related MCP server: AI Knowledge Center MCP

What This Is

Local MCP Toolbox supports development, security review, and troubleshooting workflows that need evidence without broad machine authority.

Use it to review changes in an approved repository, investigate recurring log errors, or check an unhealthy container. Each integration requires explicit configuration. Returned evidence helps you investigate; it does not establish a root cause by itself.

The exposed tools cannot edit your files, commit code, restart containers, execute arbitrary commands, or mutate remote systems. The server writes only its own security audit records. It runs locally and connects to an MCP client; there is no separate dashboard.

What You Get

Capability

What it does

Security boundary

System

Safe host metadata and developer-tool availability

No environment variables, usernames, process data, or executable paths

Filesystem

Approved-root listing, metadata, and text inspection

Canonical containment, sensitive-path blocklist, extension allowlist, bounded reads

Git

Repository status, branch, commits, and diff summaries

Explicit repository allowlist; fixed, non-interactive Git commands

GitHub

Repository, issue, and pull-request metadata

External-network opt-in, exact repository allowlist, fixed API origin, bounded GET requests

Python environments

Static virtual-environment dependency metadata audit

Separate roots, link-free fixed paths, no interpreter, Pip, subprocess, import, or network execution

Metrics

Aggregate request outcomes and latency

In-process counters only; no arguments, responses, identifiers, or secrets

Docker

Opt-in container metadata, health, and bounded logs

Official SDK only; no lifecycle, exec, mount, environment, or command access

Logs

Tails, literal search, and deterministic error grouping

Dedicated approved roots, output limits, and central redaction

Security

Bandit availability and normalized scan findings

Fixed scanner invocation; no user-controlled command arguments or fixes

Infrastructure

Project-type detection and top-level configuration inventory

Separate approved roots; no recursive content inspection

Incidents

Timestamped evidence and deterministic summaries

Read-only, bounded observations: never root-cause claims

Audit

Sanitized JSONL accountability trail

Shape-only request summaries, retention, and size limits

For parameters, output schemas, and every individual guardrail, see the full tool catalog.

Quick Start

Requirements

  • Python 3.12 or later

  • An MCP-capable client for connection after the server is validated

The default restricted profile is intentionally safe: it starts with no approved filesystem roots and no optional integrations.

Windows (PowerShell)

git clone https://github.com/chriswayneh/local-mcp-toolbox.git
Set-Location local-mcp-toolbox
python -m venv .venv
.\.venv\Scripts\python -m pip install -e ".[dev,docker]"
.\.venv\Scripts\local-mcp-toolbox doctor --config config\restricted.yml

macOS / Linux

git clone https://github.com/chriswayneh/local-mcp-toolbox.git
cd local-mcp-toolbox
python3 -m venv .venv
.venv/bin/python -m pip install -e ".[dev,docker]"
.venv/bin/local-mcp-toolbox doctor --config config/restricted.yml

doctor checks configuration and prerequisites without changing them. Its JSON status is ready only when every check passed or was skipped; a warning makes it attention, and the command still exits 0. Review any reported issues, then connect your client using the supplied template. The client starts the server when needed.

For a manual startup check, run local-mcp-toolbox serve --config config/restricted.yml using the executable in your virtual environment. It waits for MCP messages and does not open a browser. Press Ctrl+C to stop it before letting your client start its own instance.

The default profile has no approved file roots or optional integrations. Start by asking your client to call toolbox_server_status. Then follow getting started to grant only the access you need. The install above includes development and Docker support; Docker itself is optional and remains disabled until configured.

Security by Design

The design applies zero-trust principles and least privilege: each tool request is checked against local policy, and integrations receive only the access you explicitly configure. Local execution alone does not authorize access to a path, repository, or optional integration.

Control

Protection

Deny by default

The restrictive profile has no approved filesystem roots or integrations.

Approved roots

Canonical containment blocks arbitrary filesystem access and escape paths.

Read-only surface

No generic shell, mutation, commit, lifecycle, or remote-execution tool is registered.

Fixed subprocesses

External binaries use fixed argument templates, shell=False, scrubbed environments, timeouts, and output caps.

Central redaction

PEM blocks, credentials, cookies, authorization headers, connection strings, and home-directory paths are redacted before output. Email and IP redaction are off unless enabled.

Output bounds

File reads, collections, subprocess output, and responses are size-limited.

Sanitized audit

Requests record safe metadata, actual outcomes, and redaction counts. Raw secrets and tool output are excluded.

Explicit integrations

GitHub, Git, Docker, logs, scanners, infrastructure, and incident tools must be configured intentionally.

Untrusted evidence

Retrieved files, logs, commit messages, and metadata are treated as untrusted data.

The workstation, configured policy, installed dependencies, and any enabled Docker proxy remain trusted components. Read-only tools are not an operating-system sandbox, and the loopback HTTP token is not multi-user identity. These controls do not constitute a complete enterprise zero-trust architecture or security certification. See the verified limitations.

Read the security model, threat model, and the security-focused architecture decisions for the complete rationale.

How It Works

  1. An MCP client requests one registered tool.

  2. The toolbox validates typed inputs and bounded parameters.

  3. Permissions, approved roots, and integration allowlists are checked.

  4. A narrow read-only operation collects the permitted data.

  5. Results are redacted and bounded before they cross the MCP boundary.

  6. Sanitized request metadata is recorded in the audit log.

  7. The client receives a safe structured result or error.

Connect a Client

The repository includes stdio configuration templates for the clients below. Template parsing and the installed server's launch contract are tested; individual desktop application versions are not certified. Adding a client entry lets the client start the process: it does not grant the server broader permissions.

Replace the intentionally unresolved paths, then configure the smallest local policy that serves the task. See client configuration for exact installation notes and the important separation between client startup and server authorization.

See It Safely

This project is designed for evidence, not a dashboard. The synthetic demo walkthrough provides a reproducible way to see the policy boundary in action without real credentials, repositories, production logs, or a host Docker socket.

It demonstrates a safe inspection sequence:

toolbox_server_status          → verify the server and active profile
logs_tail_file                 → view redacted synthetic log evidence
logs_error_summary             → group observed errors without causal claims
infra_detect_project_types     → inspect demo project metadata
docker_unhealthy_containers    → observe an intentionally unhealthy demo service

The demo’s fabricated token is redacted, disabled integrations return a structured denial, and its audit trail contains sanitized metadata only. Follow the walkthrough to run it locally.

Architecture

flowchart LR
  Client["MCP client"] --> Transport["stdio transport"]

  subgraph Boundary["Local policy enforcement boundary"]
    Registry["MCP server / tool registry"] --> Permission{"Permission check"}
    Permission -->|Denied| Error["Safe structured error"]
    Permission -->|Allowed| Tool["Narrow read-only tool"]
    Tool --> Guard["Redaction + output limits"]
  end

  Transport --> Registry
  Guard --> Client
  Registry -. "sanitized metadata" .-> Audit["JSONL audit log"]
  Tool --> Integration["Explicitly approved local integrations"]

  classDef boundary fill:#EAF3FF,stroke:#4A78A8,color:#102A43
  classDef control fill:#E9F7EF,stroke:#2E7D32,color:#173E22
  classDef denial fill:#FDECEC,stroke:#C62828,color:#5C1111
  class Registry,Tool,Guard boundary
  class Permission,Audit,Integration control
  class Error denial

All retrieved content remains untrusted data. The full component model and trust-boundary discussion live in architecture.

Repository Structure

src/mcp_toolbox/  MCP server, tool modules, permissions, redaction, audit, config, and CLI
tests/            Unit, integration, and security regression tests
config/           Restricted, standard, and container policy profiles
docs/             Architecture, threat model, operating guides, ADRs, and tool reference
examples/         MCP client configuration templates
demo/             Synthetic services, logs, and intentionally insecure test fixtures
.github/          CI, security, documentation, release, Dependabot, and contribution templates

Documentation

Document

Purpose

Architecture

System design, components, and data flow

Security Model

Controls and trust boundaries

Threat Model

Threat analysis and mitigations

Version 1.5 Security Review

Findings, remediation, verification, and residual responsibilities

Permissions

Authorization sequence and profile behavior

Tool Catalog

Inputs, outputs, and module-level guardrails

Client Configuration

Codex, Claude, and VS Code setup

Authenticated HTTP

Optional loopback transport and bearer-token controls

Docker

Hardened container profiles and socket-proxy guidance

Demo Walkthrough

Synthetic end-to-end policy demonstration

CI and Release

Quality, security, docs, package, and release controls

Release Contract

Supported v1.5 scope, acceptance evidence, operations, and limitations

Roadmap

Completed v1.5 scope and optional future proposals

Project Status

Version 1.5 adds allowlisted GitHub inspection, content-free runtime metrics, a static Python environment auditor, authenticated loopback HTTP, crash-safe audit rotation, and hardened release controls. Kubernetes inspection and generation features were removed after security review because their effective behavior could not satisfy the inspection-only contract.

The v1.5 feature scope is complete. Version 1.5.2 closes default-parameter and container setup defects without adding capabilities. Support is limited to the local inspection contract in the acceptance record, not a hosted service, multi-user security boundary, or production availability guarantee. The package classifier Production/Stable means that local contract, not a hosted production service. Versions 2 through 4 are optional proposals, not unfinished release requirements.

Contributing and Security

Contributions are welcome when they preserve the project’s least-privilege model. Start with CONTRIBUTING.md, use the repository templates for bugs and feature proposals, and report vulnerabilities through the process in SECURITY.md.


License

Licensed under the MIT License. Use it, fork it, modify it, or build something of your own. See LICENSE for the terms.


If this project helped you, a ⭐ is appreciated.

Built with

Python · Model Context Protocol · MCP Python SDK · Pydantic · Typer · Docker

Available Tools

8 tools
disk_usageDisk usageA
Read-onlyIdempotent

Return disk capacity for the server's current working volume.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered without the description. The description's one added behavioral fact is the scope limit ('current working volume'), which is useful but thin; it says nothing about failure behavior if the volume is unavailable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that leads with the action and immediately qualifies the scope. No filler, no restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't explain return values, and a zero-parameter tool needs little else. The only gap is the absence of any hint about when this beats system_info, which for a tool with six siblings is a real but modest omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. The description correctly implies the call is argument-free.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Return') and resource ('disk capacity') and narrows the scope to the server's current working volume, which is more precise than the bare title. It does not, however, explicitly distinguish itself from siblings like system_info or toolbox_metrics_snapshot, which an agent might also reach for when asked about server health.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no exclusion of alternatives, and no note that this reports only one volume rather than all mounted filesystems. Usage is only weakly implied by the phrase 'current working volume'; an agent comparing this against system_info gets no help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

filesystem_file_metadataApproved file metadataA
Read-onlyIdempotent

Return safe metadata for an approved readable file without reading its contents.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered structurally. The description adds one useful behavioral detail — that it does not read file contents — but does not disclose what 'safe metadata' actually contains or how it handles non-approved/unreadable paths.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with the key constraint (no content read) front-loaded and no wasted words. Appropriately sized for a one-parameter read tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. The safety profile is covered by annotations. The main residual gap is the undefined term 'approved', but for a low-complexity read tool the description is essentially sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single 'path' parameter and 0% schema description coverage, the description must carry the semantics. 'Approved readable file' hints that the path must reference a pre-approved file, adding marginal meaning, but it does not clarify path format or error behavior for unapproved paths.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Return) and resource (metadata for an approved readable file), and the phrase 'without reading its contents' clearly distinguishes it from the sibling filesystem_read_text_file. It doesn't name the sibling explicitly, but the scope is unambiguous enough for selection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'approved readable file' implies a precondition (the file must be approved/readable), which gives some contextual guidance. However, there is no explicit when-to-use vs. when-not statement and no alternative tool named, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

filesystem_list_directoryList approved directoryA
Read-onlyIdempotent

List immediate entries in an approved directory with bounded pagination.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
limitNo
offsetNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered structurally. The description usefully adds 'immediate entries' (non-recursive) and 'bounded pagination', which the annotations do not convey, but says nothing about error behavior for unapproved paths or ordering guarantees.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and resource, then the two most important scope constraints. Every word earns its place with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations carry the full safety profile. The remaining gap is parameter semantics, which is minor for a three-parameter listing tool where the names are self-evident.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description provides no meaning for any of the three parameters: it never explains that 'path' is the required target, that 'limit' defaults to 50, or that 'offset' controls pagination position. 'Bounded pagination' only vaguely gestures at limit/offset without defining their semantics or bounds.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (immediate entries in a directory) with scope qualifiers ('immediate', 'approved', 'bounded pagination') that clarify it is non-recursive and sandboxed. It is clear on its own, but does not name or differentiate itself from siblings like filesystem_file_metadata or filesystem_read_text_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'approved directory' implies a precondition (only allowlisted paths are valid), which is useful implied guidance. However, there is no explicit when-to-use vs. when-to-use-an-alternative, and no mention of what to do if a path is rejected or how this differs from reading file metadata.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

filesystem_read_text_fileRead approved text fileB
Read-onlyIdempotent

Read a bounded approved text file after path, type, size, and secret controls.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false and destructiveHint=false, so the safety profile is covered. The description adds real context beyond that by revealing an approval/allowlist gate plus type, size, and secret filtering, but it is vague about what happens on rejection or truncation and gives no error/rate-limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no padding, and the gating constraint is stated immediately after the verb. It is terse to the point of being cryptic ("path, type, size, and secret controls"), which costs it the top mark.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and annotations carry the safety profile. What is still missing for a control-gated reader is what "approved" concretely means and how failures surface, leaving the agent to guess at failure modes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single path parameter has no description in the schema, so the description must compensate. It mentions "path" only as a control checkpoint and never explains whether the path is absolute or relative, how it is resolved, or whether it must fall inside an allowed root.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read) and a scoped resource (bounded approved text file), which is enough to separate it from filesystem_list_directory and filesystem_file_metadata. It does not explicitly name or contrast any sibling, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"after path, type, size, and secret controls" implies that reads are gated by approval/allowlist checks, which is a usable precondition hint. However, it never states when to prefer this tool over siblings or what to do if a file fails those controls, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

installed_developer_toolsInstalled developer toolsA
Read-onlyIdempotent

Report availability of an allowlisted set of common developer tools.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world behavior, so the safety profile is covered. The description adds the useful scope constraint that the tool reports only an 'allowlisted' set rather than all installed tools, but says nothing about which tools or how availability is determined.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no redundant clauses; every word earns its place and the scope constraint is stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only probe with an output schema present, the description need not explain return values. It is nearly complete, with the only gap being which tools are on the allowlist and any platform/availability caveats.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies. There is nothing parameter-related for the description to clarify or compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Report') and resource ('availability of an allowlisted set of common developer tools'), which is clearer than a tautology. It is distinguishable from filesystem and toolbox siblings, though it does not explicitly contrast with the nearby system_info tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use context, no prerequisites, and does not name any alternative or boundary against siblings like system_info. The agent must infer the use case entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

system_infoSystem informationA
Read-onlyIdempotent

Return sanitized OS and runtime metadata without reading environment variables.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds a genuine behavioral trait beyond that: output is 'sanitized' and environment variables are deliberately not read, which tells the agent this is a privacy-safe diagnostic call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the action and the key constraint with no filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and zero parameters means no schema gaps. The description covers purpose and the sanitization boundary; it could say marginally more about what 'runtime metadata' excludes, but nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which per the rubric sets a baseline of 4. The description correctly implies a no-argument call and needs to add no further parameter detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Return') and resource ('OS and runtime metadata'), with a distinguishing scope qualifier ('sanitized ... without reading environment variables') that separates it from data-gathering siblings like installed_developer_tools. It does not explicitly name a sibling, but the scope constraint makes the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus toolbox_server_status, toolbox_metrics_snapshot, or disk_usage, all of which also return system-level data. The 'without reading environment variables' clause hints at a privacy boundary but does not frame it as a use/don't-use condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

toolbox_metrics_snapshotLocal MCP Toolbox metricsA
Read-onlyIdempotent

Return content-free aggregate request and latency metrics.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false and destructiveHint=false, so the safety profile is covered structurally. The description adds one genuine behavioral fact beyond them: the metrics are 'content-free' aggregates, which tells the agent no per-request payloads or sensitive content are exposed. It does not disclose the aggregation window, sampling, or refresh behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero filler, front-loading the verb and the two things measured. Nothing could be removed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value structure need not be described, and there are no parameters to document. The one remaining gap is the scope of the aggregate (time window or granularity), which an agent would need to interpret the numbers, though that is plausibly captured by the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a no-parameter tool applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Return') and a specific resource ('aggregate request and latency metrics'), and the 'content-free' qualifier separates it from payload-returning siblings like filesystem_read_text_file. It does not explicitly position itself against the closest sibling, toolbox_server_status, but the resource is distinct enough to identify the tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this tool, no prerequisites, and no mention of alternatives such as toolbox_server_status. Usage is only inferable from the tool name and the word 'metrics'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

toolbox_server_statusLocal MCP Toolbox server statusB
Read-onlyIdempotent

Return server-generated, redaction-safe status metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds one genuine behavioral detail beyond that — the output is "redaction-safe" — but says nothing about freshness, caching, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no padding and the key qualifiers (server-generated, redaction-safe) front-loaded. It is efficient, though so terse that brevity shades into under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and there are no parameters to document. The remaining gap is routing: the description never says how this status differs from the several sibling status/info tools, which is the main thing an agent needs here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there is nothing for the description to disambiguate. Schema coverage is 100% and the empty argument object is unambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

It gives a verb ("Return") and a resource ("status metadata"), but the resource is abstract and largely restates the tool name and title. Nothing distinguishes it from sibling tools such as system_info or toolbox_metrics_snapshot, so an agent cannot tell which status surface this covers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no condition selecting this over system_info, disk_usage, or toolbox_metrics_snapshot, and no exclusions. The agent must guess based on the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv1.5.2
    • First observeddisk_usage
    • First observedfilesystem_file_metadata
    • First observedfilesystem_list_directory
    • First observedfilesystem_read_text_file
    • First observedinstalled_developer_tools
    • First observedsystem_info
    • First observedtoolbox_metrics_snapshot
    • First observedtoolbox_server_status

TDQS

A3.6/5.0

Scored across 8 tools

Disambiguation4/5

Most tools have clearly distinct targets: OS metadata, disk capacity, dev-tool availability, and three filesystem operations (list/metadata/read) are well separated. The only mild overlap is among system_info, toolbox_server_status, and toolbox_metrics_snapshot, which all return 'metadata/status' of a system or server, though the descriptions do differentiate them.

Naming Consistency4/5

All names use snake_case, which is consistent. However, the set mixes three prefixing conventions: unprefixed tools (system_info, disk_usage, installed_developer_tools), a filesystem_ group, and a toolbox_ group, which makes the namespace slightly less predictable than a single uniform scheme.

Tool Count5/5

Eight tools is well within the ideal 3-15 range and each one maps to a distinct read/introspection capability. Nothing feels redundant or padded, and no obvious operation is missing due to under-provisioning.

Completeness4/5

For a deliberately read-only, safety-bounded local toolbox, the surface covers system info, disk, dev tools, and the core filesystem read lifecycle (list, metadata, read). Write/search operations are absent, but the descriptions imply a read-only design intent, so this is a minor rather than significant gap.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Secure local development platform that exposes controlled developer capabilities (FS, Git, search, command execution) to AI assistants via MCP with deny-by-default security and audit logging.
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Read-only MCP server to inspect allowlisted Docker containers, systemd services, JSONL logs, and HTTP health endpoints without arbitrary shell access.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    A local-first MCP server providing guarded access to a workspace with file operations, search, commands, tests, Git helpers, checkpoints, and structured tool results. It supports multiple tool modes and emphasizes security with workspace restrictions and secret blocking.
    1
    MIT