local-mcp-toolbox
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@local-mcp-toolboxSummarize the last 50 lines of the app server logs"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Local MCP Toolbox
Read-only inspection with explicit access controls.
A local Model Context Protocol server for checking files, Git changes, application logs, container health, and Python environments. You choose what it can inspect. Every request is authorized, sensitive output is redacted, results are bounded, and activity is audited.
Current release: v1.5.2. The v1.5 scope is complete. See the release contract and acceptance record.
Quick Start · Tools · Security · How It Works · Connect a Client · Architecture · Demo · Roadmap · Contributing
Limits in force
The server enforces the YAML file passed to --config. The shipped default, config/restricted.yml (the same policy values as config/default.yml), approves nothing:
No filesystem roots. File tools deny every path.
Docker, Git, GitHub, logs, scanners, infrastructure, incidents, Python environments, and external network are off.
File reads stop at 240,000 bytes. Directory scans stop at 500 entries. Responses stop at 100 records and 262,144 bytes. Subprocess and GitHub calls time out after 10 seconds.
Readable extensions are
.md,.txt,.json,.yaml,.yml,.toml,.py, and.ts. Blocked names are.env,.env.*,id_rsa,id_ed25519,*.pem,*.key, andcredentials*.Home-directory paths are redacted. Email addresses and IP addresses are not, unless that policy option is turned on.
Audit events are capped at 8,192 bytes and segments at 8,388,608 bytes. Closed segments are kept for 30 days. A relative
audit.pathis resolved from the configuration file's directory, not the process working directory.HTTP is disabled.
profile: advancedis rejected.
Those are the running defaults, not the schema ceilings. An operator can raise several numbers inside a private profile; the ceilings, and the defaults that apply only after an integration is enabled, are in the security model. Pattern redaction is not a guarantee that every secret is recognized.
Related MCP server: AI Knowledge Center MCP
What This Is
Local MCP Toolbox supports development, security review, and troubleshooting workflows that need evidence without broad machine authority.
Use it to review changes in an approved repository, investigate recurring log errors, or check an unhealthy container. Each integration requires explicit configuration. Returned evidence helps you investigate; it does not establish a root cause by itself.
The exposed tools cannot edit your files, commit code, restart containers, execute arbitrary commands, or mutate remote systems. The server writes only its own security audit records. It runs locally and connects to an MCP client; there is no separate dashboard.
What You Get
Capability | What it does | Security boundary |
System | Safe host metadata and developer-tool availability | No environment variables, usernames, process data, or executable paths |
Filesystem | Approved-root listing, metadata, and text inspection | Canonical containment, sensitive-path blocklist, extension allowlist, bounded reads |
Git | Repository status, branch, commits, and diff summaries | Explicit repository allowlist; fixed, non-interactive Git commands |
GitHub | Repository, issue, and pull-request metadata | External-network opt-in, exact repository allowlist, fixed API origin, bounded GET requests |
Python environments | Static virtual-environment dependency metadata audit | Separate roots, link-free fixed paths, no interpreter, Pip, subprocess, import, or network execution |
Metrics | Aggregate request outcomes and latency | In-process counters only; no arguments, responses, identifiers, or secrets |
Docker | Opt-in container metadata, health, and bounded logs | Official SDK only; no lifecycle, exec, mount, environment, or command access |
Logs | Tails, literal search, and deterministic error grouping | Dedicated approved roots, output limits, and central redaction |
Security | Bandit availability and normalized scan findings | Fixed scanner invocation; no user-controlled command arguments or fixes |
Infrastructure | Project-type detection and top-level configuration inventory | Separate approved roots; no recursive content inspection |
Incidents | Timestamped evidence and deterministic summaries | Read-only, bounded observations: never root-cause claims |
Audit | Sanitized JSONL accountability trail | Shape-only request summaries, retention, and size limits |
For parameters, output schemas, and every individual guardrail, see the full tool catalog.
Quick Start
Requirements
Python 3.12 or later
An MCP-capable client for connection after the server is validated
The default restricted profile is intentionally safe: it starts with no approved filesystem roots and no optional integrations.
Windows (PowerShell)
git clone https://github.com/chriswayneh/local-mcp-toolbox.git
Set-Location local-mcp-toolbox
python -m venv .venv
.\.venv\Scripts\python -m pip install -e ".[dev,docker]"
.\.venv\Scripts\local-mcp-toolbox doctor --config config\restricted.ymlmacOS / Linux
git clone https://github.com/chriswayneh/local-mcp-toolbox.git
cd local-mcp-toolbox
python3 -m venv .venv
.venv/bin/python -m pip install -e ".[dev,docker]"
.venv/bin/local-mcp-toolbox doctor --config config/restricted.ymldoctor checks configuration and prerequisites without changing them. Its JSON status is ready only when every check passed or was skipped; a warning makes it attention, and the command still exits 0. Review any reported issues, then connect your client using the supplied template. The client starts the server when needed.
For a manual startup check, run local-mcp-toolbox serve --config config/restricted.yml using the executable in your virtual environment. It waits for MCP messages and does not open a browser. Press Ctrl+C to stop it before letting your client start its own instance.
The default profile has no approved file roots or optional integrations. Start by asking your client to call toolbox_server_status. Then follow getting started to grant only the access you need. The install above includes development and Docker support; Docker itself is optional and remains disabled until configured.
Security by Design
The design applies zero-trust principles and least privilege: each tool request is checked against local policy, and integrations receive only the access you explicitly configure. Local execution alone does not authorize access to a path, repository, or optional integration.
Control | Protection |
Deny by default | The restrictive profile has no approved filesystem roots or integrations. |
Approved roots | Canonical containment blocks arbitrary filesystem access and escape paths. |
Read-only surface | No generic shell, mutation, commit, lifecycle, or remote-execution tool is registered. |
Fixed subprocesses | External binaries use fixed argument templates, |
Central redaction | PEM blocks, credentials, cookies, authorization headers, connection strings, and home-directory paths are redacted before output. Email and IP redaction are off unless enabled. |
Output bounds | File reads, collections, subprocess output, and responses are size-limited. |
Sanitized audit | Requests record safe metadata, actual outcomes, and redaction counts. Raw secrets and tool output are excluded. |
Explicit integrations | GitHub, Git, Docker, logs, scanners, infrastructure, and incident tools must be configured intentionally. |
Untrusted evidence | Retrieved files, logs, commit messages, and metadata are treated as untrusted data. |
The workstation, configured policy, installed dependencies, and any enabled Docker proxy remain trusted components. Read-only tools are not an operating-system sandbox, and the loopback HTTP token is not multi-user identity. These controls do not constitute a complete enterprise zero-trust architecture or security certification. See the verified limitations.
Read the security model, threat model, and the security-focused architecture decisions for the complete rationale.
How It Works
An MCP client requests one registered tool.
The toolbox validates typed inputs and bounded parameters.
Permissions, approved roots, and integration allowlists are checked.
A narrow read-only operation collects the permitted data.
Results are redacted and bounded before they cross the MCP boundary.
Sanitized request metadata is recorded in the audit log.
The client receives a safe structured result or error.
Connect a Client
The repository includes stdio configuration templates for the clients below. Template parsing and the installed server's launch contract are tested; individual desktop application versions are not certified. Adding a client entry lets the client start the process: it does not grant the server broader permissions.
Client | Copy-ready template |
Codex | |
Claude Desktop | |
Claude Code | |
Visual Studio Code |
Replace the intentionally unresolved paths, then configure the smallest local policy that serves the task. See client configuration for exact installation notes and the important separation between client startup and server authorization.
See It Safely
This project is designed for evidence, not a dashboard. The synthetic demo walkthrough provides a reproducible way to see the policy boundary in action without real credentials, repositories, production logs, or a host Docker socket.
It demonstrates a safe inspection sequence:
toolbox_server_status → verify the server and active profile
logs_tail_file → view redacted synthetic log evidence
logs_error_summary → group observed errors without causal claims
infra_detect_project_types → inspect demo project metadata
docker_unhealthy_containers → observe an intentionally unhealthy demo serviceThe demo’s fabricated token is redacted, disabled integrations return a structured denial, and its audit trail contains sanitized metadata only. Follow the walkthrough to run it locally.
Architecture
flowchart LR
Client["MCP client"] --> Transport["stdio transport"]
subgraph Boundary["Local policy enforcement boundary"]
Registry["MCP server / tool registry"] --> Permission{"Permission check"}
Permission -->|Denied| Error["Safe structured error"]
Permission -->|Allowed| Tool["Narrow read-only tool"]
Tool --> Guard["Redaction + output limits"]
end
Transport --> Registry
Guard --> Client
Registry -. "sanitized metadata" .-> Audit["JSONL audit log"]
Tool --> Integration["Explicitly approved local integrations"]
classDef boundary fill:#EAF3FF,stroke:#4A78A8,color:#102A43
classDef control fill:#E9F7EF,stroke:#2E7D32,color:#173E22
classDef denial fill:#FDECEC,stroke:#C62828,color:#5C1111
class Registry,Tool,Guard boundary
class Permission,Audit,Integration control
class Error denialAll retrieved content remains untrusted data. The full component model and trust-boundary discussion live in architecture.
Repository Structure
src/mcp_toolbox/ MCP server, tool modules, permissions, redaction, audit, config, and CLI
tests/ Unit, integration, and security regression tests
config/ Restricted, standard, and container policy profiles
docs/ Architecture, threat model, operating guides, ADRs, and tool reference
examples/ MCP client configuration templates
demo/ Synthetic services, logs, and intentionally insecure test fixtures
.github/ CI, security, documentation, release, Dependabot, and contribution templatesDocumentation
Document | Purpose |
System design, components, and data flow | |
Controls and trust boundaries | |
Threat analysis and mitigations | |
Findings, remediation, verification, and residual responsibilities | |
Authorization sequence and profile behavior | |
Inputs, outputs, and module-level guardrails | |
Codex, Claude, and VS Code setup | |
Optional loopback transport and bearer-token controls | |
Hardened container profiles and socket-proxy guidance | |
Synthetic end-to-end policy demonstration | |
Quality, security, docs, package, and release controls | |
Supported v1.5 scope, acceptance evidence, operations, and limitations | |
Completed v1.5 scope and optional future proposals |
Project Status
Version 1.5 adds allowlisted GitHub inspection, content-free runtime metrics, a static Python environment auditor, authenticated loopback HTTP, crash-safe audit rotation, and hardened release controls. Kubernetes inspection and generation features were removed after security review because their effective behavior could not satisfy the inspection-only contract.
The v1.5 feature scope is complete. Version 1.5.2 closes default-parameter and container setup defects without adding capabilities. Support is limited to the local inspection contract in the acceptance record, not a hosted service, multi-user security boundary, or production availability guarantee. The package classifier Production/Stable means that local contract, not a hosted production service. Versions 2 through 4 are optional proposals, not unfinished release requirements.
Contributing and Security
Contributions are welcome when they preserve the project’s least-privilege model. Start with CONTRIBUTING.md, use the repository templates for bugs and feature proposals, and report vulnerabilities through the process in SECURITY.md.
License
Licensed under the MIT License. Use it, fork it, modify it, or build something of your own. See LICENSE for the terms.
If this project helped you, a ⭐ is appreciated.
Built with
Python · Model Context Protocol · MCP Python SDK · Pydantic · Typer · Docker
Available Tools
8 toolsdisk_usageDisk usageARead-onlyIdempotent
Return disk capacity for the server's current working volume.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered without the description. The description's one added behavioral fact is the scope limit ('current working volume'), which is useful but thin; it says nothing about failure behavior if the volume is unavailable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that leads with the action and immediately qualifies the scope. No filler, no restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description needn't explain return values, and a zero-parameter tool needs little else. The only gap is the absence of any hint about when this beats system_info, which for a tool with six siblings is a real but modest omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. The description correctly implies the call is argument-free.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return') and resource ('disk capacity') and narrows the scope to the server's current working volume, which is more precise than the bare title. It does not, however, explicitly distinguish itself from siblings like system_info or toolbox_metrics_snapshot, which an agent might also reach for when asked about server health.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no exclusion of alternatives, and no note that this reports only one volume rather than all mounted filesystems. Usage is only weakly implied by the phrase 'current working volume'; an agent comparing this against system_info gets no help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
filesystem_file_metadataApproved file metadataARead-onlyIdempotent
Return safe metadata for an approved readable file without reading its contents.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered structurally. The description adds one useful behavioral detail — that it does not read file contents — but does not disclose what 'safe metadata' actually contains or how it handles non-approved/unreadable paths.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence with the key constraint (no content read) front-loaded and no wasted words. Appropriately sized for a one-parameter read tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. The safety profile is covered by annotations. The main residual gap is the undefined term 'approved', but for a low-complexity read tool the description is essentially sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single 'path' parameter and 0% schema description coverage, the description must carry the semantics. 'Approved readable file' hints that the path must reference a pre-approved file, adding marginal meaning, but it does not clarify path format or error behavior for unapproved paths.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Return) and resource (metadata for an approved readable file), and the phrase 'without reading its contents' clearly distinguishes it from the sibling filesystem_read_text_file. It doesn't name the sibling explicitly, but the scope is unambiguous enough for selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'approved readable file' implies a precondition (the file must be approved/readable), which gives some contextual guidance. However, there is no explicit when-to-use vs. when-not statement and no alternative tool named, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
filesystem_list_directoryList approved directoryARead-onlyIdempotent
List immediate entries in an approved directory with bounded pagination.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| limit | No | ||
| offset | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered structurally. The description usefully adds 'immediate entries' (non-recursive) and 'bounded pagination', which the annotations do not convey, but says nothing about error behavior for unapproved paths or ordering guarantees.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the action and resource, then the two most important scope constraints. Every word earns its place with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations carry the full safety profile. The remaining gap is parameter semantics, which is minor for a three-parameter listing tool where the names are self-evident.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description provides no meaning for any of the three parameters: it never explains that 'path' is the required target, that 'limit' defaults to 50, or that 'offset' controls pagination position. 'Bounded pagination' only vaguely gestures at limit/offset without defining their semantics or bounds.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (immediate entries in a directory) with scope qualifiers ('immediate', 'approved', 'bounded pagination') that clarify it is non-recursive and sandboxed. It is clear on its own, but does not name or differentiate itself from siblings like filesystem_file_metadata or filesystem_read_text_file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'approved directory' implies a precondition (only allowlisted paths are valid), which is useful implied guidance. However, there is no explicit when-to-use vs. when-to-use-an-alternative, and no mention of what to do if a path is rejected or how this differs from reading file metadata.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
filesystem_read_text_fileRead approved text fileBRead-onlyIdempotent
Read a bounded approved text file after path, type, size, and secret controls.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false and destructiveHint=false, so the safety profile is covered. The description adds real context beyond that by revealing an approval/allowlist gate plus type, size, and secret filtering, but it is vague about what happens on rejection or truncation and gives no error/rate-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no padding, and the gating constraint is stated immediately after the verb. It is terse to the point of being cryptic ("path, type, size, and secret controls"), which costs it the top mark.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and annotations carry the safety profile. What is still missing for a control-gated reader is what "approved" concretely means and how failures surface, leaving the agent to guess at failure modes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single path parameter has no description in the schema, so the description must compensate. It mentions "path" only as a control checkpoint and never explains whether the path is absolute or relative, how it is resolved, or whether it must fall inside an allowed root.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and a scoped resource (bounded approved text file), which is enough to separate it from filesystem_list_directory and filesystem_file_metadata. It does not explicitly name or contrast any sibling, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"after path, type, size, and secret controls" implies that reads are gated by approval/allowlist checks, which is a usable precondition hint. However, it never states when to prefer this tool over siblings or what to do if a file fails those controls, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
installed_developer_toolsInstalled developer toolsARead-onlyIdempotent
Report availability of an allowlisted set of common developer tools.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world behavior, so the safety profile is covered. The description adds the useful scope constraint that the tool reports only an 'allowlisted' set rather than all installed tools, but says nothing about which tools or how availability is determined.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no redundant clauses; every word earns its place and the scope constraint is stated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only probe with an output schema present, the description need not explain return values. It is nearly complete, with the only gap being which tools are on the allowlist and any platform/availability caveats.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies. There is nothing parameter-related for the description to clarify or compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Report') and resource ('availability of an allowlisted set of common developer tools'), which is clearer than a tautology. It is distinguishable from filesystem and toolbox siblings, though it does not explicitly contrast with the nearby system_info tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no when-to-use context, no prerequisites, and does not name any alternative or boundary against siblings like system_info. The agent must infer the use case entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
system_infoSystem informationARead-onlyIdempotent
Return sanitized OS and runtime metadata without reading environment variables.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds a genuine behavioral trait beyond that: output is 'sanitized' and environment variables are deliberately not read, which tells the agent this is a privacy-safe diagnostic call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the action and the key constraint with no filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and zero parameters means no schema gaps. The description covers purpose and the sanitization boundary; it could say marginally more about what 'runtime metadata' excludes, but nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which per the rubric sets a baseline of 4. The description correctly implies a no-argument call and needs to add no further parameter detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Return') and resource ('OS and runtime metadata'), with a distinguishing scope qualifier ('sanitized ... without reading environment variables') that separates it from data-gathering siblings like installed_developer_tools. It does not explicitly name a sibling, but the scope constraint makes the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus toolbox_server_status, toolbox_metrics_snapshot, or disk_usage, all of which also return system-level data. The 'without reading environment variables' clause hints at a privacy boundary but does not frame it as a use/don't-use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
toolbox_metrics_snapshotLocal MCP Toolbox metricsARead-onlyIdempotent
Return content-free aggregate request and latency metrics.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false and destructiveHint=false, so the safety profile is covered structurally. The description adds one genuine behavioral fact beyond them: the metrics are 'content-free' aggregates, which tells the agent no per-request payloads or sensitive content are exposed. It does not disclose the aggregation window, sampling, or refresh behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero filler, front-loading the verb and the two things measured. Nothing could be removed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value structure need not be described, and there are no parameters to document. The one remaining gap is the scope of the aggregate (time window or granularity), which an agent would need to interpret the numbers, though that is plausibly captured by the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a no-parameter tool applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Return') and a specific resource ('aggregate request and latency metrics'), and the 'content-free' qualifier separates it from payload-returning siblings like filesystem_read_text_file. It does not explicitly position itself against the closest sibling, toolbox_server_status, but the resource is distinct enough to identify the tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this tool, no prerequisites, and no mention of alternatives such as toolbox_server_status. Usage is only inferable from the tool name and the word 'metrics'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
toolbox_server_statusLocal MCP Toolbox server statusBRead-onlyIdempotent
Return server-generated, redaction-safe status metadata.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds one genuine behavioral detail beyond that — the output is "redaction-safe" — but says nothing about freshness, caching, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no padding and the key qualifiers (server-generated, redaction-safe) front-loaded. It is efficient, though so terse that brevity shades into under-specification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and there are no parameters to document. The remaining gap is routing: the description never says how this status differs from the several sibling status/info tools, which is the main thing an agent needs here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is nothing for the description to disambiguate. Schema coverage is 100% and the empty argument object is unambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It gives a verb ("Return") and a resource ("status metadata"), but the resource is abstract and largely restates the tool name and title. Nothing distinguishes it from sibling tools such as system_info or toolbox_metrics_snapshot, so an agent cannot tell which status surface this covers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no condition selecting this over system_info, disk_usage, or toolbox_metrics_snapshot, and no exclusions. The agent must guess based on the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v1.5.2- First observed
disk_usage - First observed
filesystem_file_metadata - First observed
filesystem_list_directory - First observed
filesystem_read_text_file - First observed
installed_developer_tools - First observed
system_info - First observed
toolbox_metrics_snapshot - First observed
toolbox_server_status
TDQS
Scored across 8 tools
Most tools have clearly distinct targets: OS metadata, disk capacity, dev-tool availability, and three filesystem operations (list/metadata/read) are well separated. The only mild overlap is among system_info, toolbox_server_status, and toolbox_metrics_snapshot, which all return 'metadata/status' of a system or server, though the descriptions do differentiate them.
All names use snake_case, which is consistent. However, the set mixes three prefixing conventions: unprefixed tools (system_info, disk_usage, installed_developer_tools), a filesystem_ group, and a toolbox_ group, which makes the namespace slightly less predictable than a single uniform scheme.
Eight tools is well within the ideal 3-15 range and each one maps to a distinct read/introspection capability. Nothing feels redundant or padded, and no obvious operation is missing due to under-provisioning.
For a deliberately read-only, safety-bounded local toolbox, the surface covers system info, disk, dev tools, and the core filesystem read lifecycle (list, metadata, read). Write/search operations are absent, but the descriptions imply a read-only design intent, so this is a minor rather than significant gap.
Maintenance
Related MCP Connectors
The Cortex MCP server provides read-only access to real-time engineering context from the Cortex developer portal, allowing AI coding assistants to answer natural language questions about your organization's catalog (microservices, libraries, domains, teams, infrastructure), scorecards (engineering standards and best practices), initiatives (goals and deadlines), and Engineering Intelligence metrics. It includes tools for querying documentation, tracking personal entities, and accessing AI-assisted insights across the entire Cortex ecosystem.
Self-hosted MCP server: 26 deterministic dev, security, and EVM tools.
Guarded MCP server for agent-readable business truth, provenance, readiness, and discovery.
Security scanner for MCP servers. Detect vulnerabilities, prompt injection, and tool poisoning.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceSecure local development platform that exposes controlled developer capabilities (FS, Git, search, command execution) to AI assistants via MCP with deny-by-default security and audit logging.-
- AlicenseNot gradedqualityAmaintenanceLocal-first MCP server that provides project context, verification gates, and structured tools for coding agents to discover knowledge, run diagnostics, and execute allowlisted commands within a repository.16 npmMIT
- AlicenseNot gradedqualityCmaintenanceRead-only MCP server to inspect allowlisted Docker containers, systemd services, JSONL logs, and HTTP health endpoints without arbitrary shell access.MIT
- AlicenseNot gradedqualityCmaintenanceA local-first MCP server providing guarded access to a workspace with file operations, search, commands, tests, Git helpers, checkpoints, and structured tool results. It supports multiple tool modes and emphasizes security with workspace restrictions and secret blocking.1MIT