majrooo-mcp-devkit
This server provides safe command execution and code refactoring tools for MCP clients, with multi-project root management and a feedback system.
Run shell commands safely in project roots (default tool
run_safe_command, with destructive ops requiring explicit confirmation viarun_destructive_command, and filtered output viarun_command_grep).Read command logs saved to temp files (via
read_log_slice) instead of re-running commands.Discover and validate project roots — list allowed roots and friendly names (
list_allowed_roots) or resolve an exactcwdfor a given path (resolve_cwd).Find symbol references across workspace(s) with optional language-aware role detection (
universal_find_references).Extract code blocks (functions, structs, classes, methods) from files with precise line ranges and annotations (
extract_code_block).Split files into multiple module files based on declarations, with optional index generation and dry-run (
split_file_by_declarations).Apply batch edits to multiple files with validation-first, partial rollback on failure, and dry-run preview (
batch_apply_edits).Generate module skeletons from existing source symbols (
generate_module_skeleton).Verify refactoring safety via semantic diff (
verify_refactor_safety).Manage feedback — report bugs/improvements/feature requests, list them, and close/archive them (
report_tool_feedback,list_feedback,close_feedback).Discover tools and get help — list available tools by category (
list_tools) and get detailed parameter info (help_tool).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@majrooo-mcp-devkitRun the test suite in my project and show me any failures."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
majrooo-mcp-devkit — MCP DevKit: Safe Commands + Refactoring Tools
Repository Access: PUBLIC
Version: 0.2.1 · Tests: 354 passing · License: GPL-3.0-or-later
MCP server that provides safe command execution and code refactoring tools for Cline/Claude Desktop.
Tools
run_safe_command
Execute a shell command restricted to the active project root. Dangerous commands and writes outside the active root are automatically blocked. This is the default tool — always use this first.
Parameter | Type | Default | Description |
| string | — | Command to execute |
| string | primary root | Working directory (must be inside |
| number | 200 | Max output lines before truncation |
| number | 60000 | Command timeout (1000–600000 ms) — raise for long jest/build runs |
run_destructive_command
Execute a potentially dangerous command with explicit user confirmation. Only use when run_safe_command blocked the command and the user explicitly agreed after being informed of the specific risk.
Parameter | Type | Default | Description |
| string | — | Command to execute |
| boolean | false | Acknowledge the risk (required for dangerous commands) |
| string | primary root | Working directory (must be inside |
| number | 200 | Max output lines before truncation |
| number | 60000 | Command timeout (1000–600000 ms) |
read_log_slice
Read a portion of a previously saved log file. Use this instead of re-running a command with higher maxLines. Files are read directly via Node.js (not through the shell), so it also works for truncated logs in os.tmpdir().
Parameter | Type | Default | Description |
| string | — | Path to the log file |
| number | 0 | Starting line (0-based) |
| number | 100 | Number of lines to read |
list_allowed_roots
Return the registered roots configuration: the primary project (MCP_PROJECT_ROOT), all allowed roots (MCP_EXTRA_ROOTS, including globs), the concrete existing project directories under them (usable as cwd), and whether MCP_BLOCK_CROSS_ROOT_READS is enabled. Projects with a friendly name are returned as { path, name } — in that case you can also use the name as cwd. Call this before working in any non-primary project to discover the exact cwd value to use. Runs no commands — it only reads configuration and lists directories.
No parameters.
resolve_cwd
Verify whether a path (or a friendly project name from MCP_PROJECT_NAMES) is inside the allowed roots and get the exact cwd to use for commands. Pass the path you want to work in (e.g. your workspace folder) instead of guessing.
On success returns { ok: true, cwd, matchedRoot, exists, name? }; on failure { ok: false, error, roots }. exists tells whether the resolved directory actually exists on disk (relative cwd values are resolved against the primary project). Runs no commands — it only validates configuration.
Parameter | Type | Description |
| string | Path to verify (absolute, e.g. the project workspace folder) |
run_command_grep
Execute a command and return only lines matching a pattern (case-insensitive regex). Use instead of run_safe_command when you only care about specific lines (e.g., errors in build output). This is the replacement for Unix grep on Windows — filtering happens in-process, so grep/head/tail are not needed.
Parameter | Type | Description |
| string | Command to execute |
| string | Regex pattern to filter lines (case-insensitive) |
| string | Working directory (default: primary root) |
| number | 60000 |
universal_find_references
Find all occurrences of a symbol across a workspace. Returns structured output with file, line, column, context, and optional role annotations. Use this before any refactoring session to understand what will break when a symbol is renamed or moved.
Parameter | Type | Default | Description |
| string | — | Symbol to search for (word-boundary match) |
| string | all registered roots | Workspace root to search. When omitted, all registered roots are searched (nested roots pruned, duplicate files removed) |
| string[] | common source extensions | Restrict to these extensions |
| string[] |
| Directories to skip |
| number | 1 | Lines of context around each match |
| string | — (disabled) | Optional: |
Multi-root search (no cwd): every registered root is searched. A root nested inside another root is pruned — the outer root already covers it — and files are deduplicated by resolved real path, so the same match is never listed or counted twice. When more than one root is searched, the report states them and, for each file group, the root its relative path is based on:
Symbol: ColAlign
Total matches: 28
Searched roots (3):
- D:\W\TS
- d:\Users Data\jox\My Documents\Python
- d:\Users Data\jox\My Documents\Rust
Duplicates skipped: 2 (same file reachable through a nested root)
pixel-blaster-engine/engine_bevy/src/lib.rs (relative to d:\Users Data\jox\My Documents\Rust)
Line 72:12 [usage] — pub use ui_panel::{A path like project/src/lib.rs therefore always belongs to exactly one root — it is never a second copy of src/lib.rs. With an explicit cwd only that root is searched and the report keeps its compact form (path/file.rs:).
extract_code_block
Read the full text of a function, struct, class, or method from a file. Returns precise line range + content. Includes leading annotations (#[derive], @decorator, /// doc comments). String/comment-aware bracket matching prevents false depth counts from braces inside strings or comments.
Parameter | Type | Default | Description |
| string | — | Source file path (must resolve inside allowed root) |
| string | — | Symbol name to extract |
| number | 0 | Extra lines before/after the block |
| string | primary root | Working dir for resolving relative file paths |
split_file_by_declarations
Split a large file into multiple smaller files based on top-level declarations. Optionally generates a combining file (mod.rs / index.ts / __init__.py). Use dryRun: true (default) to preview the layout before writing.
Parameter | Type | Default | Description |
| string | — | Source file to split |
| object[] | — |
|
| string | dirname(file) | Where new files are written |
| string | auto-detect |
|
| boolean | true | Create combining file |
| boolean | true | Preview only — write nothing |
| boolean | false | Allow overwriting existing targets |
| string | primary root | Working dir for resolving relative file paths |
batch_apply_edits
Apply multiple file edits with partial rollback on failure. Validates all edits first — if any search string is not found or matches multiple times (without replaceAll), NO files are modified. On failure during application, only files modified by the failed edit and subsequent edits are reverted; earlier successful edits are preserved. Same-file chain failures revert the entire file. Relative paths without cwd are resolved against the primary root; if the file doesn't exist there, all other registered roots are searched automatically (unique match → use it; multiple matches → error with instructions to specify cwd).
Parameter | Type | Default | Description |
| object[] | — |
|
| boolean | true | Preview all changes without writing |
| string | primary root | Working dir for resolving relative file paths |
Response (explicit outcome)
The result text starts with a one-line message followed by the JSON payload:
{
"dryRun": false,
"totalEdits": 3,
"validated": 3,
"appliedEdits": 2,
"written": ["D:\\proj\\a.rs", "D:\\proj\\b.rs"],
"message": "Applied 2 of 2 edit(s) to 2 file(s) — all changes are on disk.",
"preview": [{ "file": "D:\\proj\\a.rs", "action": "edit", "matchCount": 1, "applied": true }]
}Key fields:
Field | Meaning |
| Human-readable summary — always states explicitly whether anything was written |
| Number of edits whose changes are on disk after the run ( |
| Files changed on disk after the run (empty in dry-run / when nothing was written) |
|
|
| Occurrences found — informational, not proof that anything was written (check |
On failure the response also carries error, failedAt, reason (validation_failed | write_failed), reverted (files rolled back) and nothingWritten (true = this run left no change on disk at all — either nothing was ever written, or everything written was rolled back). Validation failure example:
VALIDATION FAILED on edit #2 of 3 — NO edits were written to disk (validation-first: nothing is written until every edit validates). Reason: search string not found in D:\proj\b.rsEdits that were never evaluated get an explicit action: "error" preview entry (not evaluated — batch stopped at edit #N), so preview always maps 1:1 to the edits array. When validation fails on a file that already had a successfully applied edit in the same run (chained edits), the response switches to partial rollback: N edit(s) kept in … ; M file(s) reverted […]. If every write from the run was rolled back (e.g. all edits chained on one file), nothingWritten stays true and the message says NO net changes were left on disk: N edit(s) had been written and M file(s) were rolled back […] — an honest distinction between "never wrote" and "wrote, then undid".
generate_module_skeleton
Generate a new module file with extracted symbols from a source file. Returns error with unknownSymbols list if any symbols are not found.
Parameter | Type | Default | Description |
| string | — | Target file path |
| string[] | — | Symbol names to include |
| string | — | Original file to extract from |
| string | auto-detect |
|
| boolean | true | Preview only |
| boolean | false | Allow overwriting existing file |
| string | primary root | Working dir for resolving relative file paths |
verify_refactor_safety
Semantic diff between old and new code. Catches accidental deletions before compilation. Checks: function count, signatures, export count, imports, comment ratio. Intentionally conservative — renames appear as errors requiring explicit confirmation.
Parameter | Type | Default | Description |
| string | — | Original code text |
| string | — | New code text |
| string | auto-detect |
|
report_tool_feedback
Report a bug, improvement, or feature request about a tool of this server (see list_tools). Writes structured feedback to .mcp/FEEDBACK.md (project-specific, gitignored). Entries are idempotent — duplicate reports are skipped.
The tool name is validated against this server's tool registry: an unknown name is rejected (nothing is written) with a "did you mean …?" suggestion, so feedback about another client's built-in tool no longer lands in this log. For a missing-capability report about the server as a whole, pass allowUnknownTool: true.
Parameter | Type | Default | Description |
| string | — |
|
| string | — | Name of the MCP tool this feedback is about — must be a tool of this server (e.g. |
| string | — | Short summary (1 line) |
| string | — | Detailed description |
| string | — | Steps to reproduce (optional) |
| string | — | What you expected (optional) |
| string | — | Suggested fix or improvement (optional) |
| boolean | false | Accept a name that is not a tool of this server (missing-capability reports only) |
list_feedback
List feedback entries from .mcp/FEEDBACK.md. Optionally filter by type, tool name, or status. Use this to check existing feedback before creating new entries.
Parameter | Type | Default | Description |
| string | — | Filter: |
| string | — | Filter by tool name |
| string | — | Filter: |
| boolean | false |
|
close_feedback
Close an existing feedback entry by ID — sets status to "closed" and optionally adds resolution text. Use this to mark feedback items as resolved after fixing them.
Closing also moves every closed entry out of the active log into .mcp/FEEDBACK_ARCHIVE.md, so FEEDBACK.md keeps holding only open items and stays small. The move is append-first (archived before removed) and idempotent, and it self-heals: entries closed before this behaviour existed are migrated along with the next close. Read them back with list_feedback + archived: true.
Parameter | Type | Default | Description |
| string | — | The feedback entry ID to close (from |
| string | — | Resolution note explaining how the issue was addressed (optional) |
Closed entries are stored in two files:
File | Contents |
| Header + open entries (the active log) |
| Header + closed entries, moved verbatim (description, reproduction, resolution, |
list_tools
List all available MCP tools with descriptions. Use this to discover tools before starting a task. Filterable by category.
Parameter | Type | Default | Description |
| string | — | Filter: |
help_tool
Get detailed help for a specific MCP tool — parameters, types, defaults, and description.
Parameter | Type | Default | Description |
| string | — | Tool name to get help for |
Related MCP server: MCP Workspace Server
Supported Languages
The refactoring tools support four languages for declaration parsing, symbol extraction, and role detection:
Language | Status | Role detection | Split | Skeleton |
Rust | ✅ Fully tested | ✅ | ✅ | ✅ |
TypeScript | Experimental | ✅ | ✅ | ✅ |
Python | Experimental | ✅ | ✅ | ✅ |
C++ | Experimental | ✅ | ✅ | ❌ |
Tools without a language parameter (batch_apply_edits, extract_code_block) are language-agnostic — they operate on plain text and work with any language.
Note: Rust has been tested in production refactoring scenarios. Other languages are structurally supported but have not been tested against real-world codebases yet.
Configuration
The server supports one instance, many projects. Projects are selected per command via the cwd parameter; the active project also acts as the "lockbox" for write/read checks.
Env var | Description |
| Primary project root (default |
| Additional roots, semicolon separated. |
| Friendly names for projects, semicolon separated |
|
|
Entry forms supported in both variables:
Plain path
D:\W\TS\majrooo-mcp-devkit→ prefix: the directory itself and everything below it are allowed. RegisteringD:\Wcovers all projects under it.Glob
D:\W\TS\*(*,**,?) → any path matching the pattern (and its subtree) is allowed.
Example — one instance, many projects, with friendly names:
{
"mcpServers": {
"majrooo-mcp-devkit": {
"command": "node",
"args": ["D:\\W\\TS\\majrooo-mcp-devkit\\build\\index.js"],
"env": {
"MCP_PROJECT_ROOT": "D:\\W\\TS\\majrooo-mcp-devkit",
"MCP_EXTRA_ROOTS": "D:\\W;D:\\python",
"MCP_PROJECT_NAMES": "D:\\W\\TS\\cb=ZbaľSa;D:\\W\\TS\\nase-zasoby=Naše zásoby"
}
}
}
}MCP_PROJECT_NAMES maps a real project path to a readable name. This is useful when the folder name had to be shortened (e.g. Gradle path-length limits) or the project was renamed. The name can be used directly as cwd (e.g. "cwd": "ZbaľSa"), and list_allowed_roots will show such projects as { "path": "D:\\W\\TS\\cb", "name": "ZbaľSa" }.
Switching projects is done via the cwd parameter, never via cd in the command. cd .., cd ~, cd C:\..., and Windows cd /d D:\... are always rejected.
Running tests / long commands (the anti-freeze workflow)
Never run test suites (jest/npm test), typecheck or builds through the Cline built-in terminal — it has no timeout and can freeze the whole window. Use the MCP tools instead:
Always pass the project's
cwd(or friendly name, e.g."cwd": "ZbaľSa").To filter output (e.g. jest summary), use
run_command_grep— filtering happens in-process, socmd /c "... | findstr ... & echo DONE"is not needed and discouraged:
{
"tool": "run_command_grep",
"cwd": "ZbaľSa",
"command": "npx jest src/app/__tests__/catalog.test.tsx 2>&1",
"pattern": "Tests:|Test Suites:|FAIL|PASS|✕",
"timeoutMs": 180000
}For full output use
run_safe_commandwith a smallmaxLines— the full output is saved to a temp log forread_log_slice:
{
"tool": "run_safe_command",
"cwd": "ZbaľSa",
"command": "npm run typecheck 2>&1",
"maxLines": 100,
"timeoutMs": 180000
}If a run exceeds 10 minutes, run it in the background, redirect to a log file, and poll the log via
run_command_grep— do not watch live terminal output.
Typical workflow
Call
list_allowed_rootsto see the primary root, the allowed roots (including globs), and the concrete projects under them.If you need to confirm a specific path, call
resolve_cwdwith your workspace folder — it returns the exactcwdto use and the matched root.If the task targets a project other than the primary one, pass the resolved path as
cwdon every command (run_safe_command,run_destructive_command,run_command_grep).Otherwise, omit
cwd— commands run in the primary root.
Relative paths: File-based refactoring tools (
batch_apply_edits,extract_code_block,split_file_by_declarations,generate_module_skeleton) resolve relative paths against the primary root whencwdis omitted. If the file doesn't exist in the primary root (or in thecwdroot whencwdis provided), all other registered roots are searched automatically — a unique match is used directly, while multiple matches produce an error with instructions to specifycwd.
Safety Mechanisms
Layer | Description |
Registered roots |
|
Directory restriction | Commands execute with |
Dangerous pattern detection | Regex blacklist blocks destructive commands ( |
Write-target check | Best-effort detection of writes outside the active root ( |
Cross-root read check (opt-in) |
|
Explicit confirmation |
|
Missing destructive target | A confirmed destructive command ( |
Buffer & timeout limits | 50 MB max output, 60-second timeout. |
Output truncation | Long outputs are saved to |
Audit log | All executions logged to |
Limitations
The dangerous-pattern blacklist and the write/read target checks are best-effort layers, not security guarantees. Shell features (variables, command substitution, encoding) can bypass them. For production isolation use Docker/VM sandboxing.
Windows Notes
Commands run through
cmd.exe. Unix-only tools (grep,head,tail, ...) do not exist — the server returns a friendly error with alternatives instead of a raw "not recognized" blob.Use
run_command_grepinstead ofgrep, andread_log_sliceor PowerShell (Get-Content out.log -TotalCount 30) instead ofhead.Long-running processes (dev server, watch mode) exceed the 60s timeout — use the built-in terminal for those.
Output normalization
On Windows the server automatically prefixes commands with chcp 65001 > NUL && so the child process emits UTF-8 instead of the legacy OEM codepage (which would otherwise decode into U+FFFD replacement characters, e.g. around thousands separators in dir output). ANSI color codes from tools like vitest/jest are stripped as a fallback (NO_COLOR=1 / FORCE_COLOR=0 are also injected into the environment), and read_log_slice cleans ANSI codes defensively when reading older logs.
Redirected output reporting
When a command redirects its output into a file (npm test > test.log 2>&1, >> out.log, 2> err.log, ...), the response reports where the output went and shows the tail of the written file instead of an empty response or a bare Command failed: .... On failure the response also includes the exit code (or timeout/signal) and the captured stdout/stderr. Discard targets (NUL, /dev/null), wildcard patterns and fd-duplication tokens (2>&1, >&-) are skipped; very large files are read only from the end.
Agent Behavior Rules
The .clinerules file in the project root defines how AI agents should use these tools:
Always try
run_safe_commandfirst — never start withrun_destructive_command.run_destructive_commandonly after explicit user confirmation in the current conversation — general consent ("do what you need") is not sufficient.Never bypass
directory_escaperejections — no chaining, absolute paths, or cwd tricks. Use thecwdparameter to pick a registered project.Prefer
run_command_grep/read_log_sliceover increasingmaxLinesfor long output, and instead of Unixgrep/head/tail.Use the built-in terminal only for quick interactive checks (e.g.,
git status). Large-output commands (npm install, build, tests) must go throughrun_safe_command.Run tests before reporting task as complete.
License
This project is licensed under the GNU General Public License v3.0 or later - see the LICENSE file for details.
Development
npm run build # Compile TypeScript
npm test # Run unit tests
npm run test:watch # Watch modeThe server communicates over STDIO using the Model Context Protocol.
Available Tools
17 toolsbatch_apply_editsADestructive
Apply multiple file edits with partial rollback on failure. Validates all edits first — if validation fails, response says explicitly how many edits were applied and which files were written/reverted (NO files are modified). On failure during application, only the failed edit and later edits are reverted; earlier successful edits are preserved. Every response includes message, appliedEdits, written, and per-preview-entry applied. Use dryRun: true (default) to preview changes.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Working dir for resolving relative file paths (default: primary project root) | |
| edits | Yes | List of edits to apply | |
| dryRun | No | Preview all changes without writing (default: true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the annotations. Annotations indicate destructiveHint=true and readOnlyHint=false, but the description details the exact rollback semantics: validation failure means no files modified, and on application failure only the failed and later edits are reverted, preserving earlier ones. It also specifies the response structure (message, appliedEdits, written, per-preview applied). This is comprehensive behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise yet information-dense. Each sentence adds critical information: the core behavior, validation guarantee, rollback specifics, response fields, and dryRun default. It is front-loaded with the main purpose and uses no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with this complexity (multiple edits, validation, partial rollback, dry run), the description covers all essential aspects an agent needs: validation outcomes, failure behavior, success preservation, response format, and dryRun default. No output schema exists, so describing the response fields is necessary and well done. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters (cwd, edits, dryRun) with detailed descriptions. The tool description does not add extra meaning beyond what the schema provides; it mentions dryRun but without new details. Per the rubric, baseline 3 is appropriate when schema carries the parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Apply multiple file edits with partial rollback on failure.' It specifies the verb (apply), resource (multiple file edits), and key behavioral nuance (partial rollback). This is specific and distinguishes it from sibling tools like run_destructive_command, which focus on command execution rather than file editing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context by explaining when to use dryRun: 'Use dryRun: true (default) to preview changes.' It also describes the validation-first behavior, implying when to rely on the tool for safe multi-edit operations. However, it does not explicitly contrast with alternative tools or state when not to use it, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_feedbackADestructive
Close an existing feedback entry by ID — sets status to "closed" and optionally adds resolution text. Closed entries are automatically moved out of .mcp/FEEDBACK.md into .mcp/FEEDBACK_ARCHIVE.md (read them with list_feedback archived:true). Use this to mark feedback items as resolved after fixing them.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The feedback entry ID to close (from list_feedback output) | |
| resolution | No | Resolution note explaining how the issue was addressed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include destructiveHint: true, and the description goes beyond it by explaining exactly what happens: status is set to 'closed', resolution text is optionally added, and entries are moved from FEEDBACK.md to FEEDBACK_ARCHIVE.md. This adds valuable behavioral context (file movement) that annotations do not provide, and it does not contradict the destructive hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no fluff. The core action is front-loaded, the side effect is explained in the second sentence, and the usage instruction is in the third. Every sentence earns its place, and the structure is logical and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with a clear schema and no output schema, the description covers all essential context: what it does, the automatic file movement, how to read archived entries, and when to use it. Nothing that an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both id and resolution are already documented in the input schema. The description adds only a minor repetition of 'optionally adds resolution text,' which is already in the schema. No additional meaning is provided beyond what the schema offers, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Close'), the resource ('feedback entry'), and the key behavior (sets status to 'closed', optional resolution). It also distinguishes itself from related tools like list_feedback by explaining the archive movement, so an agent can immediately understand its role without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use this to mark feedback items as resolved after fixing them,' which provides a clear when-to-use directive. It also mentions the reading alternative (list_feedback archived:true), fulfilling the requirement for mentioning alternatives. There is no misleading guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_code_blockARead-onlyIdempotent
Read the full text of a function, struct, class, or method from a file. Returns precise line range + content. Includes leading annotations (#[derive], @decorator, /// doc comments). String/comment-aware bracket matching prevents false depth counts from braces inside strings or comments.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Working directory for resolving relative file paths (default: primary project root) | |
| file | Yes | Source file path (absolute or relative to cwd) | |
| symbol | Yes | Symbol name to extract | |
| contextLines | No | Extra lines before/after the block (default: 0) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), yet the description adds real behavioral detail beyond them: it returns a precise line range plus content, it includes leading annotations and doc comments, and it uses string/comment-aware bracket matching to avoid false depth counts from braces inside literals. That is substantive context the annotations do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler: the core action is front-loaded, followed by the return shape and then the extraction-safety note. Every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly steps in to describe the return value (line range + content) and the extraction semantics (annotation inclusion, bracket matching). It stops short of covering failure behavior, such as what happens when the symbol is not found or is ambiguous, which is the main remaining gap for a tool with a required symbol lookup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (cwd, file, symbol, contextLines) are already documented in the schema and the baseline is 3. The description adds no parameter-level detail, though 'Includes leading annotations' implicitly relates to how extraction boundaries are computed rather than to any parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: 'Read the full text of a function, struct, class, or method from a file,' which is far more precise than a generic 'read file' tool. It does not explicitly name or contrast with siblings like universal_find_references or split_file_by_declarations, so the agent must infer the boundary, but the symbol-level extraction purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the phrasing (you call it to obtain the full body of a named symbol), but there is no explicit when-to-use/when-not or reference to alternatives such as find_references or split_file_by_declarations. An agent can infer the context but gets no routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_module_skeletonADestructive
Generate a new module file with correct imports, declarations and visibility. Reads the source file, extracts the specified symbols, and writes them to the target module path. Returns error with unknownSymbols list if any symbols are not found.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Working dir for resolving relative file paths (default: primary project root) | |
| dryRun | No | Preview only (default: true) | |
| symbols | Yes | Symbol names to include | |
| language | No | Language (auto-detected) | |
| overwrite | No | Allow overwriting existing file (default: false) | |
| modulePath | Yes | Target file path (e.g. src/ai/data.rs) | |
| sourceFile | Yes | Original file to extract symbols from |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false and idempotentHint=false, so the safety profile is covered. The description adds real value beyond that: it discloses the read-extract-write flow and the specific failure mode (returns an error with an unknownSymbols list when symbols aren't found), which an agent needs to handle partial failures.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and followed by mechanism and error behavior; no filler. Slightly more could be trimmed, but each sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool with full param coverage, no output schema, and safety annotations already present, the description covers the action, the data flow, and the error case. It stops short of clarifying dryRun default behavior or overwrite interaction, which would round it out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all seven parameters (cwd, dryRun, symbols, language, overwrite, modulePath, sourceFile) are already documented in the schema. The description only loosely maps to 'symbols' via 'extracts the specified symbols' and adds no format or default semantics, so the schema does the heavy lifting — baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Generate a new module file") plus the exact mechanism (reads source, extracts symbols, writes to target path). It is clearly distinguishable in intent from siblings like split_file_by_declarations, but it never explicitly names or contrasts those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what happens mechanically but gives no when-to-use guidance: no indication of when this is preferable to extract_code_block or split_file_by_declarations, and no mention of prerequisites such as the target needing to exist or symbols needing to resolve. Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
help_toolARead-onlyIdempotent
Get detailed help for a specific MCP tool — parameters, types, defaults, description.
| Name | Required | Description | Default |
|---|---|---|---|
| tool | Yes | Tool name to get help for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful value by naming what is returned (parameters, types, defaults, description), which partially compensates for the absence of an output schema, but it says nothing about error behavior for unknown tool names.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that conveys purpose and return contents with no filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only lookup tool, the definition is essentially complete, and it describes the return contents in the absence of an output schema. Minor gap: no mention of behavior when the requested tool does not exist.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'tool' parameter is already documented as 'Tool name to get help for'. The description adds no syntax, format, or validity constraints beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get') and resource ('detailed help for a specific MCP tool'), making the purpose unambiguous. It implicitly distinguishes itself from list_tools via the word 'specific', but does not name that sibling explicitly, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: consult this when you need details about a particular tool. There is no explicit when-to-use vs. list_tools guidance or statement of prerequisites (e.g., that the tool name must be known/valid). Adequate but with a clear gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_allowed_rootsARead-onlyIdempotent
Returns the allowed-roots configuration of this MCP server: the primary project (MCP_PROJECT_ROOT), all registered roots (MCP_EXTRA_ROOTS, including globs), the existing projects under them (the list of directories you can use as "cwd") and whether MCP_BLOCK_CROSS_ROOT_READS is enabled. Use THIS tool whenever you need to find out whether — and with which "cwd" parameter — you can run a command in another project. It runs no commands — it only reads the configuration and lists directories. A project with a friendly name is shown as { path, name } and you can pass its "name" as "cwd".
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so safety is covered. The description adds genuinely useful context beyond that: it explicitly states it executes nothing ('only reads the configuration and lists directories') and documents the return shape, including that a friendly-named project appears as { path, name } and that the name is passable as cwd.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The payload is dense but front-loaded: what it returns comes first, usage second, behavioral caveat third, return-shape note last. Every clause carries information, though the sentence count and parenthetical environment-variable references make it denser than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no parameters, so the description carries the whole disclosure burden — and it does, covering return contents, the { path, name } shape, the glob handling of extra roots, the cross-root flag, and the non-executing guarantee. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description still usefully explains how a value surfaced by this tool (a project 'name') is later consumed as a 'cwd' argument, adding cross-tool semantic value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It states a specific verb and resource — returns the server's allowed-roots configuration — and enumerates exactly what that comprises (primary project, registered roots with globs, existing projects, cross-root flag). An agent immediately knows what it gets back. It does not explicitly differentiate itself from the similarly-scoped sibling `resolve_cwd`, which keeps it just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is stated as a direct trigger: 'Use THIS tool whenever you need to find out whether — and with which cwd parameter — you can run a command in another project.' That is a clear when-to-use condition with a behavioral contrast ('runs no commands'). No alternative tool is named for the same need, so no exclusion clause exists.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_feedbackARead-onlyIdempotent
List feedback entries from .mcp/FEEDBACK.md (open entries; closed ones are moved to .mcp/FEEDBACK_ARCHIVE.md). Optionally filter by type, tool name, or status — pass archived:true to read the archive of closed entries instead. Use this to check existing feedback before creating new entries, or to review reported issues.
| Name | Required | Description | Default |
|---|---|---|---|
| tool | No | Filter by tool name | |
| type | No | Filter by feedback type | |
| status | No | Filter by status | |
| archived | No | true = list archived (closed) entries from .mcp/FEEDBACK_ARCHIVE.md instead of the active log |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the bar is lower. The description adds useful behavioral context beyond annotations: open entries live in FEEDBACK.md, closed entries are moved to FEEDBACK_ARCHIVE.md, and archived:true switches to reading the archive. This is meaningful behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The core action and file location are front-loaded, followed by filtering semantics and a usage hint. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with zero required parameters and fully documented schema, the description plus annotations cover everything an agent needs: what is listed, where it is read, how filtering works, and when to call it. No critical information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with descriptions for all four parameters and enums for type and status, so the baseline is 3. The description mostly restates the filtering options and archived behavior already present in the schema, adding minimal new semantic value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the specific action and resource: 'List feedback entries from .mcp/FEEDBACK.md'. It also clarifies the open-versus-archived split, which distinguishes this from the feedback creation/closure siblings. The intended purpose is immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use the tool: 'Use this to check existing feedback before creating new entries, or to review reported issues.' It implies the alternative workflow (creating entries via report_tool_feedback) but does not explicitly name alternatives or state when not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_toolsARead-onlyIdempotent
List all available MCP tools with descriptions. Use this to discover available tools before starting a task.
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | Filter by category |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds only that results include descriptions; it says nothing about result size, ordering, or how the category filter affects output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the action stated first and the usage cue second. Nothing could be cut without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, annotation-covered read-only listing tool with no output schema, the description is nearly sufficient. It notes that tools come with descriptions, but omits any hint that results can be scoped by category, which is the only unsurfaced behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single enum parameter is documented in the schema, so the baseline of 3 applies. The description never mentions the category filter or what the enum values mean for results, adding no meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (available MCP tools) with an added scoping detail (with descriptions), so an agent can tell it apart from most siblings. It does not explicitly distinguish itself from the adjacent help_tool, which is the one plausible point of confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use this to discover available tools before starting a task" gives a clear triggering context, i.e. an initial discovery step. There are no exclusions or explicit comparisons to alternatives such as help_tool, so it stops short of the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_log_sliceARead-onlyIdempotent
Reads a slice of a log file by the given line range. Use this tool instead of re-running the same command with a higher maxLines when you already have the log path saved from a previous run_safe_command / run_destructive_command response. NOTE: the file is read directly via Node.js (not through the shell), so it also works for logs in os.tmpdir().
| Name | Required | Description | Default |
|---|---|---|---|
| logPath | Yes | Path to the log file | |
| lineCount | No | Number of lines to read (default: 100) | |
| startLine | No | Starting line (0-based, default: 0) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent and non-destructive, so the safety profile is covered. The description adds genuinely useful behavior beyond that: the file is read directly via Node.js rather than through the shell, which implies it can reach logs in os.tmpdir() that shell-based reads might not. It stops short of describing error behavior for a missing path or whether lineCount is clamped.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the core action is front-loaded ahead of the routing advice and the shell-bypass note. Every clause carries actionable information for selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool whose annotations already carry the safety profile and which has no output schema to explain, the description covers purpose, routing and the shell-bypass nuance well. Minor gaps remain around failure modes and any size limits for large slices.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so logPath, lineCount and startLine are already documented with defaults and 0-based indexing. The description only echoes the 'line range' concept and adds no format or edge-case guidance beyond the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Reads a slice of a log file') plus the scoping mechanism ('by the given line range'). It is readily distinguishable from the command-execution siblings because it names them only as the source of the log path, not as the same operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when to prefer this tool ('instead of re-running the same command with a higher maxLines') and under what precondition ('when you already have the log path saved from a previous run_safe_command / run_destructive_command response'). The named alternative and the selecting condition leave nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
report_tool_feedbackAIdempotent
Report a bug, improvement, or feature request about a tool of THIS server (see list_tools). Writes structured feedback to .mcp/FEEDBACK.md (project-specific, gitignored). Unknown tool names are rejected with a suggestion — set allowUnknownTool:true only when reporting a missing capability of this server. Use this when a tool produces unexpected results, crashes, or when you need a new capability. Entries are idempotent — duplicate reports are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| tool | Yes | Name of the MCP tool this feedback is about (must be a tool of this server, e.g. 'batch_apply_edits') | |
| type | Yes | Type of feedback | |
| title | Yes | Short summary (1 line) | |
| expected | No | What you expected to happen | |
| suggestion | No | Your suggestion for a fix or improvement | |
| description | Yes | Detailed description of the issue or request | |
| reproduction | No | Steps to reproduce the issue | |
| allowUnknownTool | No | File feedback about a name that is not a tool of this server (default false) — use only for missing-capability reports |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses the write destination and scope (.mcp/FEEDBACK.md, project-specific, gitignored), the validation behavior (unknown tool names rejected with a suggestion), the allowUnknownTool semantics, and idempotency (duplicate reports skipped). This adds context beyond the annotations and is fully consistent with idempotentHint=true and readOnlyHint=false — no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences that are dense but each earns its place: purpose, destination, validation behavior, usage triggers, and idempotency. It is front-loaded with the purpose. Slightly long for a simple feedback tool, but appropriate given the complexity of the allowUnknownTool edge case.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no output schema, the description covers purpose, when-to-use, validation rules, file destination, and idempotency. The only notable gap is that it doesn't describe what the call returns on success (confirmation, path, etc.), which is a minor omission for an agent invoking this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 8 parameters are already documented in the schema — the baseline is 3. The description adds meaningful context only for allowUnknownTool (explaining the missing-capability use case), which is a small but genuine contribution.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: 'Report a bug, improvement, or feature request about a tool of THIS server.' It clearly distinguishes itself from the feedback-management siblings (list_feedback, close_feedback) by being the creation/entry tool, and it references list_tools for discovering valid targets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggers: 'Use this when a tool produces unexpected results, crashes, or when you need a new capability.' It also defines when NOT to set allowUnknownTool (only for missing-capability reports). It could be stronger by naming list_feedback/close_feedback as the alternatives for viewing or closing feedback, but the guidance is otherwise clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_cwdARead-onlyIdempotent
Verifies whether the given path (or a friendly project name) is inside the allowed roots of this MCP server and returns the exact "cwd" to use for running commands. Use this tool when you need to find out whether — and with which "cwd" — you can work in a specific project (e.g. your workspace folder). On success: { ok: true, cwd, matchedRoot, name? }; on failure: { ok: false, error, roots }. Runs no commands — it only validates the configuration.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Path or friendly project name to verify |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/closed-world, so safety is covered. The description adds real value beyond them by disclosing that it "runs no commands — it only validates configuration" and by spelling out the success ({ ok, cwd, matchedRoot, name? }) and failure ({ ok, error, roots }) payloads in the absence of an output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and each sentence contributes: use case, return shapes, and the no-side-effects guarantee. Slightly dense with escaped quotes, but well ordered and free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by documenting both success and failure shapes plus the fact that no commands are executed. For a single-parameter validator this is sufficient for an agent to call it correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single path parameter, so the baseline is 3. The description's "or a friendly project name" adds a small nuance but essentially restates what the schema already says; no format, resolution order, or matching rules are provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource: verifies whether a path or friendly project name falls inside the server's allowed roots, and returns the exact cwd to use. An agent can distinguish this from siblings like list_allowed_roots or run_safe_command without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use this tool when you need to find out whether — and with which cwd — you can work in a specific project" gives a clear triggering context. However, it never names the obvious alternative (list_allowed_roots) or states any when-not condition, so routing between the two roots-related tools is still left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_command_grepADestructive
Runs a command and returns only lines matching the given pattern (case-insensitive regex). Use instead of run_safe_command when you know in advance that the output will be long and you only care about a specific pattern (e.g. searching for 'error' in build output, finding a specific test in test runner output). The command runs in the directory given by the cwd parameter and is subject to the same safety checks as run_safe_command. This is the replacement for Unix 'grep' on Windows — filtering happens in-process, so grep/head/tail are not needed. LIMITATION: The matches themselves can be too long — if you need more control, use run_safe_command first and then read_log_slice on the saved log file. Instead of Unix patterns like 'cmd /c ... | findstr ... & echo DONE' use THIS tool with a pattern — it filters in-process. If the task targets a project other than the primary one (MCP_PROJECT_ROOT), always pass the "cwd" parameter. Get the list of allowed roots via the "list_allowed_roots" tool.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | The directory in which the command will be run (must be inside MCP_PROJECT_ROOT / MCP_EXTRA_ROOTS). Default: the primary project (MCP_PROJECT_ROOT). | |
| command | Yes | Command to execute | |
| pattern | Yes | Pattern (regular expression) to filter lines | |
| timeoutMs | No | Timeout in milliseconds (1,000 – 600,000, default 60,000). You can extend it for longer tests/builds, e.g. 180,000 for jest. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and openWorldHint=true, so the description's job is lighter. It adds real context: it is 'subject to the same safety checks as run_safe_command', filtering is in-process, the Windows grep replacement rationale, and an explicit LIMITATION about over-long matches. It does not spell out the destructive/command-execution risk itself, but the safety-check statement covers most of what an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and purposeful, but the Windows/grep and 'use THIS tool with a pattern' points are made across multiple sentences and overlap with the LIMITATION note, so a few sentences do not fully earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description implies the return is filtered matching lines. Combined with annotations covering safety and instructions covering cwd, timeouts, and alternatives, an agent has enough to call it correctly; only the exact return shape (e.g. whether line numbers/context are included) is unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds semantics the schema lacks: the pattern is a case-insensitive regex, and cwd scoping ties to MCP_PROJECT_ROOT/MCP_EXTRA_ROOTS with a pointer to list_allowed_roots. It also notes when to always pass cwd (non-primary projects).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Runs a command and returns only lines matching the given pattern') and immediately distinguishes itself from the sibling run_safe_command by scope. An agent can pick this without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('when you know in advance that the output will be long and you only care about a specific pattern'), names the alternative (run_safe_command) and the fallback (read_log_slice). Also tells the agent what NOT to do (Unix pipes/findstr).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_destructive_commandADestructive
Use ONLY when run_safe_command rejected the command AND the user explicitly confirmed in the chat that they want to run it despite the risk. NEVER set confirm:true automatically in reaction to a rejection from run_safe_command. First restate the risk to the user in your own words (exactly what the command will do and what it could break) and wait for their explicit 'yes' or 'I confirm' in the next message. If the user is not present in the conversation (e.g. an automated run without a human), do not use this tool at all. EXAMPLE: If the user says 'do it' for a general task and you then hit a dangerous rejection, that is not sufficient confirmation — you must explain the specific risk and get a new explicit confirmation. If the task targets a project other than the primary one (MCP_PROJECT_ROOT), always pass the "cwd" parameter. Get the list of allowed roots via the "list_allowed_roots" tool.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | The directory in which the command will be run (must be inside MCP_PROJECT_ROOT / MCP_EXTRA_ROOTS). Default: the primary project (MCP_PROJECT_ROOT). | |
| command | Yes | The command to execute (runs in the directory given by the cwd parameter) | |
| confirm | No | Confirmation that you are aware of the risk (required for dangerous commands) | |
| maxLines | No | Maximum number of output lines (default: 200) | |
| timeoutMs | No | Timeout in milliseconds (1,000 – 600,000, default 60,000). You can extend it for longer tests/builds, e.g. 180,000 for jest. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, non-idempotent, open-world, but the description goes well beyond them: it discloses the mandatory confirmation protocol, the no-auto-confirm rule, the headless/no-human exclusion, and the cwd requirement for non-primary projects. That is genuine behavioral context an agent needs to invoke this safely.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the strictest rule ('Use ONLY when...'), then qualification, then a concrete counterexample. Slightly long and the risk-restatement instruction is reiterated in the EXAMPLE, but nearly every sentence carries a distinct constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, no-output-schema tool this covers everything an agent needs: precondition chain, confirmation workflow, headless behavior, and the cwd root-discovery pointer to list_allowed_roots. Timeout/output limits are already in the schema, so their absence from the description is fine.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: cwd must be set when the target project differs from MCP_PROJECT_ROOT, and confirm must never be set reflexively. The confirm parameter's social contract is only defined here, not in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description establishes that this tool executes a command that the safe path rejected, and explicitly names the sibling it is gated behind (run_safe_command). The verb 'run' is implied rather than stated outright, but the risk-restatement instruction ('exactly what the command will do and what it could break') makes the effect unambiguous. Distinguishable from all siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use (run_safe_command rejected AND user explicitly confirmed), when-not (no human present, general 'do it', automatic confirm:true), and the alternative (run_safe_command) is named. The negative case of insufficient confirmation is even exemplified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_safe_commandADestructive
Executes a safe command inside the project folder. This is the default command execution tool — use it whenever you are not sure whether a command is dangerous. Commands outside the project or that look dangerous are rejected automatically. If the tool returns isError:true with a rejection message (dangerous or directory_escape), DO NOT try to bypass it by rewriting the command or immediately switching to run_destructive_command without asking the user first. For dangerous operations use run_destructive_command with confirm:true. NOTE: 60s limit (optionally extend via timeoutMs) — not suitable for dev servers / watch mode. Run test suites (jest/npm test), typecheck and builds through this tool or run_command_grep, NOT through the built-in terminal. If the task targets a project other than the primary one (MCP_PROJECT_ROOT), always pass the "cwd" parameter. Get the list of allowed roots via the "list_allowed_roots" tool.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | The directory in which the command will be run (must be inside MCP_PROJECT_ROOT / MCP_EXTRA_ROOTS). Default: the primary project (MCP_PROJECT_ROOT). | |
| command | Yes | The command to execute (runs in the directory given by the cwd parameter) | |
| maxLines | No | Maximum number of output lines (default: 200) | |
| timeoutMs | No | Timeout in milliseconds (1,000 – 600,000, default 60,000). You can extend it for longer tests/builds, e.g. 180,000 for jest. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: it discloses automatic rejection of out-of-project or dangerous commands, the isError:true rejection messages (dangerous / directory_escape), an explicit anti-bypass policy, the 60s default timeout with timeoutMs extension, and the cwd requirement for non-primary projects. The destructiveHint:true annotation sits in slight tension with the 'safe' framing, but the description's boundary ('rejects dangerous') and escalation path resolve it rather than contradict it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and the safety boundary, and every sentence carries operational information. It is dense and heavy with emphatic capitalization (DO NOT, NOTE), which slightly hurts readability but not correctness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a command-execution tool with no output schema, it covers what an agent needs: rejection semantics, escalation path, timeout limits, working-directory rules, and how to discover allowed roots via list_allowed_roots. Nothing material is left to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: when to pass cwd (tasks targeting a project other than MCP_PROJECT_ROOT) and a concrete timeoutMs example (180,000 for jest). Only maxLines is left entirely to the schema, so it does not reach a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Executes a safe command inside the project folder') and immediately positions itself against the sibling run_destructive_command. An agent can tell the two apart without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives the default-use rule ('use it whenever you are not sure whether a command is dangerous'), names the alternative for dangerous operations (run_destructive_command with confirm:true), and states a when-not (not suitable for dev servers / watch mode), plus routing guidance for test suites and builds.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
split_file_by_declarationsADestructive
Split a large file into multiple smaller files based on top-level declarations. Optionally generates a combining file (mod.rs / index.ts / init.py). Use dryRun: true (default) to preview the layout before writing.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Working dir for resolving relative file paths (default: primary project root) | |
| file | Yes | Source file to split | |
| dryRun | No | Preview only — write nothing (default: true) | |
| grouping | Yes | Module groupings | |
| language | No | Language (auto-detected from extension) | |
| overwrite | No | Allow overwriting existing target files (default: false) | |
| targetDir | No | Where new files are written (default: dirname of file) | |
| generateIndex | No | Create combining file (default: true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered; the description adds real value beyond that by disclosing that dryRun defaults to true (preview-before-write) and that a combining file is optionally generated. It still omits explicit overwrite/destruction semantics, but the default-safe behavior is a meaningful addition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: purpose first, then the auxiliary output, then the safety-relevant default. No filler, and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutating tool with no output schema, the description plus fully-covered schema and destructiveness annotations give the agent enough to call it correctly. Minor gaps remain around failure/return behavior, but the schema and annotations carry the rest.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all 8 parameters (cwd, file, dryRun, grouping, language, overwrite, targetDir, generateIndex). The description only reinforces dryRun and the combining-file/index concept, adding nothing the schema doesn't already convey. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Split) + resource (a large file) + mechanism (based on top-level declarations), with the auxiliary combining-file behavior named. It does not explicitly differentiate from potentially adjacent siblings such as generate_module_skeleton or batch_apply_edits, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case (splitting large files) and gives safety-oriented guidance for the dryRun flag, but never states when to prefer this over alternatives like generate_module_skeleton. Usage is only implied, not routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
universal_find_referencesARead-onlyIdempotent
Find all occurrences of a symbol across a workspace. Structured output with file, line, column, context. Optional language-aware mode (rust/typescript/python/cpp) adds role annotations: declaration, import, or usage. Use this tool BEFORE any refactoring session to understand what will break when a symbol is renamed or moved. When cwd is omitted, ALL registered roots are searched: nested roots are pruned and files are deduplicated by real path, so no match is listed or counted twice.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Workspace root to search (default: all registered roots — nested roots pruned, duplicates removed) | |
| symbol | Yes | Symbol to search for (word-boundary match) | |
| language | No | Optional language-aware mode for role detection | |
| contextLines | No | Lines of context around each match (default: 1) | |
| fileExtensions | No | Restrict to these extensions (default: common source extensions) | |
| excludePatterns | No | Directories to skip (default: .git, node_modules, target, build, dist, __pycache__) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral details beyond the annotations: it explains that nested roots are pruned and files deduplicated, so no match is counted twice, and that language-aware mode adds role annotations. These details are not present in the annotations and help the agent predict output. It also notes the structured output format, which is useful. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded: it starts with the primary action, then output format, then options, then usage guidance. Each sentence adds value without redundancy. It is well-organized and avoids unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete for an agent to invoke the tool correctly: it covers the primary purpose, output structure, optional modes, default behaviors, and a clear use case. The schema and annotations provide the remaining parameter details and safety profile, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers all parameters with descriptions, so baseline is 3. The description adds extra meaning for 'cwd' by explaining the default (all registered roots) and deduplication behavior, and for 'language' by clarifying it adds role annotations. These enrich the schema without redundancy, so a 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Find') and resource ('all occurrences of a symbol across a workspace'), and mentions structured output. It clearly distinguishes the tool's core function, though it doesn't explicitly name sibling tools for differentiation. The mention of using it before refactoring provides context that helps an agent understand its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs to use this tool before any refactoring session, which is a clear usage guideline. It also explains the default behavior when cwd is omitted, which informs the agent about scope. However, it doesn't mention alternatives or conditions when not to use it, but the primary use case is well-defined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_refactor_safetyARead-onlyIdempotent
Semantic diff between old and new code. Catches accidental deletions before compilation. Checks: function count, signatures, export count, imports, comment ratio. Intentionally conservative — renames appear as errors requiring explicit confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| after | Yes | New code text | |
| before | Yes | Original code text | |
| language | No | Language (auto-detected from content) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description adds genuinely non-obvious behavior beyond that: it is deliberately conservative, and renames will surface as errors requiring explicit confirmation. That false-positive profile materially affects how an agent interprets results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: purpose, motivation, checks performed, and the conservative-bias caveat. Front-loaded with the core definition and free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the checks it reports (function count, signatures, exports, imports, comment ratio) and the rename-confirmation behavior. It stops short of describing the result shape or whether mismatches block anything, but an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so before/after/language are already documented, including language auto-detection and the enum. The description adds no syntax or format detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: a semantic diff between old and new code, plus the concrete checks it performs. This is clearly distinguishable from every sibling tool, none of which do source-diff analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Catches accidental deletions before compilation" gives a clear usage context (pre-build verification after a refactor). It does not name an alternative tool or state when not to use it, but for this toolkit there is no obvious competing option.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.2.1- Changed
batch_apply_edits1 field changed- added
Input schema / properties / edits / items / properties / excludePatternsAdded value: +{ + "description": "Patterns to exclude from replaceAll (e.g. '#[cfg(test)]' to skip test modules)", + "items": { + "type": "string" + }, + "type": "array" +}
- Changed
list_feedback1 field changed- added
Input schema / properties / archivedAdded value: +{ + "description": "true = list archived (closed) entries from .mcp/FEEDBACK_ARCHIVE.md instead of the active log", + "type": "boolean" +}
- Changed
report_tool_feedback2 fields changed- added
Input schema / properties / allowUnknownToolAdded value: +{ + "description": "File feedback about a name that is not a tool of this server (default false) — use only for missing-capability reports", + "type": "boolean" +} - changed
Input schema / properties / tool / descriptionPrevious value: -"Name of the MCP tool this feedback is about"New value: +"Name of the MCP tool this feedback is about (must be a tool of this server, e.g. 'batch_apply_edits')"
- Changed
universal_find_references1 field changed- changed
Input schema / properties / cwd / descriptionPrevious value: -"Workspace root to search (default: primary project root)"New value: +"Workspace root to search (default: all registered roots — nested roots pruned, duplicates removed)"
1 tool update
v0.1.1- Changed
batch_apply_edits1 field changed- added
Input schema / properties / cwdAdded value: +{ + "description": "Working dir for resolving relative file paths (default: primary project root)", + "type": "string" +}
17 tool updates
v0.1.0- First observed
batch_apply_edits - First observed
close_feedback - First observed
extract_code_block - First observed
generate_module_skeleton - First observed
help_tool - First observed
list_allowed_roots - First observed
list_feedback - First observed
list_tools - First observed
read_log_slice - First observed
report_tool_feedback - First observed
resolve_cwd - First observed
run_command_grep - First observed
run_destructive_command - First observed
run_safe_command - First observed
split_file_by_declarations - First observed
universal_find_references - First observed
verify_refactor_safety
TDQS
Scored across 17 tools
Most tools have clearly separated jobs (command execution vs code extraction vs refactoring vs feedback), and paired tools like run_safe_command/run_destructive_command are explicitly distinguished. The main overlap is between list_allowed_roots and resolve_cwd, which both address allowed-root/cwd discovery, though their descriptions do enough to mostly keep them apart.
All tool names follow a consistent snake_case verb_noun pattern, with predictable families like run_*, list_*, and *feedback. There is no style mixing or vague generic naming, and the command tools (run_safe_command, run_destructive_command, run_command_grep) form a particularly coherent set.
17 tools is slightly above the ideal 3-15 range, but the server covers several distinct subdomains: command execution, code analysis, refactoring, and feedback management. The count feels a bit heavy due to near-redundant helpers like list_allowed_roots and resolve_cwd, but most tools have a clear purpose.
The tool surface covers major workflows well: command execution (safe/destructive/grep/log reading), code analysis (find references/extract), refactoring (split/generate/batch edits/verify), and a full feedback lifecycle. Minor gaps exist, such as no direct arbitrary-file read/write or explicit symbol rename tool, but these can be worked around with existing tools.
Maintenance
Related MCP Connectors
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Runtime permission, approval, and audit layer for AI agent tool execution.
Develop, manage, and debug Railway projects, services, and deployments from within agents.
The trust harness for AI agents. Set what an agent can do before it acts.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides safe shell command execution capabilities for AI agents and tools like VS Code Copilot through a whitelist-based filtering system.-
- AlicenseNot gradedqualityCmaintenanceProvides secure, sandboxed file system access for AI assistants to read, write, and manage project files with controlled command execution capabilities, all confined to a designated workspace directory.MIT
- FlicenseNot gradedqualityDmaintenanceEnables safe execution of terminal commands across different shells (bash, cmd, PowerShell) with configurable timeouts, working directories, and resource limits for command-line operations through AI assistants.-
- AlicenseAqualityFmaintenanceEnables AI assistants to execute terminal commands on a host machine with configurable, granular permission controls and safety protections. It features multiple security modes, including allowlists and manual approval, to ensure safe command execution within specified directories.6Apache 2.0