Skip to main content
Glama
Majrooo

majrooo-mcp-devkit

by Majrooo

majrooo-mcp-devkit — MCP DevKit: Safe Commands + Refactoring Tools

Version License CI Tests Node.js M8ven Verified

Repository Access: PUBLIC
Version: 0.2.1 · Tests: 354 passing · License: GPL-3.0-or-later

MCP server that provides safe command execution and code refactoring tools for Cline/Claude Desktop.

Tools

run_safe_command

Execute a shell command restricted to the active project root. Dangerous commands and writes outside the active root are automatically blocked. This is the default tool — always use this first.

Parameter

Type

Default

Description

command

string

—

Command to execute

cwd

string

primary root

Working directory (must be inside MCP_PROJECT_ROOT / MCP_EXTRA_ROOTS)

maxLines

number

200

Max output lines before truncation

timeoutMs

number

60000

Command timeout (1000–600000 ms) — raise for long jest/build runs

run_destructive_command

Execute a potentially dangerous command with explicit user confirmation. Only use when run_safe_command blocked the command and the user explicitly agreed after being informed of the specific risk.

Parameter

Type

Default

Description

command

string

—

Command to execute

confirm

boolean

false

Acknowledge the risk (required for dangerous commands)

cwd

string

primary root

Working directory (must be inside MCP_PROJECT_ROOT / MCP_EXTRA_ROOTS)

maxLines

number

200

Max output lines before truncation

timeoutMs

number

60000

Command timeout (1000–600000 ms)

read_log_slice

Read a portion of a previously saved log file. Use this instead of re-running a command with higher maxLines. Files are read directly via Node.js (not through the shell), so it also works for truncated logs in os.tmpdir().

Parameter

Type

Default

Description

logPath

string

—

Path to the log file

startLine

number

0

Starting line (0-based)

lineCount

number

100

Number of lines to read

list_allowed_roots

Return the registered roots configuration: the primary project (MCP_PROJECT_ROOT), all allowed roots (MCP_EXTRA_ROOTS, including globs), the concrete existing project directories under them (usable as cwd), and whether MCP_BLOCK_CROSS_ROOT_READS is enabled. Projects with a friendly name are returned as { path, name } — in that case you can also use the name as cwd. Call this before working in any non-primary project to discover the exact cwd value to use. Runs no commands — it only reads configuration and lists directories.

No parameters.

resolve_cwd

Verify whether a path (or a friendly project name from MCP_PROJECT_NAMES) is inside the allowed roots and get the exact cwd to use for commands. Pass the path you want to work in (e.g. your workspace folder) instead of guessing.

On success returns { ok: true, cwd, matchedRoot, exists, name? }; on failure { ok: false, error, roots }. exists tells whether the resolved directory actually exists on disk (relative cwd values are resolved against the primary project). Runs no commands — it only validates configuration.

Parameter

Type

Description

path

string

Path to verify (absolute, e.g. the project workspace folder)

run_command_grep

Execute a command and return only lines matching a pattern (case-insensitive regex). Use instead of run_safe_command when you only care about specific lines (e.g., errors in build output). This is the replacement for Unix grep on Windows — filtering happens in-process, so grep/head/tail are not needed.

Parameter

Type

Description

command

string

Command to execute

pattern

string

Regex pattern to filter lines (case-insensitive)

cwd

string

Working directory (default: primary root)

timeoutMs

number

60000

universal_find_references

Find all occurrences of a symbol across a workspace. Returns structured output with file, line, column, context, and optional role annotations. Use this before any refactoring session to understand what will break when a symbol is renamed or moved.

Parameter

Type

Default

Description

symbol

string

—

Symbol to search for (word-boundary match)

cwd

string

all registered roots

Workspace root to search. When omitted, all registered roots are searched (nested roots pruned, duplicate files removed)

fileExtensions

string[]

common source extensions

Restrict to these extensions

excludePatterns

string[]

.git, node_modules, target, ...

Directories to skip

contextLines

number

1

Lines of context around each match

language

string

— (disabled)

Optional: "rust", "typescript", "python", or "cpp" — enables role detection (declaration/import/usage)

Multi-root search (no cwd): every registered root is searched. A root nested inside another root is pruned — the outer root already covers it — and files are deduplicated by resolved real path, so the same match is never listed or counted twice. When more than one root is searched, the report states them and, for each file group, the root its relative path is based on:

Symbol: ColAlign
Total matches: 28
Searched roots (3):
  - D:\W\TS
  - d:\Users Data\jox\My Documents\Python
  - d:\Users Data\jox\My Documents\Rust
Duplicates skipped: 2 (same file reachable through a nested root)

pixel-blaster-engine/engine_bevy/src/lib.rs  (relative to d:\Users Data\jox\My Documents\Rust)
  Line 72:12 [usage] — pub use ui_panel::{

A path like project/src/lib.rs therefore always belongs to exactly one root — it is never a second copy of src/lib.rs. With an explicit cwd only that root is searched and the report keeps its compact form (path/file.rs:).

extract_code_block

Read the full text of a function, struct, class, or method from a file. Returns precise line range + content. Includes leading annotations (#[derive], @decorator, /// doc comments). String/comment-aware bracket matching prevents false depth counts from braces inside strings or comments.

Parameter

Type

Default

Description

file

string

—

Source file path (must resolve inside allowed root)

symbol

string

—

Symbol name to extract

contextLines

number

0

Extra lines before/after the block

cwd

string

primary root

Working dir for resolving relative file paths

split_file_by_declarations

Split a large file into multiple smaller files based on top-level declarations. Optionally generates a combining file (mod.rs / index.ts / __init__.py). Use dryRun: true (default) to preview the layout before writing.

Parameter

Type

Default

Description

file

string

—

Source file to split

grouping

object[]

—

[{ module, symbols }] — module groupings

targetDir

string

dirname(file)

Where new files are written

language

string

auto-detect

"rust", "typescript", "python", "cpp"

generateIndex

boolean

true

Create combining file

dryRun

boolean

true

Preview only — write nothing

overwrite

boolean

false

Allow overwriting existing targets

cwd

string

primary root

Working dir for resolving relative file paths

batch_apply_edits

Apply multiple file edits with partial rollback on failure. Validates all edits first — if any search string is not found or matches multiple times (without replaceAll), NO files are modified. On failure during application, only files modified by the failed edit and subsequent edits are reverted; earlier successful edits are preserved. Same-file chain failures revert the entire file. Relative paths without cwd are resolved against the primary root; if the file doesn't exist there, all other registered roots are searched automatically (unique match → use it; multiple matches → error with instructions to specify cwd).

Parameter

Type

Default

Description

edits

object[]

—

[{ file, search, replace, description?, replaceAll?, excludePatterns? }]

dryRun

boolean

true

Preview all changes without writing

cwd

string

primary root

Working dir for resolving relative file paths

Response (explicit outcome)

The result text starts with a one-line message followed by the JSON payload:

{
  "dryRun": false,
  "totalEdits": 3,
  "validated": 3,
  "appliedEdits": 2,
  "written": ["D:\\proj\\a.rs", "D:\\proj\\b.rs"],
  "message": "Applied 2 of 2 edit(s) to 2 file(s) — all changes are on disk.",
  "preview": [{ "file": "D:\\proj\\a.rs", "action": "edit", "matchCount": 1, "applied": true }]
}

Key fields:

Field

Meaning

message

Human-readable summary — always states explicitly whether anything was written

appliedEdits

Number of edits whose changes are on disk after the run (0 in dry-run, 0 when nothing was written)

written

Files changed on disk after the run (empty in dry-run / when nothing was written)

preview[].applied

true only when that specific edit's change is on disk (never true in dry-run or after rollback)

preview[].matchCount

Occurrences found — informational, not proof that anything was written (check applied)

On failure the response also carries error, failedAt, reason (validation_failed | write_failed), reverted (files rolled back) and nothingWritten (true = this run left no change on disk at all — either nothing was ever written, or everything written was rolled back). Validation failure example:

VALIDATION FAILED on edit #2 of 3 — NO edits were written to disk (validation-first: nothing is written until every edit validates). Reason: search string not found in D:\proj\b.rs

Edits that were never evaluated get an explicit action: "error" preview entry (not evaluated — batch stopped at edit #N), so preview always maps 1:1 to the edits array. When validation fails on a file that already had a successfully applied edit in the same run (chained edits), the response switches to partial rollback: N edit(s) kept in … ; M file(s) reverted […]. If every write from the run was rolled back (e.g. all edits chained on one file), nothingWritten stays true and the message says NO net changes were left on disk: N edit(s) had been written and M file(s) were rolled back […] — an honest distinction between "never wrote" and "wrote, then undid".

generate_module_skeleton

Generate a new module file with extracted symbols from a source file. Returns error with unknownSymbols list if any symbols are not found.

Parameter

Type

Default

Description

modulePath

string

—

Target file path

symbols

string[]

—

Symbol names to include

sourceFile

string

—

Original file to extract from

language

string

auto-detect

"rust", "typescript", "python"

dryRun

boolean

true

Preview only

overwrite

boolean

false

Allow overwriting existing file

cwd

string

primary root

Working dir for resolving relative file paths

verify_refactor_safety

Semantic diff between old and new code. Catches accidental deletions before compilation. Checks: function count, signatures, export count, imports, comment ratio. Intentionally conservative — renames appear as errors requiring explicit confirmation.

Parameter

Type

Default

Description

before

string

—

Original code text

after

string

—

New code text

language

string

auto-detect

"rust", "typescript", "python", "cpp"

report_tool_feedback

Report a bug, improvement, or feature request about a tool of this server (see list_tools). Writes structured feedback to .mcp/FEEDBACK.md (project-specific, gitignored). Entries are idempotent — duplicate reports are skipped.

The tool name is validated against this server's tool registry: an unknown name is rejected (nothing is written) with a "did you mean …?" suggestion, so feedback about another client's built-in tool no longer lands in this log. For a missing-capability report about the server as a whole, pass allowUnknownTool: true.

Parameter

Type

Default

Description

type

string

—

"bug", "improvement", or "feature_request"

tool

string

—

Name of the MCP tool this feedback is about — must be a tool of this server (e.g. batch_apply_edits)

title

string

—

Short summary (1 line)

description

string

—

Detailed description

reproduction

string

—

Steps to reproduce (optional)

expected

string

—

What you expected (optional)

suggestion

string

—

Suggested fix or improvement (optional)

allowUnknownTool

boolean

false

Accept a name that is not a tool of this server (missing-capability reports only)

list_feedback

List feedback entries from .mcp/FEEDBACK.md. Optionally filter by type, tool name, or status. Use this to check existing feedback before creating new entries.

Parameter

Type

Default

Description

type

string

—

Filter: "bug", "improvement", or "feature_request"

tool

string

—

Filter by tool name

status

string

—

Filter: "open" or "closed"

archived

boolean

false

true = read .mcp/FEEDBACK_ARCHIVE.md (closed entries moved out of the active log) instead

close_feedback

Close an existing feedback entry by ID — sets status to "closed" and optionally adds resolution text. Use this to mark feedback items as resolved after fixing them.

Closing also moves every closed entry out of the active log into .mcp/FEEDBACK_ARCHIVE.md, so FEEDBACK.md keeps holding only open items and stays small. The move is append-first (archived before removed) and idempotent, and it self-heals: entries closed before this behaviour existed are migrated along with the next close. Read them back with list_feedback + archived: true.

Parameter

Type

Default

Description

id

string

—

The feedback entry ID to close (from list_feedback output)

resolution

string

—

Resolution note explaining how the issue was addressed (optional)

Closed entries are stored in two files:

File

Contents

.mcp/FEEDBACK.md

Header + open entries (the active log)

.mcp/FEEDBACK_ARCHIVE.md

Header + closed entries, moved verbatim (description, reproduction, resolution, closedAt preserved)

list_tools

List all available MCP tools with descriptions. Use this to discover tools before starting a task. Filterable by category.

Parameter

Type

Default

Description

category

string

—

Filter: "command", "refactoring", or "feedback"

help_tool

Get detailed help for a specific MCP tool — parameters, types, defaults, and description.

Parameter

Type

Default

Description

tool

string

—

Tool name to get help for

Related MCP server: MCP Workspace Server

Supported Languages

The refactoring tools support four languages for declaration parsing, symbol extraction, and role detection:

Language

Status

Role detection

Split

Skeleton

Rust

✅ Fully tested

✅

✅

✅

TypeScript

Experimental

✅

✅

✅

Python

Experimental

✅

✅

✅

C++

Experimental

✅

✅

❌

Tools without a language parameter (batch_apply_edits, extract_code_block) are language-agnostic — they operate on plain text and work with any language.

Note: Rust has been tested in production refactoring scenarios. Other languages are structurally supported but have not been tested against real-world codebases yet.

Configuration

The server supports one instance, many projects. Projects are selected per command via the cwd parameter; the active project also acts as the "lockbox" for write/read checks.

Env var

Description

MCP_PROJECT_ROOT

Primary project root (default cwd when omitted). If unset, the server's own directory is used (derived from the module location, not process.cwd()).

MCP_EXTRA_ROOTS

Additional roots, semicolon separated.

MCP_PROJECT_NAMES

Friendly names for projects, semicolon separated path=name pairs (see below). Alias paths are also included in the allowed roots, so projects registered only via this variable can be used with file-based tools.

MCP_BLOCK_CROSS_ROOT_READS

1 or true → opt-in best-effort blocking of obvious reads outside the active root.

Entry forms supported in both variables:

  • Plain path D:\W\TS\majrooo-mcp-devkit → prefix: the directory itself and everything below it are allowed. Registering D:\W covers all projects under it.

  • Glob D:\W\TS\* (*, **, ?) → any path matching the pattern (and its subtree) is allowed.

Example — one instance, many projects, with friendly names:

{
  "mcpServers": {
    "majrooo-mcp-devkit": {
      "command": "node",
      "args": ["D:\\W\\TS\\majrooo-mcp-devkit\\build\\index.js"],
      "env": {
        "MCP_PROJECT_ROOT": "D:\\W\\TS\\majrooo-mcp-devkit",
        "MCP_EXTRA_ROOTS": "D:\\W;D:\\python",
        "MCP_PROJECT_NAMES": "D:\\W\\TS\\cb=ZbaľSa;D:\\W\\TS\\nase-zasoby=Naše zásoby"
      }
    }
  }
}

MCP_PROJECT_NAMES maps a real project path to a readable name. This is useful when the folder name had to be shortened (e.g. Gradle path-length limits) or the project was renamed. The name can be used directly as cwd (e.g. "cwd": "ZbaľSa"), and list_allowed_roots will show such projects as { "path": "D:\\W\\TS\\cb", "name": "ZbaľSa" }.

Switching projects is done via the cwd parameter, never via cd in the command. cd .., cd ~, cd C:\..., and Windows cd /d D:\... are always rejected.

Running tests / long commands (the anti-freeze workflow)

Never run test suites (jest/npm test), typecheck or builds through the Cline built-in terminal — it has no timeout and can freeze the whole window. Use the MCP tools instead:

  1. Always pass the project's cwd (or friendly name, e.g. "cwd": "ZbaľSa").

  2. To filter output (e.g. jest summary), use run_command_grep — filtering happens in-process, so cmd /c "... | findstr ... & echo DONE" is not needed and discouraged:

{
  "tool": "run_command_grep",
  "cwd": "ZbaľSa",
  "command": "npx jest src/app/__tests__/catalog.test.tsx 2>&1",
  "pattern": "Tests:|Test Suites:|FAIL|PASS|✕",
  "timeoutMs": 180000
}
  1. For full output use run_safe_command with a small maxLines — the full output is saved to a temp log for read_log_slice:

{
  "tool": "run_safe_command",
  "cwd": "ZbaľSa",
  "command": "npm run typecheck 2>&1",
  "maxLines": 100,
  "timeoutMs": 180000
}
  1. If a run exceeds 10 minutes, run it in the background, redirect to a log file, and poll the log via run_command_grep — do not watch live terminal output.

Typical workflow

  1. Call list_allowed_roots to see the primary root, the allowed roots (including globs), and the concrete projects under them.

  2. If you need to confirm a specific path, call resolve_cwd with your workspace folder — it returns the exact cwd to use and the matched root.

  3. If the task targets a project other than the primary one, pass the resolved path as cwd on every command (run_safe_command, run_destructive_command, run_command_grep).

  4. Otherwise, omit cwd — commands run in the primary root.

Relative paths: File-based refactoring tools (batch_apply_edits, extract_code_block, split_file_by_declarations, generate_module_skeleton) resolve relative paths against the primary root when cwd is omitted. If the file doesn't exist in the primary root (or in the cwd root when cwd is provided), all other registered roots are searched automatically — a unique match is used directly, while multiple matches produce an error with instructions to specify cwd.

Safety Mechanisms

Layer

Description

Registered roots

MCP_PROJECT_ROOT / MCP_EXTRA_ROOTS / MCP_PROJECT_NAMES define the allowed project registry (prefix or glob). cwd must match one of them.

Directory restriction

Commands execute with cwd set to the resolved project root. cd .., cd ~, absolute-path cd, and Windows cd /d are rejected.

Dangerous pattern detection

Regex blacklist blocks destructive commands (rm -rf, format, shutdown, git push --force, fork bombs, pipe-to-shell, etc.).

Write-target check

Best-effort detection of writes outside the active root (>, >>, 2>, copy, move, mkdir, tee, curl -o, ...). Writing from project A into project B is blocked even if B is registered — pick B via cwd instead.

Cross-root read check (opt-in)

MCP_BLOCK_CROSS_ROOT_READS=1 blocks obvious reads outside the active root (type/cat/Get-Content/git -C/Node/Python path reads...). Best-effort heuristic.

Explicit confirmation

run_destructive_command requires confirm: true for dangerous commands.

Missing destructive target

A confirmed destructive command (rmdir/del/erase/Remove-Item) whose target does not exist in the active cwd is rejected with a "set the cwd parameter" message (reason missing_destructive_target) instead of a raw The system cannot find the file specified.

Buffer & timeout limits

50 MB max output, 60-second timeout.

Output truncation

Long outputs are saved to os.tmpdir() for later inspection via read_log_slice.

Audit log

All executions logged to os.tmpdir()/mcp-command-audit.log (rotated to .old once it exceeds 5 MB).

Limitations

The dangerous-pattern blacklist and the write/read target checks are best-effort layers, not security guarantees. Shell features (variables, command substitution, encoding) can bypass them. For production isolation use Docker/VM sandboxing.

Windows Notes

  • Commands run through cmd.exe. Unix-only tools (grep, head, tail, ...) do not exist — the server returns a friendly error with alternatives instead of a raw "not recognized" blob.

  • Use run_command_grep instead of grep, and read_log_slice or PowerShell (Get-Content out.log -TotalCount 30) instead of head.

  • Long-running processes (dev server, watch mode) exceed the 60s timeout — use the built-in terminal for those.

Output normalization

On Windows the server automatically prefixes commands with chcp 65001 > NUL && so the child process emits UTF-8 instead of the legacy OEM codepage (which would otherwise decode into U+FFFD replacement characters, e.g. around thousands separators in dir output). ANSI color codes from tools like vitest/jest are stripped as a fallback (NO_COLOR=1 / FORCE_COLOR=0 are also injected into the environment), and read_log_slice cleans ANSI codes defensively when reading older logs.

Redirected output reporting

When a command redirects its output into a file (npm test > test.log 2>&1, >> out.log, 2> err.log, ...), the response reports where the output went and shows the tail of the written file instead of an empty response or a bare Command failed: .... On failure the response also includes the exit code (or timeout/signal) and the captured stdout/stderr. Discard targets (NUL, /dev/null), wildcard patterns and fd-duplication tokens (2>&1, >&-) are skipped; very large files are read only from the end.

Agent Behavior Rules

The .clinerules file in the project root defines how AI agents should use these tools:

  1. Always try run_safe_command first — never start with run_destructive_command.

  2. run_destructive_command only after explicit user confirmation in the current conversation — general consent ("do what you need") is not sufficient.

  3. Never bypass directory_escape rejections — no chaining, absolute paths, or cwd tricks. Use the cwd parameter to pick a registered project.

  4. Prefer run_command_grep / read_log_slice over increasing maxLines for long output, and instead of Unix grep/head/tail.

  5. Use the built-in terminal only for quick interactive checks (e.g., git status). Large-output commands (npm install, build, tests) must go through run_safe_command.

  6. Run tests before reporting task as complete.

License

This project is licensed under the GNU General Public License v3.0 or later - see the LICENSE file for details.

Development

npm run build    # Compile TypeScript
npm test         # Run unit tests
npm run test:watch  # Watch mode

The server communicates over STDIO using the Model Context Protocol.

Available Tools

17 tools
batch_apply_editsA
Destructive

Apply multiple file edits with partial rollback on failure. Validates all edits first — if validation fails, response says explicitly how many edits were applied and which files were written/reverted (NO files are modified). On failure during application, only the failed edit and later edits are reverted; earlier successful edits are preserved. Every response includes message, appliedEdits, written, and per-preview-entry applied. Use dryRun: true (default) to preview changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoWorking dir for resolving relative file paths (default: primary project root)
editsYesList of edits to apply
dryRunNoPreview all changes without writing (default: true)

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes far beyond the annotations. Annotations indicate destructiveHint=true and readOnlyHint=false, but the description details the exact rollback semantics: validation failure means no files modified, and on application failure only the failed and later edits are reverted, preserving earlier ones. It also specifies the response structure (message, appliedEdits, written, per-preview applied). This is comprehensive behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise yet information-dense. Each sentence adds critical information: the core behavior, validation guarantee, rollback specifics, response fields, and dryRun default. It is front-loaded with the main purpose and uses no fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with this complexity (multiple edits, validation, partial rollback, dry run), the description covers all essential aspects an agent needs: validation outcomes, failure behavior, success preservation, response format, and dryRun default. No output schema exists, so describing the response fields is necessary and well done. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters (cwd, edits, dryRun) with detailed descriptions. The tool description does not add extra meaning beyond what the schema provides; it mentions dryRun but without new details. Per the rubric, baseline 3 is appropriate when schema carries the parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Apply multiple file edits with partial rollback on failure.' It specifies the verb (apply), resource (multiple file edits), and key behavioral nuance (partial rollback). This is specific and distinguishes it from sibling tools like run_destructive_command, which focus on command execution rather than file editing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context by explaining when to use dryRun: 'Use dryRun: true (default) to preview changes.' It also describes the validation-first behavior, implying when to rely on the tool for safe multi-edit operations. However, it does not explicitly contrast with alternative tools or state when not to use it, so it lacks explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

close_feedbackA
Destructive

Close an existing feedback entry by ID — sets status to "closed" and optionally adds resolution text. Closed entries are automatically moved out of .mcp/FEEDBACK.md into .mcp/FEEDBACK_ARCHIVE.md (read them with list_feedback archived:true). Use this to mark feedback items as resolved after fixing them.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe feedback entry ID to close (from list_feedback output)
resolutionNoResolution note explaining how the issue was addressed

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations include destructiveHint: true, and the description goes beyond it by explaining exactly what happens: status is set to 'closed', resolution text is optionally added, and entries are moved from FEEDBACK.md to FEEDBACK_ARCHIVE.md. This adds valuable behavioral context (file movement) that annotations do not provide, and it does not contradict the destructive hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no fluff. The core action is front-loaded, the side effect is explained in the second sentence, and the usage instruction is in the third. Every sentence earns its place, and the structure is logical and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with a clear schema and no output schema, the description covers all essential context: what it does, the automatic file movement, how to read archived entries, and when to use it. Nothing that an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both id and resolution are already documented in the input schema. The description adds only a minor repetition of 'optionally adds resolution text,' which is already in the schema. No additional meaning is provided beyond what the schema offers, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Close'), the resource ('feedback entry'), and the key behavior (sets status to 'closed', optional resolution). It also distinguishes itself from related tools like list_feedback by explaining the archive movement, so an agent can immediately understand its role without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this to mark feedback items as resolved after fixing them,' which provides a clear when-to-use directive. It also mentions the reading alternative (list_feedback archived:true), fulfilling the requirement for mentioning alternatives. There is no misleading guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_code_blockA
Read-onlyIdempotent

Read the full text of a function, struct, class, or method from a file. Returns precise line range + content. Includes leading annotations (#[derive], @decorator, /// doc comments). String/comment-aware bracket matching prevents false depth counts from braces inside strings or comments.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoWorking directory for resolving relative file paths (default: primary project root)
fileYesSource file path (absolute or relative to cwd)
symbolYesSymbol name to extract
contextLinesNoExtra lines before/after the block (default: 0)

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), yet the description adds real behavioral detail beyond them: it returns a precise line range plus content, it includes leading annotations and doc comments, and it uses string/comment-aware bracket matching to avoid false depth counts from braces inside literals. That is substantive context the annotations do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with no filler: the core action is front-loaded, followed by the return shape and then the extraction-safety note. Every sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly steps in to describe the return value (line range + content) and the extraction semantics (annotation inclusion, bracket matching). It stops short of covering failure behavior, such as what happens when the symbol is not found or is ambiguous, which is the main remaining gap for a tool with a required symbol lookup.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (cwd, file, symbol, contextLines) are already documented in the schema and the baseline is 3. The description adds no parameter-level detail, though 'Includes leading annotations' implicitly relates to how extraction boundaries are computed rather than to any parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: 'Read the full text of a function, struct, class, or method from a file,' which is far more precise than a generic 'read file' tool. It does not explicitly name or contrast with siblings like universal_find_references or split_file_by_declarations, so the agent must infer the boundary, but the symbol-level extraction purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the phrasing (you call it to obtain the full body of a named symbol), but there is no explicit when-to-use/when-not or reference to alternatives such as find_references or split_file_by_declarations. An agent can infer the context but gets no routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_module_skeletonA
Destructive

Generate a new module file with correct imports, declarations and visibility. Reads the source file, extracts the specified symbols, and writes them to the target module path. Returns error with unknownSymbols list if any symbols are not found.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoWorking dir for resolving relative file paths (default: primary project root)
dryRunNoPreview only (default: true)
symbolsYesSymbol names to include
languageNoLanguage (auto-detected)
overwriteNoAllow overwriting existing file (default: false)
modulePathYesTarget file path (e.g. src/ai/data.rs)
sourceFileYesOriginal file to extract symbols from

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false and idempotentHint=false, so the safety profile is covered. The description adds real value beyond that: it discloses the read-extract-write flow and the specific failure mode (returns an error with an unknownSymbols list when symbols aren't found), which an agent needs to handle partial failures.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and followed by mechanism and error behavior; no filler. Slightly more could be trimmed, but each sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation tool with full param coverage, no output schema, and safety annotations already present, the description covers the action, the data flow, and the error case. It stops short of clarifying dryRun default behavior or overwrite interaction, which would round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all seven parameters (cwd, dryRun, symbols, language, overwrite, modulePath, sourceFile) are already documented in the schema. The description only loosely maps to 'symbols' via 'extracts the specified symbols' and adds no format or default semantics, so the schema does the heavy lifting — baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Generate a new module file") plus the exact mechanism (reads source, extracts symbols, writes to target path). It is clearly distinguishable in intent from siblings like split_file_by_declarations, but it never explicitly names or contrasts those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what happens mechanically but gives no when-to-use guidance: no indication of when this is preferable to extract_code_block or split_file_by_declarations, and no mention of prerequisites such as the target needing to exist or symbols needing to resolve. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

help_toolA
Read-onlyIdempotent

Get detailed help for a specific MCP tool — parameters, types, defaults, description.

ParametersJSON Schema
NameRequiredDescriptionDefault
toolYesTool name to get help for

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful value by naming what is returned (parameters, types, defaults, description), which partially compensates for the absence of an output schema, but it says nothing about error behavior for unknown tool names.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that conveys purpose and return contents with no filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, read-only lookup tool, the definition is essentially complete, and it describes the return contents in the absence of an output schema. Minor gap: no mention of behavior when the requested tool does not exist.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'tool' parameter is already documented as 'Tool name to get help for'. The description adds no syntax, format, or validity constraints beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get') and resource ('detailed help for a specific MCP tool'), making the purpose unambiguous. It implicitly distinguishes itself from list_tools via the word 'specific', but does not name that sibling explicitly, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: consult this when you need details about a particular tool. There is no explicit when-to-use vs. list_tools guidance or statement of prerequisites (e.g., that the tool name must be known/valid). Adequate but with a clear gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_allowed_rootsA
Read-onlyIdempotent

Returns the allowed-roots configuration of this MCP server: the primary project (MCP_PROJECT_ROOT), all registered roots (MCP_EXTRA_ROOTS, including globs), the existing projects under them (the list of directories you can use as "cwd") and whether MCP_BLOCK_CROSS_ROOT_READS is enabled. Use THIS tool whenever you need to find out whether — and with which "cwd" parameter — you can run a command in another project. It runs no commands — it only reads the configuration and lists directories. A project with a friendly name is shown as { path, name } and you can pass its "name" as "cwd".

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive and closed-world, so safety is covered. The description adds genuinely useful context beyond that: it explicitly states it executes nothing ('only reads the configuration and lists directories') and documents the return shape, including that a friendly-named project appears as { path, name } and that the name is passable as cwd.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The payload is dense but front-loaded: what it returns comes first, usage second, behavioral caveat third, return-shape note last. Every clause carries information, though the sentence count and parenthetical environment-variable references make it denser than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no parameters, so the description carries the whole disclosure burden — and it does, covering return contents, the { path, name } shape, the glob handling of extra roots, the cross-root flag, and the non-executing guarantee. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description still usefully explains how a value surfaced by this tool (a project 'name') is later consumed as a 'cwd' argument, adding cross-tool semantic value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

It states a specific verb and resource — returns the server's allowed-roots configuration — and enumerates exactly what that comprises (primary project, registered roots with globs, existing projects, cross-root flag). An agent immediately knows what it gets back. It does not explicitly differentiate itself from the similarly-scoped sibling `resolve_cwd`, which keeps it just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is stated as a direct trigger: 'Use THIS tool whenever you need to find out whether — and with which cwd parameter — you can run a command in another project.' That is a clear when-to-use condition with a behavioral contrast ('runs no commands'). No alternative tool is named for the same need, so no exclusion clause exists.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_feedbackA
Read-onlyIdempotent

List feedback entries from .mcp/FEEDBACK.md (open entries; closed ones are moved to .mcp/FEEDBACK_ARCHIVE.md). Optionally filter by type, tool name, or status — pass archived:true to read the archive of closed entries instead. Use this to check existing feedback before creating new entries, or to review reported issues.

ParametersJSON Schema
NameRequiredDescriptionDefault
toolNoFilter by tool name
typeNoFilter by feedback type
statusNoFilter by status
archivedNotrue = list archived (closed) entries from .mcp/FEEDBACK_ARCHIVE.md instead of the active log

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the bar is lower. The description adds useful behavioral context beyond annotations: open entries live in FEEDBACK.md, closed entries are moved to FEEDBACK_ARCHIVE.md, and archived:true switches to reading the archive. This is meaningful behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The core action and file location are front-loaded, followed by filtering semantics and a usage hint. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with zero required parameters and fully documented schema, the description plus annotations cover everything an agent needs: what is listed, where it is read, how filtering works, and when to call it. No critical information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with descriptions for all four parameters and enums for type and status, so the baseline is 3. The description mostly restates the filtering options and archived behavior already present in the schema, adding minimal new semantic value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the specific action and resource: 'List feedback entries from .mcp/FEEDBACK.md'. It also clarifies the open-versus-archived split, which distinguishes this from the feedback creation/closure siblings. The intended purpose is immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use the tool: 'Use this to check existing feedback before creating new entries, or to review reported issues.' It implies the alternative workflow (creating entries via report_tool_feedback) but does not explicitly name alternatives or state when not to use it, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_toolsA
Read-onlyIdempotent

List all available MCP tools with descriptions. Use this to discover available tools before starting a task.

ParametersJSON Schema
NameRequiredDescriptionDefault
categoryNoFilter by category

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds only that results include descriptions; it says nothing about result size, ordering, or how the category filter affects output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the action stated first and the usage cue second. Nothing could be cut without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, annotation-covered read-only listing tool with no output schema, the description is nearly sufficient. It notes that tools come with descriptions, but omits any hint that results can be scoped by category, which is the only unsurfaced behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single enum parameter is documented in the schema, so the baseline of 3 applies. The description never mentions the category filter or what the enum values mean for results, adding no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (available MCP tools) with an added scoping detail (with descriptions), so an agent can tell it apart from most siblings. It does not explicitly distinguish itself from the adjacent help_tool, which is the one plausible point of confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use this to discover available tools before starting a task" gives a clear triggering context, i.e. an initial discovery step. There are no exclusions or explicit comparisons to alternatives such as help_tool, so it stops short of the top band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_log_sliceA
Read-onlyIdempotent

Reads a slice of a log file by the given line range. Use this tool instead of re-running the same command with a higher maxLines when you already have the log path saved from a previous run_safe_command / run_destructive_command response. NOTE: the file is read directly via Node.js (not through the shell), so it also works for logs in os.tmpdir().

ParametersJSON Schema
NameRequiredDescriptionDefault
logPathYesPath to the log file
lineCountNoNumber of lines to read (default: 100)
startLineNoStarting line (0-based, default: 0)

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent and non-destructive, so the safety profile is covered. The description adds genuinely useful behavior beyond that: the file is read directly via Node.js rather than through the shell, which implies it can reach logs in os.tmpdir() that shell-based reads might not. It stops short of describing error behavior for a missing path or whether lineCount is clamped.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, and the core action is front-loaded ahead of the routing advice and the shell-bypass note. Every clause carries actionable information for selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool whose annotations already carry the safety profile and which has no output schema to explain, the description covers purpose, routing and the shell-bypass nuance well. Minor gaps remain around failure modes and any size limits for large slices.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so logPath, lineCount and startLine are already documented with defaults and 0-based indexing. The description only echoes the 'line range' concept and adds no format or edge-case guidance beyond the schema, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Reads a slice of a log file') plus the scoping mechanism ('by the given line range'). It is readily distinguishable from the command-execution siblings because it names them only as the source of the log path, not as the same operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when to prefer this tool ('instead of re-running the same command with a higher maxLines') and under what precondition ('when you already have the log path saved from a previous run_safe_command / run_destructive_command response'). The named alternative and the selecting condition leave nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report_tool_feedbackA
Idempotent

Report a bug, improvement, or feature request about a tool of THIS server (see list_tools). Writes structured feedback to .mcp/FEEDBACK.md (project-specific, gitignored). Unknown tool names are rejected with a suggestion — set allowUnknownTool:true only when reporting a missing capability of this server. Use this when a tool produces unexpected results, crashes, or when you need a new capability. Entries are idempotent — duplicate reports are skipped.

ParametersJSON Schema
NameRequiredDescriptionDefault
toolYesName of the MCP tool this feedback is about (must be a tool of this server, e.g. 'batch_apply_edits')
typeYesType of feedback
titleYesShort summary (1 line)
expectedNoWhat you expected to happen
suggestionNoYour suggestion for a fix or improvement
descriptionYesDetailed description of the issue or request
reproductionNoSteps to reproduce the issue
allowUnknownToolNoFile feedback about a name that is not a tool of this server (default false) — use only for missing-capability reports

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses the write destination and scope (.mcp/FEEDBACK.md, project-specific, gitignored), the validation behavior (unknown tool names rejected with a suggestion), the allowUnknownTool semantics, and idempotency (duplicate reports skipped). This adds context beyond the annotations and is fully consistent with idempotentHint=true and readOnlyHint=false — no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences that are dense but each earns its place: purpose, destination, validation behavior, usage triggers, and idempotency. It is front-loaded with the purpose. Slightly long for a simple feedback tool, but appropriate given the complexity of the allowUnknownTool edge case.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with no output schema, the description covers purpose, when-to-use, validation rules, file destination, and idempotency. The only notable gap is that it doesn't describe what the call returns on success (confirmation, path, etc.), which is a minor omission for an agent invoking this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 8 parameters are already documented in the schema — the baseline is 3. The description adds meaningful context only for allowUnknownTool (explaining the missing-capability use case), which is a small but genuine contribution.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: 'Report a bug, improvement, or feature request about a tool of THIS server.' It clearly distinguishes itself from the feedback-management siblings (list_feedback, close_feedback) by being the creation/entry tool, and it references list_tools for discovering valid targets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit triggers: 'Use this when a tool produces unexpected results, crashes, or when you need a new capability.' It also defines when NOT to set allowUnknownTool (only for missing-capability reports). It could be stronger by naming list_feedback/close_feedback as the alternatives for viewing or closing feedback, but the guidance is otherwise clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_cwdA
Read-onlyIdempotent

Verifies whether the given path (or a friendly project name) is inside the allowed roots of this MCP server and returns the exact "cwd" to use for running commands. Use this tool when you need to find out whether — and with which "cwd" — you can work in a specific project (e.g. your workspace folder). On success: { ok: true, cwd, matchedRoot, name? }; on failure: { ok: false, error, roots }. Runs no commands — it only validates the configuration.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath or friendly project name to verify

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/closed-world, so safety is covered. The description adds real value beyond them by disclosing that it "runs no commands — it only validates configuration" and by spelling out the success ({ ok, cwd, matchedRoot, name? }) and failure ({ ok, error, roots }) payloads in the absence of an output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and each sentence contributes: use case, return shapes, and the no-side-effects guarantee. Slightly dense with escaped quotes, but well ordered and free of filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by documenting both success and failure shapes plus the fact that no commands are executed. For a single-parameter validator this is sufficient for an agent to call it correctly and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single path parameter, so the baseline is 3. The description's "or a friendly project name" adds a small nuance but essentially restates what the schema already says; no format, resolution order, or matching rules are provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource: verifies whether a path or friendly project name falls inside the server's allowed roots, and returns the exact cwd to use. An agent can distinguish this from siblings like list_allowed_roots or run_safe_command without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use this tool when you need to find out whether — and with which cwd — you can work in a specific project" gives a clear triggering context. However, it never names the obvious alternative (list_allowed_roots) or states any when-not condition, so routing between the two roots-related tools is still left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_command_grepA
Destructive

Runs a command and returns only lines matching the given pattern (case-insensitive regex). Use instead of run_safe_command when you know in advance that the output will be long and you only care about a specific pattern (e.g. searching for 'error' in build output, finding a specific test in test runner output). The command runs in the directory given by the cwd parameter and is subject to the same safety checks as run_safe_command. This is the replacement for Unix 'grep' on Windows — filtering happens in-process, so grep/head/tail are not needed. LIMITATION: The matches themselves can be too long — if you need more control, use run_safe_command first and then read_log_slice on the saved log file. Instead of Unix patterns like 'cmd /c ... | findstr ... & echo DONE' use THIS tool with a pattern — it filters in-process. If the task targets a project other than the primary one (MCP_PROJECT_ROOT), always pass the "cwd" parameter. Get the list of allowed roots via the "list_allowed_roots" tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoThe directory in which the command will be run (must be inside MCP_PROJECT_ROOT / MCP_EXTRA_ROOTS). Default: the primary project (MCP_PROJECT_ROOT).
commandYesCommand to execute
patternYesPattern (regular expression) to filter lines
timeoutMsNoTimeout in milliseconds (1,000 – 600,000, default 60,000). You can extend it for longer tests/builds, e.g. 180,000 for jest.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and openWorldHint=true, so the description's job is lighter. It adds real context: it is 'subject to the same safety checks as run_safe_command', filtering is in-process, the Windows grep replacement rationale, and an explicit LIMITATION about over-long matches. It does not spell out the destructive/command-execution risk itself, but the safety-check statement covers most of what an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and purposeful, but the Windows/grep and 'use THIS tool with a pattern' points are made across multiple sentences and overlap with the LIMITATION note, so a few sentences do not fully earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description implies the return is filtered matching lines. Combined with annotations covering safety and instructions covering cwd, timeouts, and alternatives, an agent has enough to call it correctly; only the exact return shape (e.g. whether line numbers/context are included) is unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds semantics the schema lacks: the pattern is a case-insensitive regex, and cwd scoping ties to MCP_PROJECT_ROOT/MCP_EXTRA_ROOTS with a pointer to list_allowed_roots. It also notes when to always pass cwd (non-primary projects).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Runs a command and returns only lines matching the given pattern') and immediately distinguishes itself from the sibling run_safe_command by scope. An agent can pick this without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it ('when you know in advance that the output will be long and you only care about a specific pattern'), names the alternative (run_safe_command) and the fallback (read_log_slice). Also tells the agent what NOT to do (Unix pipes/findstr).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_destructive_commandA
Destructive

Use ONLY when run_safe_command rejected the command AND the user explicitly confirmed in the chat that they want to run it despite the risk. NEVER set confirm:true automatically in reaction to a rejection from run_safe_command. First restate the risk to the user in your own words (exactly what the command will do and what it could break) and wait for their explicit 'yes' or 'I confirm' in the next message. If the user is not present in the conversation (e.g. an automated run without a human), do not use this tool at all. EXAMPLE: If the user says 'do it' for a general task and you then hit a dangerous rejection, that is not sufficient confirmation — you must explain the specific risk and get a new explicit confirmation. If the task targets a project other than the primary one (MCP_PROJECT_ROOT), always pass the "cwd" parameter. Get the list of allowed roots via the "list_allowed_roots" tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoThe directory in which the command will be run (must be inside MCP_PROJECT_ROOT / MCP_EXTRA_ROOTS). Default: the primary project (MCP_PROJECT_ROOT).
commandYesThe command to execute (runs in the directory given by the cwd parameter)
confirmNoConfirmation that you are aware of the risk (required for dangerous commands)
maxLinesNoMaximum number of output lines (default: 200)
timeoutMsNoTimeout in milliseconds (1,000 – 600,000, default 60,000). You can extend it for longer tests/builds, e.g. 180,000 for jest.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, non-idempotent, open-world, but the description goes well beyond them: it discloses the mandatory confirmation protocol, the no-auto-confirm rule, the headless/no-human exclusion, and the cwd requirement for non-primary projects. That is genuine behavioral context an agent needs to invoke this safely.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the strictest rule ('Use ONLY when...'), then qualification, then a concrete counterexample. Slightly long and the risk-restatement instruction is reiterated in the EXAMPLE, but nearly every sentence carries a distinct constraint.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, no-output-schema tool this covers everything an agent needs: precondition chain, confirmation workflow, headless behavior, and the cwd root-discovery pointer to list_allowed_roots. Timeout/output limits are already in the schema, so their absence from the description is fine.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: cwd must be set when the target project differs from MCP_PROJECT_ROOT, and confirm must never be set reflexively. The confirm parameter's social contract is only defined here, not in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description establishes that this tool executes a command that the safe path rejected, and explicitly names the sibling it is gated behind (run_safe_command). The verb 'run' is implied rather than stated outright, but the risk-restatement instruction ('exactly what the command will do and what it could break') makes the effect unambiguous. Distinguishable from all siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use (run_safe_command rejected AND user explicitly confirmed), when-not (no human present, general 'do it', automatic confirm:true), and the alternative (run_safe_command) is named. The negative case of insufficient confirmation is even exemplified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_safe_commandA
Destructive

Executes a safe command inside the project folder. This is the default command execution tool — use it whenever you are not sure whether a command is dangerous. Commands outside the project or that look dangerous are rejected automatically. If the tool returns isError:true with a rejection message (dangerous or directory_escape), DO NOT try to bypass it by rewriting the command or immediately switching to run_destructive_command without asking the user first. For dangerous operations use run_destructive_command with confirm:true. NOTE: 60s limit (optionally extend via timeoutMs) — not suitable for dev servers / watch mode. Run test suites (jest/npm test), typecheck and builds through this tool or run_command_grep, NOT through the built-in terminal. If the task targets a project other than the primary one (MCP_PROJECT_ROOT), always pass the "cwd" parameter. Get the list of allowed roots via the "list_allowed_roots" tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoThe directory in which the command will be run (must be inside MCP_PROJECT_ROOT / MCP_EXTRA_ROOTS). Default: the primary project (MCP_PROJECT_ROOT).
commandYesThe command to execute (runs in the directory given by the cwd parameter)
maxLinesNoMaximum number of output lines (default: 200)
timeoutMsNoTimeout in milliseconds (1,000 – 600,000, default 60,000). You can extend it for longer tests/builds, e.g. 180,000 for jest.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: it discloses automatic rejection of out-of-project or dangerous commands, the isError:true rejection messages (dangerous / directory_escape), an explicit anti-bypass policy, the 60s default timeout with timeoutMs extension, and the cwd requirement for non-primary projects. The destructiveHint:true annotation sits in slight tension with the 'safe' framing, but the description's boundary ('rejects dangerous') and escalation path resolve it rather than contradict it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and the safety boundary, and every sentence carries operational information. It is dense and heavy with emphatic capitalization (DO NOT, NOTE), which slightly hurts readability but not correctness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a command-execution tool with no output schema, it covers what an agent needs: rejection semantics, escalation path, timeout limits, working-directory rules, and how to discover allowed roots via list_allowed_roots. Nothing material is left to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: when to pass cwd (tasks targeting a project other than MCP_PROJECT_ROOT) and a concrete timeoutMs example (180,000 for jest). Only maxLines is left entirely to the schema, so it does not reach a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Executes a safe command inside the project folder') and immediately positions itself against the sibling run_destructive_command. An agent can tell the two apart without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives the default-use rule ('use it whenever you are not sure whether a command is dangerous'), names the alternative for dangerous operations (run_destructive_command with confirm:true), and states a when-not (not suitable for dev servers / watch mode), plus routing guidance for test suites and builds.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

split_file_by_declarationsA
Destructive

Split a large file into multiple smaller files based on top-level declarations. Optionally generates a combining file (mod.rs / index.ts / init.py). Use dryRun: true (default) to preview the layout before writing.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoWorking dir for resolving relative file paths (default: primary project root)
fileYesSource file to split
dryRunNoPreview only — write nothing (default: true)
groupingYesModule groupings
languageNoLanguage (auto-detected from extension)
overwriteNoAllow overwriting existing target files (default: false)
targetDirNoWhere new files are written (default: dirname of file)
generateIndexNoCreate combining file (default: true)

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered; the description adds real value beyond that by disclosing that dryRun defaults to true (preview-before-write) and that a combining file is optionally generated. It still omits explicit overwrite/destruction semantics, but the default-safe behavior is a meaningful addition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: purpose first, then the auxiliary output, then the safety-relevant default. No filler, and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter mutating tool with no output schema, the description plus fully-covered schema and destructiveness annotations give the agent enough to call it correctly. Minor gaps remain around failure/return behavior, but the schema and annotations carry the rest.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all 8 parameters (cwd, file, dryRun, grouping, language, overwrite, targetDir, generateIndex). The description only reinforces dryRun and the combining-file/index concept, adding nothing the schema doesn't already convey. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (Split) + resource (a large file) + mechanism (based on top-level declarations), with the auxiliary combining-file behavior named. It does not explicitly differentiate from potentially adjacent siblings such as generate_module_skeleton or batch_apply_edits, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case (splitting large files) and gives safety-oriented guidance for the dryRun flag, but never states when to prefer this over alternatives like generate_module_skeleton. Usage is only implied, not routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

universal_find_referencesA
Read-onlyIdempotent

Find all occurrences of a symbol across a workspace. Structured output with file, line, column, context. Optional language-aware mode (rust/typescript/python/cpp) adds role annotations: declaration, import, or usage. Use this tool BEFORE any refactoring session to understand what will break when a symbol is renamed or moved. When cwd is omitted, ALL registered roots are searched: nested roots are pruned and files are deduplicated by real path, so no match is listed or counted twice.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoWorkspace root to search (default: all registered roots — nested roots pruned, duplicates removed)
symbolYesSymbol to search for (word-boundary match)
languageNoOptional language-aware mode for role detection
contextLinesNoLines of context around each match (default: 1)
fileExtensionsNoRestrict to these extensions (default: common source extensions)
excludePatternsNoDirectories to skip (default: .git, node_modules, target, build, dist, __pycache__)

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral details beyond the annotations: it explains that nested roots are pruned and files deduplicated, so no match is counted twice, and that language-aware mode adds role annotations. These details are not present in the annotations and help the agent predict output. It also notes the structured output format, which is useful. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded: it starts with the primary action, then output format, then options, then usage guidance. Each sentence adds value without redundancy. It is well-organized and avoids unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for an agent to invoke the tool correctly: it covers the primary purpose, output structure, optional modes, default behaviors, and a clear use case. The schema and annotations provide the remaining parameter details and safety profile, so nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers all parameters with descriptions, so baseline is 3. The description adds extra meaning for 'cwd' by explaining the default (all registered roots) and deduplication behavior, and for 'language' by clarifying it adds role annotations. These enrich the schema without redundancy, so a 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Find') and resource ('all occurrences of a symbol across a workspace'), and mentions structured output. It clearly distinguishes the tool's core function, though it doesn't explicitly name sibling tools for differentiation. The mention of using it before refactoring provides context that helps an agent understand its role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs to use this tool before any refactoring session, which is a clear usage guideline. It also explains the default behavior when cwd is omitted, which informs the agent about scope. However, it doesn't mention alternatives or conditions when not to use it, but the primary use case is well-defined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_refactor_safetyA
Read-onlyIdempotent

Semantic diff between old and new code. Catches accidental deletions before compilation. Checks: function count, signatures, export count, imports, comment ratio. Intentionally conservative — renames appear as errors requiring explicit confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
afterYesNew code text
beforeYesOriginal code text
languageNoLanguage (auto-detected from content)

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds genuinely non-obvious behavior beyond that: it is deliberately conservative, and renames will surface as errors requiring explicit confirmation. That false-positive profile materially affects how an agent interprets results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: purpose, motivation, checks performed, and the conservative-bias caveat. Front-loaded with the core definition and free of filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the checks it reports (function count, signatures, exports, imports, comment ratio) and the rename-confirmation behavior. It stops short of describing the result shape or whether mismatches block anything, but an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so before/after/language are already documented, including language auto-detection and the enum. The description adds no syntax or format detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: a semantic diff between old and new code, plus the concrete checks it performs. This is clearly distinguishable from every sibling tool, none of which do source-diff analysis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Catches accidental deletions before compilation" gives a clear usage context (pre-build verification after a refactor). It does not name an alternative tool or state when not to use it, but for this toolkit there is no obvious competing option.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.2.1
    • Changedbatch_apply_edits1 field changed
      • addedInput schema / properties / edits / items / properties / excludePatterns
        Added value: +{
        +  "description": "Patterns to exclude from replaceAll (e.g. '#[cfg(test)]' to skip test modules)",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
    • Changedlist_feedback1 field changed
      • addedInput schema / properties / archived
        Added value: +{
        +  "description": "true = list archived (closed) entries from .mcp/FEEDBACK_ARCHIVE.md instead of the active log",
        +  "type": "boolean"
        +}
    • Changedreport_tool_feedback2 fields changed
      • addedInput schema / properties / allowUnknownTool
        Added value: +{
        +  "description": "File feedback about a name that is not a tool of this server (default false) — use only for missing-capability reports",
        +  "type": "boolean"
        +}
      • changedInput schema / properties / tool / description
        Previous value: -"Name of the MCP tool this feedback is about"New value: +"Name of the MCP tool this feedback is about (must be a tool of this server, e.g. 'batch_apply_edits')"
    • Changeduniversal_find_references1 field changed
      • changedInput schema / properties / cwd / description
        Previous value: -"Workspace root to search (default: primary project root)"New value: +"Workspace root to search (default: all registered roots — nested roots pruned, duplicates removed)"
  2. 1 tool updatev0.1.1
    • Changedbatch_apply_edits1 field changed
      • addedInput schema / properties / cwd
        Added value: +{
        +  "description": "Working dir for resolving relative file paths (default: primary project root)",
        +  "type": "string"
        +}
  3. 17 tool updatesv0.1.0
    • First observedbatch_apply_edits
    • First observedclose_feedback
    • First observedextract_code_block
    • First observedgenerate_module_skeleton
    • First observedhelp_tool
    • First observedlist_allowed_roots
    • First observedlist_feedback
    • First observedlist_tools
    • First observedread_log_slice
    • First observedreport_tool_feedback
    • First observedresolve_cwd
    • First observedrun_command_grep
    • First observedrun_destructive_command
    • First observedrun_safe_command
    • First observedsplit_file_by_declarations
    • First observeduniversal_find_references
    • First observedverify_refactor_safety

TDQS

A4/5.0

Scored across 17 tools

Disambiguation4/5

Most tools have clearly separated jobs (command execution vs code extraction vs refactoring vs feedback), and paired tools like run_safe_command/run_destructive_command are explicitly distinguished. The main overlap is between list_allowed_roots and resolve_cwd, which both address allowed-root/cwd discovery, though their descriptions do enough to mostly keep them apart.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun pattern, with predictable families like run_*, list_*, and *feedback. There is no style mixing or vague generic naming, and the command tools (run_safe_command, run_destructive_command, run_command_grep) form a particularly coherent set.

Tool Count4/5

17 tools is slightly above the ideal 3-15 range, but the server covers several distinct subdomains: command execution, code analysis, refactoring, and feedback management. The count feels a bit heavy due to near-redundant helpers like list_allowed_roots and resolve_cwd, but most tools have a clear purpose.

Completeness4/5

The tool surface covers major workflows well: command execution (safe/destructive/grep/log reading), code analysis (find references/extract), refactoring (split/generate/batch edits/verify), and a full feedback lifecycle. Minor gaps exist, such as no direct arbitrary-file read/write or explicit symbol rename tool, but these can be worked around with existing tools.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides secure, sandboxed file system access for AI assistants to read, write, and manage project files with controlled command execution capabilities, all confined to a designated workspace directory.
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables safe execution of terminal commands across different shells (bash, cmd, PowerShell) with configurable timeouts, working directories, and resource limits for command-line operations through AI assistants.
    -
  • A
    license
    A
    quality
    F
    maintenance
    Enables AI assistants to execute terminal commands on a host machine with configurable, granular permission controls and safety protections. It features multiple security modes, including allowlists and manual approval, to ensure safe command execution within specified directories.
    6
    Apache 2.0