Skip to main content
Glama
mlzoo

amazon-quick-exec-mcp

by mlzoo

amazon-quick-exec-mcp — a real shell for Amazon Quick

中文

Amazon Quick's built-in agent tools (run_python, ripgrep, and friends) execute inside quickwork-sandbox: they cannot reach localhost and cannot see most of your filesystem. Local MCP servers are different — Quick launches them as ordinary child processes of the desktop app, outside that sandbox.

So this server runs commands on the real machine, as you. Sync execution, background jobs, process-tree cleanup, an audit log, and a narrow guard against the handful of commands nobody means to run.

Commands go through a login shell (zsh -lc), so PATH matches Terminal. That matters more than it sounds: anything installed under ~/.local/bin, ~/.toolbox/bin, or by a version manager only exists once the profile is sourced.

Install

git clone <this repo> && cd quick-exec-mcp

uv sync                                # create the venv
./.venv/bin/python test_exec_mcp.py    # 70 checks over real MCP stdio
python install.py install              # register with Quick

Then in Quick: Settings → Capabilities → Connectors → Local Shell (exec) → Refresh. Ask it call shell_info as a first message — it reports the user, host and PATH your commands actually get, which is how you confirm you are outside the sandbox.

Requires Python 3.10+ and uv. install.py is stdlib-only, so it runs with any python3.

python install.py status       what Quick currently has registered
python install.py install      add this server, enabled
python install.py disable      keep the entry, turn it off
python install.py enable       turn it back on
python install.py uninstall    remove the entry

It edits the active profile's mcp_config.json — the same file the Connectors UI writes — and backs it up to mcp_config.json.bak before every write. A reinstall preserves secret:// references and any env you set by hand. See docs/how-quick-loads-mcp.md for the file layout, the schema, and the gotchas that cost the most time.

Related MCP server: cli-mcp

Tools

Tool

What it does

shell_execute

Run and wait. Returns exit_code, stdout, stderr, duration_s, timed_out

shell_start

Run in the background, returns a job_id. For builds, installs, dev servers

shell_job_output

Snapshot of a job's output so far. Safe to call repeatedly

shell_job_wait

Block until a job exits, then return its output. Beats polling

shell_job_kill

Kill a job and its whole process tree

shell_list_jobs

Every job started this session, running and finished

shell_info

Configuration plus a live whoami / hostname / PATH probe

shell_execute takes cwd (defaults to $HOME, ~ expands), timeout (default 120s), stdin, and env (merged over the inherited environment).

Configuration

Set these in the env block of the exec entry in Quick's mcp_config.json.

Variable

Default

Meaning

QUICK_EXEC_SHELL

/bin/zsh

Which shell to use

QUICK_EXEC_SHELL_ARGS

-lc

-l sources the profile; drop it to skip that

QUICK_EXEC_DEFAULT_CWD

$HOME

Working directory when cwd is omitted

QUICK_EXEC_DEFAULT_TIMEOUT

120

Default timeout, seconds

QUICK_EXEC_MAX_TIMEOUT

3600

Ceiling on the timeout argument

QUICK_EXEC_MAX_BYTES

60000

Per-stream output cap, roughly 15k tokens

QUICK_EXEC_READER_GRACE

2

Seconds to keep draining the pipes after the child exits

QUICK_EXEC_MAX_FINISHED_JOBS

50

Finished jobs retained before the oldest are dropped

QUICK_EXEC_AUDIT_LOG

~/.quick-exec-mcp/audit.jsonl

Audit log path

QUICK_EXEC_LOG_LEVEL

WARNING

INFO logs every call to stderr

QUICK_EXEC_ALLOW_DANGEROUS

unset

1 turns the command guard off

Audit log

Every command appends one JSON line to ~/.quick-exec-mcp/audit.jsonl — time, command, cwd, exit code, duration — and refusals are logged too. This is how you find out what Quick actually ran on your machine.

tail -20 ~/.quick-exec-mcp/audit.jsonl | jq -c '[.ts, .kind, .exit_code, .command]'

Command guard

Refused by default (QUICK_EXEC_ALLOW_DANGEROUS=1 disables it):

  • recursive delete of /, $HOME, /System, /Applications, /usr, and peers

  • mkfs*, diskutil eraseDisk/reformat/partitionDisk, dd of=/dev/disk*

  • fork bombs

  • shutdown / reboot / halt in command position

  • csrutil disable, spctl --master-disable (turning off SIP / Gatekeeper)

Patterns are matched against the command both verbatim and with quotes stripped, because rm -rf "$HOME" is how a model writes that as often as not. They stay deliberately narrow, so ordinary work goes through: rm -rf on a specific directory is allowed, rm -rf "$HOME/project" is allowed, and so is grep shutdown /etc/hosts — shutdown there is an argument, not a command.

This guards against a model slipping, not against malice. It is pattern matching, so any indirection ($(echo rm) -rf /) defeats it, and anyone who can reach this server can already run anything you can.

Behaviour worth knowing

  • Timeouts kill the whole process tree. Children run in their own process group (start_new_session=True); on timeout the group gets SIGTERM then SIGKILL, so a timed-out npm install leaves no orphans. Output produced before the timeout is kept, with timed_out: true.

  • timeout is a real bound, even against a process that escapes. Anything that calls setsid — a daemon, mostly — leaves the process group, survives the kill, and keeps the inherited stdout pipe open. Waiting for end-of-file on that pipe would hang the call forever and make the connector look dead, so the drain gets its own QUICK_EXEC_READER_GRACE budget and the reply comes back with output_incomplete: true. Expect roughly timeout + 5s in that case: SIGTERM, the escalation to SIGKILL, then the grace period.

  • timeout: 0 or a negative value means "use the default", not "give up immediately".

  • Truncation keeps the head and the tail. Over the cap, the first third and last two thirds survive with a note about how many bytes went missing in between. A failing build puts the useful part at the end; keeping only the head would discard exactly what you need.

  • stdout and stderr come back separately, never interleaved.

  • Interactive pagers are disabled via PAGER=cat, GIT_PAGER=cat, TERM=dumb, NO_COLOR=1, so git log cannot hang in less or return a wall of ANSI escapes. Override any of them through env.

Tests

./.venv/bin/python test_exec_mcp.py

70 checks driving the server over real MCP stdio — initialize, tools/list, tools/call, the whole JSON-RPC exchange Quick makes — rather than importing the module and calling functions. Tools that look fine in Python and break over the wire are exactly the failure this catches. Covered: handshake, tool schemas, login-shell PATH (compared against an actual zsh -lc), stream separation, stdin, env overrides, invalid UTF-8, timeout with orphan detection via pgrep, an orphan that escapes the process group entirely, head-and-tail truncation, background job lifecycle, job pruning, and the audit log.

The destructive guard cases — rm -rf "$HOME" and friends — are asserted against _guard() in-process, so they are never handed to a shell. A suite that sends them to a live server and trusts a regex to stop them is one bad regex away from deleting the home directory of whoever ran it. The one case that does go over the wire runs with HOME pointed at a throwaway directory and then asserts that directory is still intact, so a future regression destroys a temp dir instead.

Security

This gives Quick full shell access as your user — read and write any file you can, read credential files, call the AWS CLI, push code, make network requests. Bypassing the sandbox is the point of it, not a bug.

  • Leave Quick's global_default tool permission on prompt so calls are confirmed before they run.

  • Read ~/.quick-exec-mcp/audit.jsonl now and then.

  • Do not register this on a machine or client you do not trust.

  • Keep credentials out of the config file; Quick's secret store (secret://) exists for that, and a reinstall preserves those references.

License

MIT — see LICENSE.

Available Tools

7 tools
shell_executeA

Run a shell command on this Mac and wait for it to finish.

Runs through a login shell, so pipes, globs, redirects, &&, $VARS, heredocs and shell functions all work — pass the command exactly as you would type it in Terminal. Use this for anything that finishes within the timeout.

For work that outlives the timeout (builds, installs, dev servers) use shell_start instead.

Args: command: Shell command line, e.g. "ls -la ~/Documents | head -20". cwd: Working directory. Defaults to the user's home. ~ is expanded. timeout: Seconds before the process tree is killed. Default 120, max 3600. Zero or negative means "use the default", not "give up immediately". stdin: Text piped to the command's stdin. env: Extra environment variables, merged over the inherited environment.

Returns: exit_code, stdout, stderr, duration_s, cwd, timed_out, and stdout_truncated / stderr_truncated flags.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNo
envNo
stdinNo
commandYes
timeoutNo

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden well: it discloses that execution goes through a login shell, that timeout kills the whole process tree, that timeout <= 0 falls back to the default, and that stdout/stderr may be truncated. It omits the security/permission posture of executing arbitrary commands and any output-size limits, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core behavior and the sibling routing before any detail, then cleanly separated Args/Returns sections. Slightly longer than necessary — 'wait for it to finish' is restated by the timeout sentence — but every block earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description supplies the return shape (exit_code, stdout, stderr, duration_s, cwd, timed_out, truncation flags), so an agent knows what to expect. Missing only side-context such as permission requirements and truncation limits for an arbitrary-code-execution tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and it does: it documents all five parameters, including the non-obvious timeout edge case (zero/negative means default, not immediate abort), ~ expansion for cwd, and env merge-over-inherited semantics. This adds meaning well beyond the bare schema types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Run a shell command on this Mac and wait for it to finish') and immediately contrasts its blocking behavior with the non-blocking sibling shell_start. An agent can distinguish it from all six siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly scopes usage: 'Use this for anything that finishes within the timeout' and names the alternative ('For work that outlives the timeout (builds, installs, dev servers) use shell_start instead'). That is a clear when/when-not with the routing alternative called out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shell_infoA

Report how this server is configured and what environment commands see.

Worth calling first in a session: it proves the server is running on the real host as the real user rather than inside Quick's sandbox.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses what a call proves (real host vs sandbox), but says nothing about what configuration fields are returned or whether the output format is stable. Adequate for a read-only info tool, but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core purpose followed by the practical rationale for calling it. No filler, though the line breaks split a single thought mid-sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only diagnostic tool with no output schema and no annotations, the description covers purpose, timing, and the meaning of a successful call. Only the shape of the returned configuration data is left unspecified, which is minor here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Report') and resource ('how this server is configured and what environment commands see'), which is clearly distinct from the execution-oriented siblings like shell_execute and shell_start. It does not name a sibling explicitly, but the informational nature is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly recommends calling it first in a session and gives the reason (verifying the server runs on the real host as the real user, not in Quick's sandbox). It lacks an explicit 'when not to use' or named alternative, but the timing guidance is concrete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shell_job_killA

Kill a background job and its whole process tree.

Args: job_id: Id returned by shell_start.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that the whole process tree is killed (not just the parent), which is important destructive behavior. It does not state error behavior for an unknown/finished job_id, permissions required, or whether the kill is graceful or forced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence states the action and scope, followed by a compact args block. Nothing is wasted, though the args line largely restates the single parameter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, no-output-schema tool, the description covers the essential action, scope, and where the id comes from. Missing details are minor (return value, error cases), so it is complete enough for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the only schema detail is a title 'Job Id'. The description compensates by explaining the id's provenance ('Id returned by shell_start'), which is genuinely useful for obtaining the value. It adds meaning beyond the schema but not full syntax or validation details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Kill') and resource ('a background job and its whole process tree'), with the scope of the kill made explicit. It doesn't explicitly differentiate itself from siblings like shell_job_wait or shell_job_output, but the destructive action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The note that job_id is 'Id returned by shell_start' implies the lifecycle context (start then kill), which is a form of usage guidance. However, there is no explicit when-to-use or when-not-to-use guidance against alternatives like shell_job_wait or shell_job_output.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shell_job_outputA

Read the output captured so far from a background job.

Safe to call repeatedly while the job runs — it returns a snapshot and does not consume the buffer.

Args: job_id: Id returned by shell_start. tail_lines: Return only the last N lines of each stream. 0 means all.

Returns: running, exit_code (null while running), stdout, stderr, elapsed_s.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes
tail_linesNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that the call is idempotent/non-destructive ('does not consume the buffer'), that it is safe to call repeatedly, and that exit_code is null while running. It omits permissions/auth and any rate-limit or buffer-eviction caveats, keeping it just short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line purpose followed by Args/Returns blocks; every section earns its place. Slightly formulaic but no wasted sentences, and the behavioral caveat appears before the parameter detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the Returns block (running, exit_code, stdout, stderr, elapsed_s) is genuinely needed and supplied. Combined with the non-consuming guarantee, an agent has enough to poll and interpret results; only cross-tool routing to shell_job_wait is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and largely does: job_id is explained as the id returned by shell_start, and tail_lines is explained as 'last N lines of each stream, 0 means all.' It fails to mention the schema's default of 200, so an agent cannot fully predict default behavior, but the added semantic meaning is substantial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Read the output captured so far from a background job') and clarifies the snapshot semantics. It names shell_start as the source of job_id, which implicitly differentiates it from siblings, but never explicitly contrasts it with shell_job_wait, the closest alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Safe to call repeatedly while the job runs' implies a polling usage pattern, and the non-consuming behavior hints at when this is preferred. However, it never explicitly says when to use this versus shell_job_wait (blocking) or how to know the job is finished beyond the returned 'running' field. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shell_job_waitA

Block until a background job exits, then return its output.

Cheaper and more reliable than polling shell_job_output in a loop. If the timeout expires the job keeps running and running comes back true — call again or kill it.

Args: job_id: Id returned by shell_start. timeout: Max seconds to wait. The job is NOT killed on timeout. tail_lines: Return only the last N lines of each stream. 0 means all.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes
timeoutNo
tail_linesNo

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and mostly succeeds: it discloses that the call blocks, that timeout does NOT kill the job, and the `running` flag on expiry. It is softer on return shape beyond 'output' and any auth/concurrency constraints, but the important non-obvious semantics are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose sentence followed by a compact args block; every line carries information, no filler. Easily scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema or annotations, but the description covers the essential return semantics (output plus the `running` flag) and the timeout contract. Slightly thin on exact return format, but sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and does: it documents all three params with meaning beyond the schema, notably that timeout does not kill the job and that tail_lines=0 means all lines, plus the provenance of job_id from shell_start.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource with scope: block until a background job exits then return its output. It explicitly distinguishes itself from the sibling shell_job_output by contrasting blocking versus polling, so an agent can select it without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (polling shell_job_output in a loop) and states the advantage, then gives the recovery path on timeout: the job keeps running, `running` returns true, call again or kill it. When-to-use, tradeoff, and follow-up action are all explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shell_list_jobsB

List background jobs started this session, running and finished.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose meaningful behavior: results are session-scoped and include both running and finished jobs. However, it omits that this is a non-destructive read operation, whether job IDs from other sessions are visible, or any ordering/limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the scope qualifier ('started this session') and state filter ('running and finished') both earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool with no output schema, the description covers scope and state but says nothing about the returned shape—notably whether job identifiers (needed by sibling tools like shell_job_output and shell_job_kill) are included. That gap matters because this tool is the natural entry point to the job-management workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to document; baseline 4 applies. No ambiguity about inputs is possible.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (background jobs) with explicit scope: jobs started this session, both running and finished. This clearly distinguishes it from shell_job_output/shell_job_wait/shell_job_kill, which imply a known job ID, though it doesn't name those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is a discovery step (to find jobs before using output/wait/kill), but never states when to use it versus the other six shell tools, nor any prerequisites. No explicit when/when-not guidance is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shell_startA

Start a long-running command in the background and return a job id.

Output is buffered as it arrives; read it with shell_job_output, block on it with shell_job_wait, stop it with shell_job_kill. Use this for builds, installs, test suites and dev servers that would blow shell_execute's timeout.

Args: command: Shell command line. cwd: Working directory. Defaults to the user's home. env: Extra environment variables.

Returns: job_id, pid, command, cwd.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNo
envNo
commandYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does well: it discloses background execution, output buffering as it arrives, and that the operation is non-blocking with a returned job id. It does not cover environment/permission caveats or what happens if the command exits immediately, but the core behavioral model is clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose, then the buffering/lifecycle sentence, then compact Args and Returns blocks. Every line is useful; the Args/Returns sections restate the schema property names but earn their place given the 0% schema coverage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a job-launching tool with no output schema, the description provides the return fields (job_id, pid, command, cwd), the background/buffering semantics, and the lifecycle routing to siblings. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: all three parameters are documented (command, cwd with its home-directory default, env as extra variables). Slightly less than a 5 because per-parameter types/format nuances (e.g., env key/value semantics) are not spelled out.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: starts a long-running command in the background and returns a job id. The name and sibling set (shell_execute, shell_job_*) make the async/background distinction explicit, so an agent can tell it apart without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it (builds, installs, test suites, dev servers that would blow shell_execute's timeout) and names the follow-up tools for the full lifecycle: shell_job_output for reading, shell_job_wait for blocking, shell_job_kill for stopping. Both the trigger condition and the alternative are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv1.0.0
    • First observedshell_execute
    • First observedshell_info
    • First observedshell_job_kill
    • First observedshell_job_output
    • First observedshell_job_wait
    • First observedshell_list_jobs
    • First observedshell_start

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Provides secure execution of terminal commands (PowerShell, CMD, shell) with configurable security policies including command blocking, path restrictions, and timeout.
    1
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables safe shell command execution with configurable directory and command restrictions, allowing Claude Desktop to run shell commands securely.
    MIT
  • F
    license
    A
    quality
    D
    maintenance
    Enables secure file system operations (read, write, delete) and simulated command execution with server-enforced permission policies, risk assessment, and human-in-the-loop approval.
    5
    -