amazon-quick-exec-mcp
Integrates with Amazon Quick to provide a real local shell outside its normal sandbox. It exposes tools for synchronous and background command execution, streaming/job monitoring, waiting on or killing entire process trees, inspecting runtime identity/config/paths, passing stdin/env/cwd/timeouts, truncating large outputs safely, and writing an audit log.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@amazon-quick-exec-mcprun whoami and hostname to confirm I'm outside the sandbox"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
amazon-quick-exec-mcp — a real shell for Amazon Quick
Amazon Quick's built-in agent tools (run_python, ripgrep, and friends) execute
inside quickwork-sandbox: they cannot reach localhost and cannot see most of
your filesystem. Local MCP servers are different — Quick launches them as ordinary
child processes of the desktop app, outside that sandbox.
So this server runs commands on the real machine, as you. Sync execution, background jobs, process-tree cleanup, an audit log, and a narrow guard against the handful of commands nobody means to run.
Commands go through a login shell (zsh -lc), so PATH matches Terminal. That
matters more than it sounds: anything installed under ~/.local/bin,
~/.toolbox/bin, or by a version manager only exists once the profile is sourced.
Install
git clone <this repo> && cd quick-exec-mcp
uv sync # create the venv
./.venv/bin/python test_exec_mcp.py # 70 checks over real MCP stdio
python install.py install # register with QuickThen in Quick: Settings → Capabilities → Connectors → Local Shell (exec) →
Refresh. Ask it call shell_info as a first message — it reports the user, host
and PATH your commands actually get, which is how you confirm you are outside
the sandbox.
Requires Python 3.10+ and uv. install.py is
stdlib-only, so it runs with any python3.
python install.py status what Quick currently has registered
python install.py install add this server, enabled
python install.py disable keep the entry, turn it off
python install.py enable turn it back on
python install.py uninstall remove the entryIt edits the active profile's mcp_config.json — the same file the Connectors UI
writes — and backs it up to mcp_config.json.bak before every write. A reinstall
preserves secret:// references and any env you set by hand. See
docs/how-quick-loads-mcp.md for the file layout,
the schema, and the gotchas that cost the most time.
Related MCP server: cli-mcp
Tools
Tool | What it does |
| Run and wait. Returns |
| Run in the background, returns a |
| Snapshot of a job's output so far. Safe to call repeatedly |
| Block until a job exits, then return its output. Beats polling |
| Kill a job and its whole process tree |
| Every job started this session, running and finished |
| Configuration plus a live |
shell_execute takes cwd (defaults to $HOME, ~ expands), timeout
(default 120s), stdin, and env (merged over the inherited environment).
Configuration
Set these in the env block of the exec entry in Quick's mcp_config.json.
Variable | Default | Meaning |
|
| Which shell to use |
|
|
|
|
| Working directory when |
|
| Default timeout, seconds |
|
| Ceiling on the |
|
| Per-stream output cap, roughly 15k tokens |
|
| Seconds to keep draining the pipes after the child exits |
|
| Finished jobs retained before the oldest are dropped |
|
| Audit log path |
|
|
|
| unset |
|
Audit log
Every command appends one JSON line to ~/.quick-exec-mcp/audit.jsonl — time,
command, cwd, exit code, duration — and refusals are logged too. This is how you
find out what Quick actually ran on your machine.
tail -20 ~/.quick-exec-mcp/audit.jsonl | jq -c '[.ts, .kind, .exit_code, .command]'Command guard
Refused by default (QUICK_EXEC_ALLOW_DANGEROUS=1 disables it):
recursive delete of
/,$HOME,/System,/Applications,/usr, and peersmkfs*,diskutil eraseDisk/reformat/partitionDisk,dd of=/dev/disk*fork bombs
shutdown/reboot/haltin command positioncsrutil disable,spctl --master-disable(turning off SIP / Gatekeeper)
Patterns are matched against the command both verbatim and with quotes stripped,
because rm -rf "$HOME" is how a model writes that as often as not. They stay
deliberately narrow, so ordinary work goes through: rm -rf on a specific
directory is allowed, rm -rf "$HOME/project" is allowed, and so is
grep shutdown /etc/hosts — shutdown there is an argument, not a command.
This guards against a model slipping, not against malice. It is pattern
matching, so any indirection ($(echo rm) -rf /) defeats it, and anyone who can
reach this server can already run anything you can.
Behaviour worth knowing
Timeouts kill the whole process tree. Children run in their own process group (
start_new_session=True); on timeout the group gets SIGTERM then SIGKILL, so a timed-outnpm installleaves no orphans. Output produced before the timeout is kept, withtimed_out: true.timeoutis a real bound, even against a process that escapes. Anything that callssetsid— a daemon, mostly — leaves the process group, survives the kill, and keeps the inherited stdout pipe open. Waiting for end-of-file on that pipe would hang the call forever and make the connector look dead, so the drain gets its ownQUICK_EXEC_READER_GRACEbudget and the reply comes back withoutput_incomplete: true. Expect roughlytimeout + 5sin that case: SIGTERM, the escalation to SIGKILL, then the grace period.timeout: 0or a negative value means "use the default", not "give up immediately".Truncation keeps the head and the tail. Over the cap, the first third and last two thirds survive with a note about how many bytes went missing in between. A failing build puts the useful part at the end; keeping only the head would discard exactly what you need.
stdout and stderr come back separately, never interleaved.
Interactive pagers are disabled via
PAGER=cat,GIT_PAGER=cat,TERM=dumb,NO_COLOR=1, sogit logcannot hang inlessor return a wall of ANSI escapes. Override any of them throughenv.
Tests
./.venv/bin/python test_exec_mcp.py70 checks driving the server over real MCP stdio — initialize,
tools/list, tools/call, the whole JSON-RPC exchange Quick makes — rather than
importing the module and calling functions. Tools that look fine in Python and
break over the wire are exactly the failure this catches. Covered: handshake, tool
schemas, login-shell PATH (compared against an actual zsh -lc), stream
separation, stdin, env overrides, invalid UTF-8, timeout with orphan detection via
pgrep, an orphan that escapes the process group entirely, head-and-tail
truncation, background job lifecycle, job pruning, and the audit log.
The destructive guard cases — rm -rf "$HOME" and friends — are asserted against
_guard() in-process, so they are never handed to a shell. A suite that sends
them to a live server and trusts a regex to stop them is one bad regex away from
deleting the home directory of whoever ran it. The one case that does go over the
wire runs with HOME pointed at a throwaway directory and then asserts that
directory is still intact, so a future regression destroys a temp dir instead.
Security
This gives Quick full shell access as your user — read and write any file you can, read credential files, call the AWS CLI, push code, make network requests. Bypassing the sandbox is the point of it, not a bug.
Leave Quick's
global_defaulttool permission onpromptso calls are confirmed before they run.Read
~/.quick-exec-mcp/audit.jsonlnow and then.Do not register this on a machine or client you do not trust.
Keep credentials out of the config file; Quick's secret store (
secret://) exists for that, and a reinstall preserves those references.
License
MIT — see LICENSE.
Available Tools
7 toolsshell_executeA
Run a shell command on this Mac and wait for it to finish.
Runs through a login shell, so pipes, globs, redirects, &&, $VARS, heredocs and shell functions all work — pass the command exactly as you would type it in Terminal. Use this for anything that finishes within the timeout.
For work that outlives the timeout (builds, installs, dev servers) use shell_start instead.
Args: command: Shell command line, e.g. "ls -la ~/Documents | head -20". cwd: Working directory. Defaults to the user's home. ~ is expanded. timeout: Seconds before the process tree is killed. Default 120, max 3600. Zero or negative means "use the default", not "give up immediately". stdin: Text piped to the command's stdin. env: Extra environment variables, merged over the inherited environment.
Returns: exit_code, stdout, stderr, duration_s, cwd, timed_out, and stdout_truncated / stderr_truncated flags.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | ||
| env | No | ||
| stdin | No | ||
| command | Yes | ||
| timeout | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden well: it discloses that execution goes through a login shell, that timeout kills the whole process tree, that timeout <= 0 falls back to the default, and that stdout/stderr may be truncated. It omits the security/permission posture of executing arbitrary commands and any output-size limits, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core behavior and the sibling routing before any detail, then cleanly separated Args/Returns sections. Slightly longer than necessary — 'wait for it to finish' is restated by the timeout sentence — but every block earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description supplies the return shape (exit_code, stdout, stderr, duration_s, cwd, timed_out, truncation flags), so an agent knows what to expect. Missing only side-context such as permission requirements and truncation limits for an arbitrary-code-execution tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and it does: it documents all five parameters, including the non-obvious timeout edge case (zero/negative means default, not immediate abort), ~ expansion for cwd, and env merge-over-inherited semantics. This adds meaning well beyond the bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Run a shell command on this Mac and wait for it to finish') and immediately contrasts its blocking behavior with the non-blocking sibling shell_start. An agent can distinguish it from all six siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scopes usage: 'Use this for anything that finishes within the timeout' and names the alternative ('For work that outlives the timeout (builds, installs, dev servers) use shell_start instead'). That is a clear when/when-not with the routing alternative called out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_infoA
Report how this server is configured and what environment commands see.
Worth calling first in a session: it proves the server is running on the real host as the real user rather than inside Quick's sandbox.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses what a call proves (real host vs sandbox), but says nothing about what configuration fields are returned or whether the output format is stable. Adequate for a read-only info tool, but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core purpose followed by the practical rationale for calling it. No filler, though the line breaks split a single thought mid-sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only diagnostic tool with no output schema and no annotations, the description covers purpose, timing, and the meaning of a successful call. Only the shape of the returned configuration data is left unspecified, which is minor here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Report') and resource ('how this server is configured and what environment commands see'), which is clearly distinct from the execution-oriented siblings like shell_execute and shell_start. It does not name a sibling explicitly, but the informational nature is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly recommends calling it first in a session and gives the reason (verifying the server runs on the real host as the real user, not in Quick's sandbox). It lacks an explicit 'when not to use' or named alternative, but the timing guidance is concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_job_killA
Kill a background job and its whole process tree.
Args: job_id: Id returned by shell_start.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that the whole process tree is killed (not just the parent), which is important destructive behavior. It does not state error behavior for an unknown/finished job_id, permissions required, or whether the kill is graceful or forced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence states the action and scope, followed by a compact args block. Nothing is wasted, though the args line largely restates the single parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, no-output-schema tool, the description covers the essential action, scope, and where the id comes from. Missing details are minor (return value, error cases), so it is complete enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the only schema detail is a title 'Job Id'. The description compensates by explaining the id's provenance ('Id returned by shell_start'), which is genuinely useful for obtaining the value. It adds meaning beyond the schema but not full syntax or validation details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Kill') and resource ('a background job and its whole process tree'), with the scope of the kill made explicit. It doesn't explicitly differentiate itself from siblings like shell_job_wait or shell_job_output, but the destructive action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The note that job_id is 'Id returned by shell_start' implies the lifecycle context (start then kill), which is a form of usage guidance. However, there is no explicit when-to-use or when-not-to-use guidance against alternatives like shell_job_wait or shell_job_output.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_job_outputA
Read the output captured so far from a background job.
Safe to call repeatedly while the job runs — it returns a snapshot and does not consume the buffer.
Args: job_id: Id returned by shell_start. tail_lines: Return only the last N lines of each stream. 0 means all.
Returns: running, exit_code (null while running), stdout, stderr, elapsed_s.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| tail_lines | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that the call is idempotent/non-destructive ('does not consume the buffer'), that it is safe to call repeatedly, and that exit_code is null while running. It omits permissions/auth and any rate-limit or buffer-eviction caveats, keeping it just short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line purpose followed by Args/Returns blocks; every section earns its place. Slightly formulaic but no wasted sentences, and the behavioral caveat appears before the parameter detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the Returns block (running, exit_code, stdout, stderr, elapsed_s) is genuinely needed and supplied. Combined with the non-consuming guarantee, an agent has enough to poll and interpret results; only cross-tool routing to shell_job_wait is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate and largely does: job_id is explained as the id returned by shell_start, and tail_lines is explained as 'last N lines of each stream, 0 means all.' It fails to mention the schema's default of 200, so an agent cannot fully predict default behavior, but the added semantic meaning is substantial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Read the output captured so far from a background job') and clarifies the snapshot semantics. It names shell_start as the source of job_id, which implicitly differentiates it from siblings, but never explicitly contrasts it with shell_job_wait, the closest alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Safe to call repeatedly while the job runs' implies a polling usage pattern, and the non-consuming behavior hints at when this is preferred. However, it never explicitly says when to use this versus shell_job_wait (blocking) or how to know the job is finished beyond the returned 'running' field. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_job_waitA
Block until a background job exits, then return its output.
Cheaper and more reliable than polling shell_job_output in a loop. If the
timeout expires the job keeps running and running comes back true — call
again or kill it.
Args: job_id: Id returned by shell_start. timeout: Max seconds to wait. The job is NOT killed on timeout. tail_lines: Return only the last N lines of each stream. 0 means all.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| timeout | No | ||
| tail_lines | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and mostly succeeds: it discloses that the call blocks, that timeout does NOT kill the job, and the `running` flag on expiry. It is softer on return shape beyond 'output' and any auth/concurrency constraints, but the important non-obvious semantics are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence followed by a compact args block; every line carries information, no filler. Easily scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations, but the description covers the essential return semantics (output plus the `running` flag) and the timeout contract. Slightly thin on exact return format, but sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate and does: it documents all three params with meaning beyond the schema, notably that timeout does not kill the job and that tail_lines=0 means all lines, plus the provenance of job_id from shell_start.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource with scope: block until a background job exits then return its output. It explicitly distinguishes itself from the sibling shell_job_output by contrasting blocking versus polling, so an agent can select it without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative (polling shell_job_output in a loop) and states the advantage, then gives the recovery path on timeout: the job keeps running, `running` returns true, call again or kill it. When-to-use, tradeoff, and follow-up action are all explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_list_jobsB
List background jobs started this session, running and finished.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose meaningful behavior: results are session-scoped and include both running and finished jobs. However, it omits that this is a non-destructive read operation, whether job IDs from other sessions are visible, or any ordering/limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the scope qualifier ('started this session') and state filter ('running and finished') both earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool with no output schema, the description covers scope and state but says nothing about the returned shape—notably whether job identifiers (needed by sibling tools like shell_job_output and shell_job_kill) are included. That gap matters because this tool is the natural entry point to the job-management workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing to document; baseline 4 applies. No ambiguity about inputs is possible.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (background jobs) with explicit scope: jobs started this session, both running and finished. This clearly distinguishes it from shell_job_output/shell_job_wait/shell_job_kill, which imply a known job ID, though it doesn't name those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is a discovery step (to find jobs before using output/wait/kill), but never states when to use it versus the other six shell tools, nor any prerequisites. No explicit when/when-not guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shell_startA
Start a long-running command in the background and return a job id.
Output is buffered as it arrives; read it with shell_job_output, block on it with shell_job_wait, stop it with shell_job_kill. Use this for builds, installs, test suites and dev servers that would blow shell_execute's timeout.
Args: command: Shell command line. cwd: Working directory. Defaults to the user's home. env: Extra environment variables.
Returns: job_id, pid, command, cwd.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | ||
| env | No | ||
| command | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does well: it discloses background execution, output buffering as it arrives, and that the operation is non-blocking with a returned job id. It does not cover environment/permission caveats or what happens if the command exits immediately, but the core behavioral model is clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then the buffering/lifecycle sentence, then compact Args and Returns blocks. Every line is useful; the Args/Returns sections restate the schema property names but earn their place given the 0% schema coverage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a job-launching tool with no output schema, the description provides the return fields (job_id, pid, command, cwd), the background/buffering semantics, and the lifecycle routing to siblings. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: all three parameters are documented (command, cwd with its home-directory default, env as extra variables). Slightly less than a 5 because per-parameter types/format nuances (e.g., env key/value semantics) are not spelled out.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: starts a long-running command in the background and returns a job id. The name and sibling set (shell_execute, shell_job_*) make the async/background distinction explicit, so an agent can tell it apart without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it (builds, installs, test suites, dev servers that would blow shell_execute's timeout) and names the follow-up tools for the full lifecycle: shell_job_output for reading, shell_job_wait for blocking, shell_job_kill for stopping. Both the trigger condition and the alternative are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v1.0.0- First observed
shell_execute - First observed
shell_info - First observed
shell_job_kill - First observed
shell_job_output - First observed
shell_job_wait - First observed
shell_list_jobs - First observed
shell_start
Related MCP Connectors
Securely control computers you explicitly pair through files, terminals, processes, screenshots, desktop UI/input, clipboard, browser automation, diagnostics, and document tools.
Scoped, audited SSH exec, sessions, and SFTP on your saved servers without exposing credentials
I run shell commands on your private cloud environment (bash, sh, zsh)
Remote shell and detached long-running jobs on your own machines — no SSH, open ports or VPN.
Related MCP Servers
- FlicenseBqualityNot gradedmaintenanceProvides unrestricted access to your development environment with filesystem operations and shell command execution capabilities, including sudo support for local development machines.49-
- FlicenseNot gradedqualityDmaintenanceProvides secure execution of terminal commands (PowerShell, CMD, shell) with configurable security policies including command blocking, path restrictions, and timeout.1-
- AlicenseNot gradedqualityDmaintenanceEnables safe shell command execution with configurable directory and command restrictions, allowing Claude Desktop to run shell commands securely.MIT
- FlicenseAqualityDmaintenanceEnables secure file system operations (read, write, delete) and simulated command execution with server-enforced permission policies, risk assessment, and human-in-the-loop approval.5-