Skip to main content
Glama

Run Shell Script

exec_shell
Destructive

Run scripts through bash, sh, fish, zsh, cmd, or PowerShell to execute shell syntax like pipes, redirects, &&, and globs on remote machines; set a timeout.

Instructions

Run one script through one shell (POSIX: bash/sh/fish/zsh -c; Windows: cmd /c, powershell -c). Only for shell syntax: pipes, redirects, &&, globs; otherwise exec. timeout in seconds, default 120, clamped to [1, 1800]. Invalid UTF-8 in stdout/stderr becomes U+FFFD.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
cwdNoChild working directory; omit to inherit.
shellNoPOSIX: bash (default), sh, fish, zsh. Windows: cmd (default), powershell.
scriptYesScript passed to the shell, e.g. "du -sh /var/log | sort -h".
timeoutNoKill after this many seconds (1..1800, default 120).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
okYes
errorNo
stderrNo
stdoutNo
timeoutNo
exit_codeNo
duration_msNo
duration_usNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true and idempotentHint=false, so the safety profile is covered. The description adds real behavior beyond that: the timeout is defaulted at 120s and clamped to [1, 1800], and invalid UTF-8 in stdout/stderr is silently converted to U+FFFD — a data-mangling detail an agent could not learn elsewhere.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the verb and routing rule before the mechanics. No filler; the telegraphic fragments (timeout semantics, encoding note) each carry distinct, actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and annotations carry the destructive/open-world safety signal. Combined with shell selection, timeout bounds, cwd inheritance (via schema), and encoding behavior, nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including shell defaults and the timeout range. The description largely restates those values; its only marginal addition is showing how the script string is handed to the shell (-c / /c), which is the baseline 3 case when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource ('Run one script through one shell') and immediately pins down the mechanism per platform (POSIX -c vs Windows /c / -c). It also names the sibling it is not for the non-shell case ('otherwise exec'), so an agent can route between exec_shell and exec without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly scopes usage to shell syntax (pipes, redirects, &&, globs) and names the alternative tool 'exec' for everything else — a genuine when/when-not pair. It does not, however, address the exec_start/exec_poll/exec_wait background family, so the agent must still infer whether long-running or interactive work belongs here.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.