Skip to main content
Glama

run_script

Destructive

Execute Python or Bash code in a temporary file and return stdout and stderr. Use for quick security scripting tasks when full code execution is explicitly enabled.

Instructions

Write and execute a Python or Bash script, returning its output.

Writes the code to a temporary file, executes it via python3/bash, and returns stdout/stderr. The temp file is deleted after execution. Requires CYBERSEC_MCP_ALLOW_SCRIPTS=1. This is an explicit full-code execution opt-in: scripts are not OS-sandboxed and are not constrained by CYBERSEC_MCP_ALLOW_EXTERNAL.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
codeYesThe script source code to execute.
venvNoOptional Python venv name from ~/.ctf-venvs/ (e.g. "pwntools"). Allows using a different Python with specific packages installed. Ignored for language="bash". If not set, uses the MCP server's Python.
timeoutNoMaximum execution time in seconds (default 120, max 300).
languageNo"python" (default) or "bash".python
working_dirNoWorking directory for the script (default: system temp dir).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.2.0

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive/openWorld/non-idempotent, and the description meaningfully extends them: it discloses the temp-file write-then-delete lifecycle, the interpreter used, and two independent gating/containment flags. The non-sandboxing warning is exactly the kind of consequence an agent needs before calling a mutating execution tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with what the tool does and the effect, followed by execution mechanics and then the opt-in/containment caveats. Every sentence contributes information an agent would not otherwise have.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value explanation is unnecessary, and the description covers the remaining risk surface: execution mechanism, cleanup, and the two environment flags governing permission and containment. Nothing material is missing for a correct call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the venv, timeout, language, and working_dir semantics are already fully documented in the schema, including defaults and the venv/bash interaction. The description adds no parameter-level detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb+resource pair ('Write and execute a Python or Bash script') plus the scope of the effect ('returning its output'), which is enough to separate it from sibling run_tool and run_pipeline. Nothing about the action is left ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states the precondition for use (CYBERSEC_MCP_ALLOW_SCRIPTS=1) and the security posture (not OS-sandboxed, not bound by CYBERSEC_MCP_ALLOW_EXTERNAL). It stops short of naming when to reach for run_tool/run_pipeline instead of scripting an ad-hoc program, so it lacks an explicit alternative-routing clause.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.