Skip to main content
Glama

Adversarial test of your own MCP server

adversarial_test

Tests an MCP server you own by calling non-destructive tools with command-injection and path-traversal payloads to reveal unchecked shell or filesystem access. Skips destructive tools unless enabled.

Instructions

CALLS every non-destructive tool of one configured server with command-injection payloads (each would only create an empty canary file) and path-traversal payloads, then reports which parameters reach a shell or the file system unchecked. Only for servers the user develops or operates, ideally a test instance in a container. Destructive tools are skipped unless include_destructive is true.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
serverYesServer name, optionally "scope:name".
canary_dirNoWritable directory as seen by the server (default: the OS temp directory). For a container, mount a host directory and pass host_canary_dir too.
project_dirNo
confirm_launchYesMust be true: the user agreed that the server is started and its tools are called.
host_canary_dirNo
i_own_this_serverYesMust be true: the user confirmed they own or operate this server.
include_destructiveNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.7.0

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (readOnlyHint=false, openWorldHint=true, destructiveHint=false) by disclosing side effects (each payload would only create an empty canary file), the default safety scope (non-destructive tools only), and the ownership/consent prerequisite. This is exactly the behavioral context an agent needs before invoking a tool that launches a server and executes its tools.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with zero filler: the action and payload types come first, followed by the safety constraint and exception. Every clause carries information the agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, no-output-schema tool with real blast radius, the description covers the mechanism, side effects, scope limits, and consent prerequisites, and even summarizes what the run reports ('which parameters reach a shell or the file system unchecked'). Nothing essential to safe invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is moderate (57%) and the description adds rationale for canary_dir (the canary file the payloads attempt to create) and for include_destructive. However, project_dir and host_canary_dir are left unexplained in both description and schema, so the description does not fully compensate for the coverage gap. Baseline 3 for moderate coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (calls every non-destructive tool of one configured server with injection/traversal payloads) plus the outcome reported (which parameters reach a shell or file system unchecked). This is clearly distinguishable from siblings like audit_server_tools or analyze_tool_definitions, which inspect definitions rather than actively probe them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit gating conditions: 'Only for servers the user develops or operates, ideally a test instance in a container,' and states destructive tools are skipped unless include_destructive is true. It does not name a sibling alternative for the safer static-analysis case, but the when/when-not context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.