Skip to main content
Glama

mutation_test

Apply an explicit mutant, prove it compiles, run scoped tests, classify KILLED/SURVIVED/INVALID, and restore the file to tell a real assertion from a vacuous one.

Instructions

Mutation-test your own assertions: apply an explicit mutant, prove it still COMPILES, run a scoped test set, classify the result, and restore the file — the check that tells a real assertion from a vacuous one. Takes explicit mutants only (file_path + exact-once old_string/new_string, like edit_file); it does not generate them. Three outcomes: KILLED (mutant compiled and a test failed — the assertion is real), SURVIVED (mutant compiled and every test still passed — the assertion is VACUOUS, the finding that matters), and INVALID (the mutant did not apply, did not compile, could not be started, or timed out — it proves nothing and is NEVER reported as a kill; that false kill is why the compile gate exists). Scope the run with test_target ({target}; topology_affected says which package) and test_run (the {run} test-name filter). Commands are the stored, trust-gated [tasks.] slots run_task uses; you cannot pass a command line. They run from the git work-tree holding the mutated file: a file in another worktree of the commands' repository re-roots them there (same relative working_dir). Mutants spanning work-trees, or in a linked worktree they can't move into, are refused, never run on the wrong tree. Restoration is guaranteed on every exit path (including panic and cancellation): the pre-mutation bytes are snapshotted, rewritten under the per-path lock, and SHA-256-verified before the run is reported clean. It REFUSES a file with uncommitted changes (untracked included), no override — a clean file means git checkout recovers it if the daemon dies mid-run. It also refuses to start unless the workspace BUILDS and its tests PASS unmutated: a kill means "green before, red after", so against an already-red suite every mutant reads as killed for a reason unrelated to it. The refusal says which: suite red, timed out, or could not start — only the first is about your code. One mutation run at a time per daemon; a second call is refused rather than queued.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
mutantsNoThe mutants to test, applied and restored ONE AT A TIME. Each is an exact-once str_replace in the style of edit_file.
test_runNoOptional test-name filter for the test command's {run} placeholder (go -run, pytest -k), to run only the tests that should kill the mutants, e.g. TestFoo|TestBar. Same rules as run_task's run.
test_taskNoWhich stored [tasks.<lang>] slot runs the tests. Default "test". The built-ins are build, lint, test, e2e and verify; a project-defined slot works here too.
test_targetNoOptional value for the test command's {target} placeholder — THE way to scope the run to the affected package or test instead of the whole suite (ask topology_affected which). The shipped go/python/rust test defaults carry a defaulted placeholder, so this works with no config edit; a hand-written test command needs a {target} token of its own or the target is refused. Scoping matters: each mutant costs a full compile+test cycle, so the whole suite per mutant is the difference between minutes and tens of minutes. One shell-safe argument ([A-Za-z0-9._/:@-]).
compile_taskNoWhich stored slot proves the mutant COMPILES before its tests are trusted. Default "build". It always runs unscoped (no {target}) — a whole-module compile catches breakage a scoped test never reaches. Cannot be disabled: without it a non-compiling mutant looks exactly like a kill. The built-ins are build, lint, test, e2e and verify; a project-defined slot works here too.
timeout_secondsNoPer-step timeout for the compile and test commands. Default 600.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.21.0
    • addedInput schema / properties / test_run
      Added value: +{
      +  "description": "Optional test-name filter for the test command's {run} placeholder (go -run, pytest -k), to run only the tests that should kill the mutants, e.g. TestFoo|TestBar. Same rules as run_task's run.",
      +  "type": "string"
      +}
  2. Changed4 schema fields changedv0.17.2
    • changedInput schema / properties / compile_task / description
      Previous value: -"Which stored slot proves the mutant COMPILES before its tests are trusted. Default \"build\". It always runs unscoped (no {target}) — a whole-module compile catches breakage a scoped test never reaches. Cannot be disabled: without it a non-compiling mutant looks exactly like a kill."New value: +"Which stored slot proves the mutant COMPILES before its tests are trusted. Default \"build\". It always runs unscoped (no {target}) — a whole-module compile catches breakage a scoped test never reaches. Cannot be disabled: without it a non-compiling mutant looks exactly like a kill. The built-ins are build, lint, test, e2e and verify; a project-defined slot works here too."
    • removedInput schema / properties / compile_task / enum
      Removed value: -[
      -  "build",
      -  "lint",
      -  "test",
      -  "e2e",
      -  "verify"
      -]
    • changedInput schema / properties / test_task / description
      Previous value: -"Which stored [tasks.<lang>] slot runs the tests. Default \"test\"."New value: +"Which stored [tasks.<lang>] slot runs the tests. Default \"test\". The built-ins are build, lint, test, e2e and verify; a project-defined slot works here too."
    • removedInput schema / properties / test_task / enum
      Removed value: -[
      -  "build",
      -  "lint",
      -  "test",
      -  "e2e",
      -  "verify"
      -]
  3. Addedv0.17.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses the three outcomes (KILLED/SURVIVED/INVALID), the pre-run refusal on uncommitted changes, the build-and-test-green precondition, guaranteed restoration with SHA-256 verification under a per-path lock, worktree re-rooting and refusal rules, and the one-run-at-a-time concurrency limit. It also explains why the compile gate exists, which is genuine behavioral insight rather than restating a flag.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action is front-loaded in the first clause, and nearly every sentence carries distinct information (outcomes, gate rationale, restoration, refusals). However it is a single dense block with some redundancy between the opening framing and the SURVIVED gloss, so it is more thorough than tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, stateful, mutation-based tool with no output schema, the description covers outcomes, failure modes, refusal conditions, restoration guarantees, and scoping strategy. An agent has everything needed to invoke it correctly and interpret KILLED/SURVIVED/INVALID.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter meaning beyond the schema: that commands cannot be passed ('you cannot pass a command line'), that test_target interoperates with the defaulted {target} placeholder in shipped commands, and that compile_task 'cannot be disabled'. These nesting and interaction details go past the per-field schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('mutation-test your own assertions') and enumerates the exact workflow: apply mutant, prove compile, run scoped tests, classify, restore. It explicitly distinguishes itself from siblings by noting mutants are 'like edit_file' and that commands are the 'run_task' slots, so an agent can place it without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('the check that tells a real assertion from a vacuous one'), an explicit when-not ('it does not generate them', takes explicit mutants only), and routes to siblings ('ask topology_affected which' package to scope). Prerequisites and refusals are spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.