Skip to main content
Glama

Check whether a decision function ignores a given input

check_mutation_invariance

Check if a function's output changes when a protected attribute, geography, or price is mutated. Re-runs your function per scenario and flags vacuous mutations.

Instructions

Checks -- for the scenarios you supply, not a formal proof for every possible input -- whether a decision, score, or ranking function's output changes depending on a variable it claims not to depend on: a protected attribute (name, inferred ethnicity/gender/age signal), geography, price, or any axis you name. Re-runs your actual function once per named mutation scenario and confirms the output is byte-identical to the unmutated baseline; a scenario whose mutation didn't actually change the input is flagged 'vacuous' rather than silently counting as a pass. This is the general form of check_payout_invariance -- use this one for hiring/lending/insurance/housing-style fairness claims or any other 'should not depend on X' claim; use check_payout_invariance specifically for the payout/commission axis (it also has a static-import-grep mode this tool doesn't need). Pass JS source for the function under test and each mutation -- this runs in a worker thread with a bounded timeout, not a sandbox, so only pass code you wrote or trust. A pass covers only the mutations you ran: it says nothing about values you didn't try, fields changed one at a time but never together, or a proxy field you never touched (a ZIP code standing in for race, a graduation year for age). Re-run this in CI whenever the function changes.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fnSourceYesJS source for the pure function `(input) => output` under test, e.g. "(applicant) => scoreApplicant(applicant)". Built and run in a worker thread with a bounded timeout (default 10s, see README).
baseInputYesThe baseline input to fn.
scenariosYesNamed mutation scenarios to re-run fn under.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden and does so thoroughly: it re-runs the function once per scenario, confirms byte-identical output, flags vacuous mutations, runs in a worker thread with bounded timeout that is NOT a sandbox, and states the trust constraint on passed code. It also discloses coverage limitations (untried values, one-at-a-time fields, proxy fields).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and the sibling routing, and almost every sentence earns its place given the tool's complexity. The downside is a single dense paragraph stitched with em dashes that is harder to skim than a structured list would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with no output schema, it explains result semantics (byte-identical = pass, vacuous flag) and the exact scope of what a pass covers. It stops short of describing the concrete result object shape, but the caveats and constraints an agent needs to call it correctly are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds meaning beyond the schema, notably the 'vacuous' scenario flagging and the notion that a mutation must actually change the input to count. It reinforces that scenario names appear in failure output and that mutations rewrite one axis on a copy.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('checks whether ... output changes') and resource (a decision/score/ranking function's dependence on a named variable), and explicitly differentiates itself from sibling check_payout_invariance. An agent can tell this apart from siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('hiring/lending/insurance/housing-style fairness claims or any other should-not-depend-on-X claim') versus when to prefer the alternative ('use check_payout_invariance specifically for the payout/commission axis'). Also tells the agent to re-run in CI whenever the function changes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.