Skip to main content
Glama

Check whether a ranking engine ignores payout

check_payout_invariance

Check whether a ranking engine's output order changes when payouts or commissions change. Run mutation scenarios or scan source for payout references to catch bias.

Instructions

Checks -- for the specific scenarios you supply, not a formal proof for every possible payout configuration -- whether a ranking/recommendation/comparison engine's output ordering changes depending on which option pays the operator more (affiliate commission, sponsored placement, referral fee). Two modes: runtime re-runs your actual ranking function under adversarial payout-mutation scenarios you name and confirms the result is byte-identical to the unmutated baseline (pass JS source for the ranking function and each mutation -- this runs in a worker thread with a bounded timeout, not a sandbox, so only pass code you wrote or trust; a scenario whose mutation didn't actually change the input is flagged 'vacuous' rather than silently counting as a pass). static-imports instead greps a set of source files for any reference to payout-related identifiers, to assert the ranking engine's code never even has payout data in scope -- no code execution needed for this mode, and it's a best-effort text/regex grep, not a real parser (it won't catch a dynamically-built import specifier or a re-export under an aliased name). Use runtime when you can call the ranking function directly; use static-imports as a cheaper, complementary check on the engine's source. A passing result means no difference in the scenarios tested, not that the function is payout-neutral in general -- write adversarial and boundary scenarios, not one easy case, and re-run this in CI whenever the ranking logic changes.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeYesWhich check to run.
filesNo[static-imports mode, required] Source files to scan for payout references.
baseInputNo[runtime mode, required] The baseline input to rankFn -- e.g. an array of candidate objects each carrying a payout/commission field.
mutationsNo[runtime mode, required, non-empty] Named adversarial payout-mutation scenarios.
rankFnSourceNo[runtime mode, required] JS source for a pure ranking function `(input) => result`, e.g. "(candidates) => candidates.slice().sort((a, b) => b.score - a.score)". Built and run in a worker thread with a bounded timeout (default 10s, see README).
stripCommentsNo[static-imports mode] Strip comments before matching, so a mention in a comment doesn't count. Default true.
payoutIdentifiersNo[static-imports mode, required, non-empty] Identifiers that must never appear in the ranking engine's source, e.g. "commission", "payout", "affiliateRate".
caseInsensitiveMatchNo[static-imports mode] Case-insensitive identifier matching. Default true.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: it discloses that runtime code executes in a worker thread with a bounded timeout and is NOT a sandbox (only pass trusted code), that static-imports is a best-effort grep rather than a real parser with named blind spots, and that vacuous mutations are flagged rather than counted as passes. It also scopes what a 'pass' actually means.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded, but the body is one very dense multi-clause sentence stuffed with parentheticals and em-dash asides, which makes it hard to scan for the mode-selection rules. The information mostly earns its place, but the structure hurts retrieval rather than helping it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, two-mode tool with no output schema, the description covers execution model, trust boundaries, mode tradeoffs, and pass semantics well. It does not describe the shape of the returned result beyond 'shown in failure output', which is the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 8 parameters (baseline 3). The description adds genuine meaning beyond it: an example shape for baseInput, guidance that mutateSource must be pure and adversarial with concrete examples, an example rankFnSource signature, and the security caveat on passing code. It stops short of enumerating defaults/mode-requirements per parameter, which the schema does cover.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (checks) and a precisely scoped resource (whether a ranking/recommendation/comparison engine's output ordering depends on operator payout), and immediately bounds the claim ('for the specific scenarios you supply, not a formal proof'). It is distinguishable from the sibling check_mutation_invariance by its explicit payout/commission focus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes mode selection: 'Use `runtime` when you can call the ranking function directly; use `static-imports` as a cheaper, complementary check on the engine's source.' It also states the ongoing workflow expectation ('re-run this in CI whenever the ranking logic changes') and warns to write adversarial/boundary scenarios rather than one easy case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.