Skip to main content
Glama

honesty-mcp

An MCP (Model Context Protocol) server that exposes eleven already-built, zero-runtime-dependency "honesty SDK" TypeScript libraries as tools any MCP-compatible coding agent (Claude Code, Claude Desktop, or any other MCP client) can call while building or auditing a product — grounded-citation checking, evidence corroboration grading, payout/mutation-invariance checks (scenario-tested, not a formal proof for every possible input), trust/ authenticity scoring, provenance-claim validation, a claims registry, a tamper-evident audit chain, and recommendation/decision grading plus engine-vs-human divergence, plus scaffolding for two runtime-library kits that don't fit the "check this content" shape.

This server is thin by design: every tool is a Zod input schema plus a handler that imports and calls the real, unmodified export from the wrapped kit. It does not reimplement any of the underlying logic.

Install

honesty-mcp is published on npm and wraps its eleven sibling "honesty SDK" kits as ordinary registry dependencies (semver ranges, not local file: paths) — npm install (or a plain npx) resolves everything from the public registry, with nothing else to check out first.

The fastest way to run it, with no install step at all:

npx honesty-mcp

Or add it to a project:

npm install honesty-mcp

Adding it to an MCP client

Claude Code:

claude mcp add honesty-mcp -- npx -y honesty-mcp

Any other MCP client that reads an mcpServers config (Claude Desktop's claude_desktop_config.json uses the same shape):

{
  "mcpServers": {
    "honesty-mcp": {
      "command": "npx",
      "args": ["-y", "honesty-mcp"]
    }
  }
}

The server runs on the stdio transport — it reads JSON-RPC requests from stdin and writes responses to stdout, logging only a one-line startup banner to stderr so stdout stays a clean JSON-RPC channel. This is the transport Claude Code, Claude Desktop, and most local/CLI MCP client integrations expect for a locally-run server; it is not an HTTP server and has no port to browse to.

Listing in the MCP registry

honesty-mcp also ships a server.json (validated against the official schema) so it can be listed on the official MCP registry, under the name io.github.lkopietz3-byte/honesty-mcp. Listing is a manual, one-time step the maintainer runs after this package is on npm — it is not run as part of this repo's CI or npm publish. Exact commands, from the registry's own docs (publishing guide, CLI reference):

# 1. Install the publisher CLI (macOS/Linux)
curl -L "https://github.com/modelcontextprotocol/registry/releases/latest/download/mcp-publisher_$(uname -s | tr '[:upper:]' '[:lower:]')_$(uname -m | sed 's/x86_64/amd64/;s/aarch64/arm64/').tar.gz" | tar xz mcp-publisher
sudo mv mcp-publisher /usr/local/bin/

# 2. Log in with GitHub (proves ownership of the io.github.lkopietz3-byte namespace)
mcp-publisher login github

# 3. Publish (validates server.json against the schema, then submits it)
mcp-publisher publish

package.json's mcpName field (io.github.lkopietz3-byte/honesty-mcp) must match server.json's name field exactly — that's how the registry verifies the npm package and the registry listing belong to the same publisher (see package-types.mdx, "Ownership Verification").

Related MCP server: Agent Trust Stack MCP Server

Tools

9 content-checking tools' worth of kits, 12 tools (kits 1–9 — hand one of these real content or data, get back a structured verdict):

Tool

Wraps

One-line purpose

check_grounding

grounding-kit

Detects ungrounded or forged citations in AI-generated text, sentence by sentence.

corroborate_evidence

corroboration-kit

Grades confidence in a claim from independent evidence signals — not a vote count.

check_payout_invariance

payout-invariance-kit

Checks, for the scenarios you supply, whether a ranking engine's output changes with who pays more (runtime mode) or whether payout identifiers appear in its source at all (static-imports mode). Not a formal proof for every possible payout configuration.

check_mutation_invariance

mutation-invariance-kit

Checks, for the scenarios you supply, whether a decision/score/ranking function's output depends on a variable it claims not to (protected attribute, geography, price, or anything you name). Not a formal proof for every possible input.

score_trust_identified

trust-core (identified)

Scores an entity from known, identified contributors (reviewer accounts, raters, inspectors).

assess_anonymous_authenticity

trust-core (anonymous)

Scores unattributed/scraped sentiment on how organic it looks, with a heuristic discount for two specific manipulation patterns (source concentration, suspiciously uniform sentiment) — not a fraud detector.

check_provenance_claims

provenance-kit

Flags certainty-implying language ("(verified)", "guaranteed") not backed by an appropriate provenance tier.

check_claims_registry

claims-registry-kit

Buckets public-facing claims as current / stale / unverified against their linked evidence and last-verified date.

append_audit_entry

audit-chain-kit

Appends one entry to a hash-chained, tamper-evident audit log.

verify_audit_chain

audit-chain-kit

Independently re-verifies a hash chain from genesis; detects mutation and severed links unconditionally, plain tail deletion with expectedMinLength, and truncate-and-re-append (or a full rewrite) only with anchor — expectedMinLength alone does NOT catch a truncate-and-re-append.

grade_decision

advice-ledger-kit

Grades one recommendation-and-decision pair against a before/after observation log — exposure-aligned, with separate floors per window and machine-readable refusal codes. holding requires both the exposed bad count to stay under refuteThreshold AND its bad rate to not exceed the baseline's (exact comparison, not the rounded badRate); either failing is not-holding (see the tool description).

compute_divergence

advice-ledger-kit

Measures how often an engine and a human disagreed, and where a later outcome exists, reports engine-right/human-right as separate, never-blended counts.

2 scaffolding tools (agent-receipt-kit, cost-governor-kit — these are runtime libraries a project installs and imports into its own running code, not something checked on-demand the same way; see the code comments in src/tools/agentReceiptScaffold.ts / costGovernorScaffold.ts for why a "check this content" shape would misrepresent them):

Tool

Wraps

One-line purpose

scaffold_agent_receipts

agent-receipt-kit

Explains the issue-a-packet / verify-the-claim authorization pattern for AI agents, with an install step, a starter snippet, and (optionally) a live accepted/rejected worked example run against the real kit.

scaffold_cost_governor

cost-governor-kit

Explains the estimated pre-call spend check + cache-aware pricing math (default 0.1x cache-read rate, overridable per model via cacheReadPerMillion) + advisory (not concurrency-safe) check-then-commit usage-counting pattern (a commit failure after a successful call returns commitError instead of discarding the result), with an install step, a starter snippet, and (optionally) a live worked example.

That's 9 kits, 12 content-checking tool slots (payout-invariance-kit, audit-chain-kit, and advice-ledger-kit each got 2 tools instead of 1, since each bundles two genuinely distinct operations — runtime check vs. static grep; append vs. verify; grading one decision vs. measuring divergence across a whole population — that read better as separate, single-purpose MCP tools than as one tool with a mode switch) plus 2 scaffolding tools = 14 tools total.

A note on the code-execution tools

check_payout_invariance (runtime mode) and check_mutation_invariance both wrap kit functions whose real signature takes an actual JavaScript function (a ranking function, a mutation closure) — there is no way to represent a function as MCP JSON arguments. Both tools accept that function as a source-code string and build the real function via the Function constructor (src/lib/buildFunction.ts) before calling the unmodified kit export with it. compute_divergence's optional config.groupBySource uses the same mechanism for advice-ledger-kit's groupBy config function, but only when a caller actually supplies it — omitting it (the common case) needs no code execution at all, since the library's own default groups by each pair's group field. This is equivalent to eval for that one string. It is appropriate here because this is a local, stdio-only dev tool a coding agent runs against its own project's code — the same trust model as that agent running node -e, vitest run, or any other local code-execution tool — and it is not designed to be exposed to untrusted, remote, or adversarial input. Don't wire this server up to accept tool arguments from anyone other than the trusted local agent driving it.

All three of these run inside a worker_threads Worker with a bounded timeout (src/lib/runFunctionJob.ts), not on the server's own main thread. Earlier versions called the built function synchronously in-process, so a while (true) {} source string would block the entire stdio server forever — every other in-flight or future tool call along with it. Now, if the worker doesn't finish within the timeout (default 10 seconds, override with the HONESTY_MCP_WORKER_TIMEOUT_MS environment variable, in milliseconds), it is forcibly terminate()d and the tool call returns a clear timeout error instead of hanging. This is not a sandbox — the worker has this process's full OS-level privileges (filesystem, network, environment variables) — it only bounds time. It does not stop caller code from reading your filesystem, making network requests, or doing anything else this Node process itself can do. The trust model above still applies in full: only pass code you wrote or trust, and don't expose this server to untrusted, remote, or adversarial input.

Honest limits

  • This server adds no guarantee beyond what the kit it wraps already documents. Every tool description was checked against its kit's own README "Honest limits"/"Limits" section as of the sibling-kit commits this branch was built against — read that kit's README for the full detail behind any one-line tool description or table row here. If a kit's own limits change later, this server's wording can drift out of sync again; nothing here re-checks that automatically.

  • check_payout_invariance, check_mutation_invariance, and compute_divergence's worker-thread timeout bounds time only, not behavior — see "A note on the code-execution tools" above. It is not a sandbox.

  • SERVER_VERSION (src/server.ts) and package.json's version are two separate values with no automated sync. They happen to agree today; a future release could forget to bump one.

Development

To work on this server itself (rather than just use it), clone the repo and build from source:

git clone https://github.com/lkopietz3-byte/honesty-mcp.git
cd honesty-mcp
npm install          # resolves the 11 wrapped kits from the npm registry
npm run build         # compiles src/ -> dist/
npm start              # runs dist/index.js on stdio
npm run lint         # eslint . --max-warnings=0
npm run typecheck   # tsc --noEmit over src/ + test/
npm test             # vitest — spins up a real MCP Client/Server pair
                      # over the SDK's InMemoryTransport and calls tools
                      # end-to-end (see test/*.test.ts)
npm run build        # tsc -p tsconfig.build.json -- compiles src/ -> dist/
npm run smoke        # npm run build first, then node dist/smoke.js --
                      # a standalone script that does the same real
                      # client/server handshake, lists all tools, and
                      # calls one end-to-end
npm run verify       # lint + typecheck + test + build + smoke, in order
npm run audit:dependencies   # npm audit --package-lock-only --include=dev
                              # --ignore-scripts --audit-level=low

Project layout

src/
  index.ts                    entrypoint: connects the server to stdio
  server.ts                   createServer() -- builds the McpServer and registers every tool
  smoke.ts                    standalone smoke-test script (see npm run smoke)
  lib/
    result.ts                 CallToolResult helpers (jsonResult / errorResult)
    buildFunction.ts           turns a JS source string into a callable function
    runFunctionJob.ts          runs that function (or the divergence groupBySource case)
                               inside a worker_threads Worker with a bounded timeout
  tools/
    grounding.ts
    corroboration.ts
    payoutInvariance.ts
    mutationInvariance.ts
    trustIdentified.ts
    trustAnonymous.ts
    provenance.ts
    claimsRegistry.ts
    auditChain.ts               (append_audit_entry + verify_audit_chain)
    adviceLedger.ts              (grade_decision + compute_divergence)
    agentReceiptScaffold.ts
    costGovernorScaffold.ts
test/
  server.test.ts               end-to-end tests over a real in-memory MCP client/server pair,
                                covering every tool's known-good and known-bad input
  buildFunction.test.ts        unit tests for the source-string -> function builder
  runFunctionJob.test.ts       unit tests for the worker-thread job runner (normal jobs,
                                error propagation, timeout enforcement, env var parsing)
  workerTimeout.test.ts        end-to-end proof that the three code-execution tools stay
                                bounded by the worker-thread timeout instead of hanging

Available Tools

14 tools
append_audit_entryAppend a tamper-evident audit-chain entryA

Appends one entry to an append-only, hash-chained audit log and returns the new (longer) chain. Each entry's hash binds its payload, index, createdAt, a fixed format tag (formatVersion, currently "audit-chain-kit/v1"), and the previous entry's hash together, so any later tampering with an earlier entry -- or an entry produced by an incompatible version of audit-chain-kit -- breaks the chain in a way verify_audit_chain will detect. Use this wherever you need a tamper-evident record of events (agent actions, approvals, state transitions) that a skeptical third party can later verify independently. This function does not persist anything itself -- store the returned chain (or just the new entry) in whatever your app already uses.

ParametersJSON Schema
NameRequiredDescriptionDefault
chainYesThe existing chain, in order, exactly as previously returned/stored. Pass [] to start a new chain.
payloadYesCaller data for the new entry -- any JSON-serializable value.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses the append-only hash-chain mechanism, what each entry hash binds, the fixed format tag, that tampering or an incompatible format version will break the chain in a detectable way, and crucially that the function does not persist anything itself. This is exactly the behavioral context an agent needs before invoking it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and return value, then mechanism, then usage guidance, then the important non-persistence caveat. It is somewhat long, but nearly every sentence carries useful technical information for an annotation-free tool handling a hash chain.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a tamper-evident hash chain, the absence of annotations, and the lack of an output schema, the description is complete enough: it explains the return value, persistence expectations, verification path, and versioning constraints. An agent has what it needs to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents both parameters and their nested fields thoroughly. The description adds thematic explanation of how the chain and payload are bound by hashes, but does not add parameter-level syntax or format details beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb and resource: appending one entry to an append-only, hash-chained audit log and returning the new longer chain. It also names verify_audit_chain and explains how tampering or version mismatch will be detected there, which clearly distinguishes it from the verification sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit usage context: 'Use this wherever you need a tamper-evident record of events (agent actions, approvals, state transitions) that a skeptical third party can later verify independently.' It also clarifies the non-persistence behavior and what to do with the return value, though it does not explicitly state when not to use it or name a direct alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assess_anonymous_authenticityScore how organic unattributed signals lookA

Scores how organic a corpus of UNATTRIBUTED, scraped sentiment signals (crawled mentions, imported reviews with no verifiable identity, aggregator feeds) looks, weighing a positive composite of consensus/diversity/volume/recency against a heuristic penalty for two specific, cheap manipulation patterns: evidence concentrated in a single source, and suspiciously uniform sentiment (near-maximal with near-zero variance -- the fingerprint of copy-pasted or purchased praise). This is NOT a fraud or astroturf detector: it cannot show that sentiment is fabricated or that any reviewer is fake, and a campaign that varies its wording/sentiment and spreads across several sources isn't caught by these two checks. Treat a low score as 'looks statistically unusual in a specific way worth a human look', not as a fraud finding. Use this for reviews/mentions/buzz with no identity behind them. For signals from known, identified contributors, use score_trust_identified instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowYesISO 'now' timestamp recency decay is computed against. Pass a fixed value for determinism.
configNoPartial override merged over the library's illustrative EXAMPLE_ANONYMOUS_CONFIG -- omit to use the example config as-is.
signalsYesThe unattributed signals to assess. May be empty.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it explains the positive composite, the two specific penalty patterns, and the tool's limits (it cannot show fabrication and misses varied, multi-source campaigns). It does not explicitly state side-effect purity or deterministic behavior beyond the schema's now parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, and the description is dense but largely earns its length by covering scope, interpretation, and limitations. Some sentences are long, but the complexity of the tool justifies most of the space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description gives rich context for a complex scoring tool, including scope, interpretation, and caveats. Because there is no output schema, it would be stronger if it specified the return score's range or shape, but it is still sufficient to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters and nested config properties thoroughly. The description adds conceptual meaning about consensus/diversity/volume/recency and the two manipulation checks, but no additional parameter-level semantics beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: it scores how organic a corpus of unattributed, scraped sentiment signals looks. It clearly distinguishes itself from the sibling score_trust_identified and explicitly states what kind of signals it handles.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance (reviews/mentions/buzz with no identity behind them) and names the alternative tool for identified contributors. It also clarifies when not to over-interpret the result, saying a low score is not a fraud finding.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_claims_registryCheck public claims for stale or missing evidenceA

Keeps public-facing product claims honest over time by checking each one against its own claimed evidence -- the SOC2-control-evidence pattern applied to marketing/product copy instead of compliance controls. Buckets every claim into 'current' (evidence present, review within policy), 'stale' (evidence present but verifiedAt is older than maxAgeDays, or unparseable), or 'unverified' (no evidenceRef at all -- this always wins over staleness, since a fresh date next to an empty reference proves nothing). Use this as a periodic 'Monday-morning' review or a CI gate on a claims registry, to catch marketing copy that drifted out of sync with what the product actually does after a refactor. Note: this only checks that a reference EXISTS and is fresh, not that the thing it points to still actually supports the claim's text -- pair with check_grounding for that.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowNoISO 'now' timestamp to evaluate against. Defaults to the current time.
claimsYesThe claims to evaluate.
maxAgeDaysYesStaleness policy: evidence older than this many days is flagged stale. 0 is valid (every claim must have been verified today or it's stale) -- the kit requires a finite number >= 0.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden and does so well: it defines the three buckets, the precedence rule ('unverified always wins over staleness'), and the scope limit (existence + freshness only, not grounding). It stops short of stating side effects or that it is a pure read, but the semantic disclosure is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action, then usage, then the caveat, and every section is relevant. The SOC2-controls analogy costs a few words but meaningfully conveys the pattern, so the size is justified for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and no annotations, the description supplies the missing return semantics by naming the three classification outcomes and their tie-breaking. Nothing essential for correct invocation or interpretation appears to be absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and every parameter is already documented in the schema (including the empty/whitespace evidenceRef rule and the maxAgeDays semantics). The description reinforces the maxAgeDays/verifiedAt relationship in prose but adds no format or syntax detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (checking) applied to a specific resource (public-facing product claims registry) and immediately defines the three output buckets. It is clearly distinguishable from siblings, naming check_grounding as the adjacent tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames when to use it ('periodic Monday-morning review or a CI gate') and states the boundary condition and the alternative tool to pair with ('not that the thing it points to still actually supports the claim's text -- pair with check_grounding'). This is exactly the when/when-not/alternative guidance expected.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_groundingCheck citation groundingA

Detects ungrounded or forged citations in AI-generated text. Splits text into sentences and classifies each one against evidence: 'grounded' (cites a marker whose evidence plausibly supports it), 'placeholder' (an honest 'TBD'/unknown gap, no fake citation), 'ungrounded' (a claim with no citation at all), or 'invalid' (cites a marker id that is missing from evidence, or whose evidence doesn't plausibly support the sentence under the default matcher -- i.e. a forged or hallucinated citation; this outranks every other status). This is a mechanical/structural check, not a truth checker: the default support check is naive substring/word-overlap matching, not semantic entailment -- it can pass a coincidental word match and can fail a genuine paraphrase, and it cannot verify that the evidence itself is true. Use this before shipping any AI-written report, summary, or answer that cites sources, to catch a model inventing or misattributing a citation. Treat any 'invalid' sentence as a hard stop; treat 'ungrounded' sentences as claims that should probably cite something but currently don't. For higher-stakes content, use grounding-kit directly with a custom supports() function (embedding-similarity or NLI-based) instead of the default matcher.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe AI-generated text to check. Citation markers use the kit's default convention `[[cite:id]]` (e.g. "The bridge opened in 1932 [[cite:source-a]]."). This tool uses the default marker/placeholder patterns; if your generator emits a different citation syntax (e.g. "[1]"), rewrite markers to `[[cite:1]]` before calling, or use grounding-kit directly with a custom markerPattern.
evidenceYesMap of citation marker id -> the evidence text/span it claims to support. This is the closed world: a marker cited in `text` whose id is NOT a key here is flagged invalid, and a marker whose evidence text doesn't plausibly support the sentence is also flagged invalid. Pass {} if there is no evidence at all (every citation will then be invalid, and uncited claims will be 'ungrounded').

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses that this is a mechanical/structural check, that the default matcher is naive substring/word-overlap rather than semantic entailment, that it can both false-positive (coincidental word match) and false-negative (genuine paraphrase), and that it cannot verify the truth of the evidence itself. It also states the priority rule ('invalid outranks every other status').

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then classification definitions, matcher caveat, and escalation path in a logical order where each sentence carries distinct information. It is long, but the length is justified by the four-status taxonomy and the non-obvious matcher limitations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description fully enumerates the return classifications and their meanings, and covers the matcher's failure modes and parameter interplay. For a 2-parameter tool with a nested evidence map and no annotations, nothing an agent needs to invoke and interpret it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema by framing `evidence` as a closed world and explaining the parameter interaction (a marker cited in `text` missing from `evidence` is invalid; empty evidence makes all citations invalid). It reinforces the `[[cite:id]]` marker convention and notes what to do with alternate syntax.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource ('Detects ungrounded or forged citations in AI-generated text') and the body enumerates the four classification outcomes. It is clearly distinguishable from siblings like corroborate_evidence or check_provenance_claims by naming the exact artifact it inspects (sentence-level citation markers vs. evidence map).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it ('before shipping any AI-written report, summary, or answer that cites sources') and when to escalate ('for higher-stakes content, use grounding-kit directly with a custom supports() function'). It also names the alternative tool and the condition that selects it, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_mutation_invarianceCheck whether a decision function ignores a given inputA

Checks -- for the scenarios you supply, not a formal proof for every possible input -- whether a decision, score, or ranking function's output changes depending on a variable it claims not to depend on: a protected attribute (name, inferred ethnicity/gender/age signal), geography, price, or any axis you name. Re-runs your actual function once per named mutation scenario and confirms the output is byte-identical to the unmutated baseline; a scenario whose mutation didn't actually change the input is flagged 'vacuous' rather than silently counting as a pass. This is the general form of check_payout_invariance -- use this one for hiring/lending/insurance/housing-style fairness claims or any other 'should not depend on X' claim; use check_payout_invariance specifically for the payout/commission axis (it also has a static-import-grep mode this tool doesn't need). Pass JS source for the function under test and each mutation -- this runs in a worker thread with a bounded timeout, not a sandbox, so only pass code you wrote or trust. A pass covers only the mutations you ran: it says nothing about values you didn't try, fields changed one at a time but never together, or a proxy field you never touched (a ZIP code standing in for race, a graduation year for age). Re-run this in CI whenever the function changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
fnSourceYesJS source for the pure function `(input) => output` under test, e.g. "(applicant) => scoreApplicant(applicant)". Built and run in a worker thread with a bounded timeout (default 10s, see README).
baseInputYesThe baseline input to fn.
scenariosYesNamed mutation scenarios to re-run fn under.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden and does so thoroughly: it re-runs the function once per scenario, confirms byte-identical output, flags vacuous mutations, runs in a worker thread with bounded timeout that is NOT a sandbox, and states the trust constraint on passed code. It also discloses coverage limitations (untried values, one-at-a-time fields, proxy fields).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and the sibling routing, and almost every sentence earns its place given the tool's complexity. The downside is a single dense paragraph stitched with em dashes that is harder to skim than a structured list would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with no output schema, it explains result semantics (byte-identical = pass, vacuous flag) and the exact scope of what a pass covers. It stops short of describing the concrete result object shape, but the caveats and constraints an agent needs to call it correctly are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds meaning beyond the schema, notably the 'vacuous' scenario flagging and the notion that a mutation must actually change the input to count. It reinforces that scenario names appear in failure output and that mutations rewrite one axis on a copy.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('checks whether ... output changes') and resource (a decision/score/ranking function's dependence on a named variable), and explicitly differentiates itself from sibling check_payout_invariance. An agent can tell this apart from siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('hiring/lending/insurance/housing-style fairness claims or any other should-not-depend-on-X claim') versus when to prefer the alternative ('use check_payout_invariance specifically for the payout/commission axis'). Also tells the agent to re-run in CI whenever the function changes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_payout_invarianceCheck whether a ranking engine ignores payoutA

Checks -- for the specific scenarios you supply, not a formal proof for every possible payout configuration -- whether a ranking/recommendation/comparison engine's output ordering changes depending on which option pays the operator more (affiliate commission, sponsored placement, referral fee). Two modes: runtime re-runs your actual ranking function under adversarial payout-mutation scenarios you name and confirms the result is byte-identical to the unmutated baseline (pass JS source for the ranking function and each mutation -- this runs in a worker thread with a bounded timeout, not a sandbox, so only pass code you wrote or trust; a scenario whose mutation didn't actually change the input is flagged 'vacuous' rather than silently counting as a pass). static-imports instead greps a set of source files for any reference to payout-related identifiers, to assert the ranking engine's code never even has payout data in scope -- no code execution needed for this mode, and it's a best-effort text/regex grep, not a real parser (it won't catch a dynamically-built import specifier or a re-export under an aliased name). Use runtime when you can call the ranking function directly; use static-imports as a cheaper, complementary check on the engine's source. A passing result means no difference in the scenarios tested, not that the function is payout-neutral in general -- write adversarial and boundary scenarios, not one easy case, and re-run this in CI whenever the ranking logic changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesWhich check to run.
filesNo[static-imports mode, required] Source files to scan for payout references.
baseInputNo[runtime mode, required] The baseline input to rankFn -- e.g. an array of candidate objects each carrying a payout/commission field.
mutationsNo[runtime mode, required, non-empty] Named adversarial payout-mutation scenarios.
rankFnSourceNo[runtime mode, required] JS source for a pure ranking function `(input) => result`, e.g. "(candidates) => candidates.slice().sort((a, b) => b.score - a.score)". Built and run in a worker thread with a bounded timeout (default 10s, see README).
stripCommentsNo[static-imports mode] Strip comments before matching, so a mention in a comment doesn't count. Default true.
payoutIdentifiersNo[static-imports mode, required, non-empty] Identifiers that must never appear in the ranking engine's source, e.g. "commission", "payout", "affiliateRate".
caseInsensitiveMatchNo[static-imports mode] Case-insensitive identifier matching. Default true.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: it discloses that runtime code executes in a worker thread with a bounded timeout and is NOT a sandbox (only pass trusted code), that static-imports is a best-effort grep rather than a real parser with named blind spots, and that vacuous mutations are flagged rather than counted as passes. It also scopes what a 'pass' actually means.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded, but the body is one very dense multi-clause sentence stuffed with parentheticals and em-dash asides, which makes it hard to scan for the mode-selection rules. The information mostly earns its place, but the structure hurts retrieval rather than helping it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, two-mode tool with no output schema, the description covers execution model, trust boundaries, mode tradeoffs, and pass semantics well. It does not describe the shape of the returned result beyond 'shown in failure output', which is the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 8 parameters (baseline 3). The description adds genuine meaning beyond it: an example shape for baseInput, guidance that mutateSource must be pure and adversarial with concrete examples, an example rankFnSource signature, and the security caveat on passing code. It stops short of enumerating defaults/mode-requirements per parameter, which the schema does cover.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (checks) and a precisely scoped resource (whether a ranking/recommendation/comparison engine's output ordering depends on operator payout), and immediately bounds the claim ('for the specific scenarios you supply, not a formal proof'). It is distinguishable from the sibling check_mutation_invariance by its explicit payout/commission focus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes mode selection: 'Use `runtime` when you can call the ranking function directly; use `static-imports` as a cheaper, complementary check on the engine's source.' It also states the ongoing workflow expectation ('re-run this in CI whenever the ranking logic changes') and warns to write adversarial/boundary scenarios rather than one easy case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_provenance_claimsCheck claims for unbacked certainty languageA

Scans reader-facing claims for certainty-implying language ("(verified)", "independently verified", "guaranteed", "fact-checked", "100% accurate", ...) that isn't backed by an appropriate provenance tier, plus tiers that require a sourceRef but don't have one, and claims carrying an unrecognized tier. This is the exact pattern that caught ~150 false '(verified)' labels on a live site after they had already shipped -- run it on any copy, marketing page, or AI-drafted content that makes factual-sounding claims before it ships, not after. The default phrase list is a starting point drawn from that one incident, not a taxonomy -- it will miss phrases it doesn't know about (e.g. 'clinically proven', 'third-party tested'); extend certaintyPhrases for your domain. Negation detection is a fixed character window before a match, not a parser, so it can miss a negation in an earlier clause or over-suppress one further away. An empty result means every claim's certainty language (that this tool's phrase list and negation window caught) is backed by its tier.

ParametersJSON Schema
NameRequiredDescriptionDefault
claimsYesThe claims to check.
caseSensitiveNoCase-sensitive phrase matching. Default false.
negationWindowNoCharacters before a phrase match to scan for a negation word ('not', 'without', ...). Default 40.
certaintyPhrasesNoOverride the certainty-phrase vocabulary. Defaults to a generic starter list ('(verified)', 'independently verified', 'proprietary dataset', 'guaranteed', 'fact-checked', '100% accurate', ...).
certaintyRequiresTierNoTiers strong enough to back certainty language. Default ['verified'].
requireSourceRefForTiersNoTiers that must carry a non-empty sourceRef. Default ['verified'].

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden and does so unusually well: it discloses the incident-derived provenance of the default phrase list, the phrase-list coverage gap ('will miss phrases it doesn't know about'), the negated-window limitation ('fixed character window... not a parser'), and the precise meaning of an empty result. These are exactly the caveats an agent needs to avoid over-trusting the output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core detection behavior, then limitations. Dense and every sentence carries signal, though the incident anecdote is longer than strictly required and could be tightened without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description does most of the work, covering scope, limits, and empty-result semantics. Its one gap is the shape of the returned offenses (it references 'offense messages' but never describes the result structure), which matters for a tool with no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description adds value beyond the schema by explaining that certaintyPhrases defaults are a starting point drawn from one incident and should be extended per domain, and by clarifying the tier semantics context for certaintyRequiresTier/requireSourceRefForTiers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('scans reader-facing claims for certainty-implying language') and enumerates the three detection classes (unbacked certainty phrases, tiers requiring a sourceRef, unrecognized tiers). This is clearly distinct from siblings like check_grounding or check_claims_registry without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong workflow context ('run it on any copy, marketing page, or AI-drafted content that makes factual-sounding claims before it ships, not after') and a domain-extension hint for certaintyPhrases. It stops short of naming a sibling alternative or the condition that would route to one, so it falls just below the top bar.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compute_divergenceMeasure where two judges disagree, and who was rightA

Computes how often two independent judges (e.g. an automated engine and a human override) rated the same object differently, and -- where a later outcome exists to settle it -- reports engineRight and humanRight as two SEPARATE counts, never combined into one blended accuracy number (a system can be right most of the time overall and still wrong every single time a human bothers to overrule it, and that second fact is the one worth acting on). Divergence needs BOTH a count floor (minDivergentCount) and a rate floor (minDivergentRate) before it's 'reportable' -- either alone lets noise through. Calibration (who was right) carries its own separate floor (minResolvedDivergent) and its own status, so a population can legitimately have reportable divergence and refused calibration at the same time: plenty of disagreements, not enough of them settled by a later outcome yet. Supply group on each pair (or a custom config.groupBySource) to also get a sorted per-group breakdown alongside the overall report. Use this whenever a system has both an automated judgment and a human override/correction recorded for the same objects, to find out whether overruling the system is actually earning its keep. For grading one recommendation's own before/after outcome window instead of a whole population of engine-vs-human calls, use grade_decision.

ParametersJSON Schema
NameRequiredDescriptionDefault
pairsYesEvery occasion both judges could have spoken on. Pairs missing one judgment are still counted, just not comparable.
configNoFloors, thresholds, and optional custom grouping, every field optional with a library default (see advice-ledger-kit's DEFAULT_DIVERGENCE_CONFIG). Omit entirely to use the library's defaults.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses that engineRight/humanRight are never blended, that divergence requires BOTH a count and rate floor, that calibration has its own separate floor and status (so divergence can be reportable while calibration is refused), and that groupBySource runs in a worker thread with a bounded timeout and a code-trust caveat.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is front-loaded with the core computation, but the body is dense with nested parentheticals and run-on sentences that carry multiple ideas each. For a complex statistical tool the length is partly justified, yet several clauses could be trimmed without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description must carry the load; it explains the computed counts, the floor semantics, the separate calibration status, and the grouping behavior. It stops short of describing the report's full shape (e.g. examples, overall structure), which is the one remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds conceptual meaning beyond the per-field text: why two separate counts exist (system can be right overall yet wrong on every overrule) and why both floors must clear together. Some of this overlaps the schema's own phrasing, keeping it below a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('computes how often two independent judges rated the same object differently') and even names the two output counts (engineRight/humanRight). It explicitly distinguishes itself from the sibling grade_decision by scope (population of engine-vs-human calls vs one recommendation's window).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger ('use this whenever a system has both an automated judgment and a human override/correction recorded for the same objects') and an explicit alternative with its own condition ('for grading one recommendation's own before/after outcome window ... use grade_decision'). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

corroborate_evidenceGrade evidence corroboration for a claimA

Grades confidence in a claim from a set of independent evidence signals -- NOT a vote count. 'confirmed' requires 2+ distinct supporting sources AND at least one non-textual signal; disagreement among sources surfaces as 'mixed' rather than being averaged away; a null result is 'not-found' only under adequate coverage, otherwise 'inconclusive' (a thin sample can't prove a negative); and thin coverage caps the verdict below 'confirmed' no matter how clean the signals look. Use this whenever several pieces of evidence were gathered for a claim (by you, another tool, or a research/verification pass) and you need an honest, non-inflated verdict instead of eyeballing how many checks 'passed'.

ParametersJSON Schema
NameRequiredDescriptionDefault
signalsYesThe evidence signals gathered about the claim. May be empty.
coverageYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the verdict logic (2+ distinct sources plus a non-textual signal for 'confirmed', disagreement surfacing as 'mixed', null results gated by coverage, thin coverage capping the verdict). It does not describe error behavior or the shape of the returned verdict object, leaving a modest gap for a scoring tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The key constraint ('NOT a vote count') is front-loaded and the grading rules follow in a tight, information-dense sentence. It is longer than average but nearly every clause carries a distinct rule; minor density without wasted sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description conveys the decision rules an agent needs to invoke it correctly and anticipate verdict semantics. It stops short of stating return-value structure, but the embedded verdict names give sufficient orientation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%, but the description compensates by explaining how the parameters drive outcomes -- distinct `source` counting for independence, non-textual `kind` unlocking 'confirmed', and coverage capping the verdict. This adds real semantic meaning beyond the field-level schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('grades confidence in a claim from a set of independent evidence signals') and immediately differentiates itself from a naive alternative with 'NOT a vote count'. An agent can distinguish this from siblings like check_grounding or score_trust_identified without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it: 'whenever several pieces of evidence were gathered for a claim (by you, another tool, or a research/verification pass) and you need an honest, non-inflated verdict'. It gives a clear triggering context but does not name or contrast against specific sibling tools, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grade_decisionGrade a decision against what happened after itA

Grades one recommendation-and-decision pair against a before/after observation log. Compares a pre-decision BASELINE window to a post-decision, exposure-aligned RESULT window on the same subjectId + checkKey, and returns 'holding' (the exposed post-decision window stayed under refuteThreshold bad observations AND its bad rate is not higher than the baseline's), 'not-holding' (EITHER the exposed post-decision window hit refuteThreshold bad observations, OR its bad rate is higher than the baseline's, even below that count -- 1 bad of 10 before, 1 bad of 3 exposed since is 'not-holding', not 'holding', despite the count staying under the default bar of 2), or 'refused' (the evidence didn't clear a floor -- see refusalCodes for exactly which one, never a vague 'unproven'). The rate comparison is exact -- it cross-multiplies the raw counts rather than comparing the rounded badRate fields or the sign of badRateDelta, so a not-holding can occur even when both windows display the same 3-place badRate. Always show badRateDelta and the raw baseline/result counts next to the verdict, don't quote 'holding' on its own. The verdict grades the DECISION, not just whether advice was taken: a DISMISSED recommendation whose problem later surfaced also grades 'not-holding', because the evidence sided with the advice either way. Two things this tool refuses to let slide: (1) 'nothing has gone wrong since' only counts as a result if the baseline shows the problem existed before -- otherwise it's 'baseline_lacks_negative_signal'; (2) only observations where the advice could actually have applied (exposed === true) can create or reverse the headline verdict -- everything else is reported separately as secondary (whose own wouldBeVerdict can itself be 'refused'), labelled non-headline, and can never become the headline. Use this to close the loop on any recommender/advisor system that logs recommendations, human accept/reject calls, and later outcomes, instead of letting advice quality go unmeasured forever. For grading a whole population of engine-vs-human calls at once (rather than one decision against its own before/after window), use compute_divergence instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
configNoFloors and thresholds, every field optional with a library default (see advice-ledger-kit's DEFAULT_GRADE_CONFIG). Omit entirely to use the library's defaults.
decisionYesThe human's call on the recommendation.
observationsYesEvery observation available. This tool does the filtering by subjectId + checkKey itself, so passing the whole ledger (observations for other subjects/checks too) is fine.
recommendationYesThe advice being graded.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discharges it: it enumerates the three verdicts, the exact cross-multiplication semantics, the exposure (exposed === true) rule, the baseline_negative_signal refusal, the secondary/non-headline handling, and the downgrade of a dismissed-but-correct recommendation. This is far beyond what any structured field provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and routing are front-loaded in the first sentence, and every sentence is load-bearing (verdict logic, exposure rule, secondary handling, alternative tool). It is dense and long, however, with the parenthetical worked example pushing it toward the verbose side, so it is efficient but not maximally tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex nested-object tool with no annotations and no output schema, the description supplies the missing return semantics: the three verdict strings, the fields to surface (badRateDelta, baseline/result counts), secondary.wouldBeVerdict, and refusalCodes. An agent has enough to interpret results without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description genuinely adds meaning: it explains how refuteThreshold drives 'holding' vs 'not-holding', that the rate comparison cross-multiplies raw counts rather than badRate fields, and how exposure gates the headline verdict. It stops short of restating per-field defaults, which is fine since the schema already covers them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb (grades) plus resource (a recommendation-and-decision pair against a before/after observation log) and immediately defines the scope: same subjectId + checkKey, baseline vs result window. It is trivially distinguishable from siblings like compute_divergence, which it names as the population-level alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this to close the loop on any recommender/advisor system that logs recommendations, human accept/reject calls, and later outcomes' and names when NOT to use it: 'For grading a whole population of engine-vs-human calls at once ... use compute_divergence instead.' Both the when and the alternative are stated outright.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scaffold_agent_receiptsScaffold the agent-receipt-kit authorization patternA

agent-receipt-kit is a RUNTIME LIBRARY your own agent-orchestration code imports and calls at two specific moments (issue a WorkPacket before an agent runs, verify its AgentClaim after) -- it is not something this MCP server can 'check' on demand the way it checks a document's citations, because the packet only exists inside your application's own runtime. Call this tool to get the pattern explained, an install step, a copy-pasteable starter snippet for wiring issuePacket/verifyReceipt into your own agent loop, and (optionally) a live worked example run against the real kit -- one accepted claim, one rejected claim -- so you can see actual output before wiring it in. Use this when you're building or reviewing anything that lets an AI agent report back what it did (a coding agent, a browser-automation agent, a data-processing agent) and you don't yet trust that report by construction.

ParametersJSON Schema
NameRequiredDescriptionDefault
includeWorkedExampleNoAlso run a live accepted/rejected example against the real issuePacket/verifyReceipt, not just show a snippet.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full burden, and it does substantial work: it clarifies this is documentation/scaffold output rather than a live check, that the worked example runs against the real kit producing one accepted and one rejected claim, and that the snippet is copy-pasteable. It stops short of saying whether the worked example has side effects or resource cost, which is the one remaining behavioral gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a long single paragraph, but it is front-loaded with the most important framing (runtime library, not a server-side check) before the deliverables. A few phrases restate ('so you can see actual output before wiring it in'), which keeps it just under a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter scaffolding tool with no output schema, the description fully covers what the agent receives (explanation, install step, snippet, optional live example), when to invoke it, and how it differs from the verification siblings. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single boolean, and the schema already states it runs a live accepted/rejected example. The description's '(optionally) a live worked example ... one accepted claim, one rejected claim' largely restates that, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific deliverable (pattern explanation, install step, copy-pasteable starter snippet, optional worked example) and explicitly contrasts it with siblings: 'it is not something this MCP server can check on demand the way it checks a document's citations.' An agent can distinguish this scaffolding tool from check_grounding / corroborate_evidence without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use: 'Use this when you're building or reviewing anything that lets an AI agent report back what it did ... and you don't yet trust that report by construction,' with concrete example domains. It also implicitly excludes the runtime-check interpretation by explaining the packet only exists in the caller's own runtime.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scaffold_cost_governorScaffold the cost-governor-kit spend-safety patternA

cost-governor-kit is a RUNTIME LIBRARY your app installs and calls at the moment it's about to make an AI API call -- it is not something this MCP server can 'check' on demand the way it checks a document's citations. Its advisory check-then-commit piece (withReserveConfirm) needs a live UsageLedger backed by YOUR database and an async callback making the real call, neither of which exist as content to hand this tool. Call this to get all three pieces explained (estimated pre-call spend check, cache-aware pricing math, advisory successful-call usage counting), an install step, a copy-pasteable starter snippet, and (optionally) a live worked example: real checkPreCallCeiling calls (one allowed, one blocked), a real withReserveConfirm run against an in-memory demo ledger (one call under the limit, one over, one whose commit fails after a successful call), and a real estimateCostUsd comparison of the default 0.1x cache-read rate against a caller-supplied cacheReadPerMillion override. Use this when you're building or reviewing anything that calls a paid AI API and want a pre-call spend estimate plus usage counting that skips a call that threw -- NOT a strict concurrent limit and not proof a timed-out call was never billed by the provider (see concurrency_and_recovery in the output, and use withCapacityReservation instead if you need a real reservation).

ParametersJSON Schema
NameRequiredDescriptionDefault
includeWorkedExampleNoAlso run live examples against the real checkPreCallCeiling/withReserveConfirm, not just show a snippet.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well: it explains that the withReserveConfirm piece needs a live UsageLedger and async callback that cannot be handed to this tool, and that the worked example runs against an in-memory demo ledger (so no external data dependency). It also scopes the tool's limits (advisory only, not a hard concurrency limit, not proof of provider billing). It stops short of stating side effects of running the live example (latency/cost/credentials), keeping it just under a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The content is substantive and mostly earns its place, but it is delivered as one dense, meandering paragraph with heavy parenthetical nesting for a tool that takes one optional boolean. It is not front-loaded on the single most important fact, and readability suffers; it would be far stronger as a short lead sentence plus structured bullets.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description must carry the full load, and it does enumerate what is returned (three explained pieces, install step, snippet, optional live examples with specific cases) and names the concurrency_and_recovery section of the output. It is nearly complete, only leaving unclear whether the live examples incur latency or require credentials.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% description coverage for the single includeWorkedExample parameter, so the baseline is 3. The description restates the same idea ('optionally... a live worked example') and enumerates what the examples contain, which adds content detail but no syntax or format meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (explain the three cost-governor-kit pieces, provide an install step, a starter snippet, and optionally live worked examples) for a specific resource. It also distinguishes itself from the sibling check_* tools by stressing it is a runtime library, not an on-demand checker. The purpose is clear, though it is buried in a long first sentence rather than stated crisply up front.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit when (building or reviewing anything that calls a paid AI API and wanting a pre-call spend estimate plus usage counting that skips a call that threw). It names exclusions (NOT a strict concurrent limit, not proof a timed-out call was never billed) and routes to an alternative explicitly: use withCapacityReservation instead if you need a real reservation. This is model usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

score_trust_identifiedScore an entity from known contributorsA

Scores one entity (or one dimension of one entity -- quality, reliability, communication, ...) from signals contributed by KNOWN, identified sources: reviewer accounts, raters, inspectors, verified buyers. Weighs each signal by tier x source x proof-strength x reputation x recency decay, sums to an effective (credibility-weighted) sample size, and shrinks the result toward a domain baseline ('prior') by a configurable dial -- thin evidence stays close to the prior, deep evidence overrides it. Use this for trust/reputation scores backed by attributable evidence. For unattributed/scraped signals with no identity behind them, use assess_anonymous_authenticity instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
nowYesISO 'now' timestamp recency decay is computed against. Pass a fixed value for determinism.
dialNoShrinkage strength: a named preset, or a raw phantom-prior-signal count. Defaults to 'balanced'.
priorYesThe domain/category baseline the score shrinks toward when evidence is thin.
configNoPartial override merged over the library's illustrative EXAMPLE_IDENTIFIED_CONFIG (tiers new/standard/verified/expert). Real callers should supply their own tier/source/proof vocabulary for their domain -- omit to use the example config as-is.
signalsYesThe identified signals to score from. May be empty (yields the prior).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it does substantial work: it discloses the weighting model (tier x source x proof-strength x reputation x recency decay), the effective-sample-size computation, and the shrinkage-toward-prior behavior with thin vs. deep evidence. It also notes the empty-signals edge case yields the prior. It stops short of stating what the response actually contains or any determinism caveat (the `now` param hint lives in the schema, not here).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the verb and scope, then the mechanism, then the routing rule — a sensible order with no filler sentences. It is dense and somewhat long-winded in the mechanism sentence, but each clause carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, nested-schema tool with no output schema and no annotations, the description explains the computation well but never says what the tool returns (score, confidence band, effective sample size?). The config's confidence thresholds imply a confidence output, yet the return shape is left entirely to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds conceptual meaning the per-param descriptions do not: it explains that `dial` is the shrinkage strength and that `prior` is the baseline thin evidence shrinks toward, framing how the two required params interact. It does not add anything for `signals` item fields beyond the schema's own vocabulary guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('scores one entity or one dimension') and scopes it to signals from KNOWN, identified sources. It explicitly distinguishes itself from the sibling assess_anonymous_authenticity by naming the unattributed/scraped case, so an agent can route without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit use case ('trust/reputation scores backed by attributable evidence') and an explicit exclusion with the alternative tool named ('for unattributed/scraped signals ... use assess_anonymous_authenticity instead'). The condition that selects each tool is stated, not implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_audit_chainVerify a tamper-evident audit chainA

Independently re-verifies a hash chain from genesis: recomputes every entry's hash and confirms each entry's prevHash matches the preceding entry's entryHash. Detects a mutated payload (stored entryHash no longer matches recomputed hash) and a severed link or spliced-out middle entry (prevHash mismatch) with no options at all. Two DIFFERENT truncation attacks need two DIFFERENT options, and neither is on by default: (1) entries deleted from the end and NOT replaced -- pass expectedMinLength to catch a chain that's simply shorter than it should be (a tail deletion alone leaves no broken pointer for the walk to find); this does NOT catch truncate-and-re-append (deleting tail entries and then appending new, self-consistent ones to reach the same length again -- expectedMinLength sees the 'right' length and passes). (2) to catch truncate-and-re-append, pass anchor: an { index, entryHash } checkpoint you saved earlier from a source the chain's writer cannot edit (your own log, a separate append-only store). The chain is invalid unless the entry at anchor.index has exactly that entryHash -- this also catches a full rewrite from genesis, but entries after the anchor are only checked for internal consistency, not against anything external. Passing neither option means a truncate-and-re-append (or a full rewrite from genesis) can pass as 'valid'. Proves the chain is internally consistent and unaltered since it was hashed (and, with anchor, unaltered up to the anchor point); it does NOT prove nobody with write access ever rewrote the whole chain from genesis when no anchor is supplied (tamper-evident, not tamper-proof). Use this to audit a chain you did not produce yourself before trusting it.

ParametersJSON Schema
NameRequiredDescriptionDefault
chainYesThe chain to verify, exactly as stored/received.
anchorNoA checkpoint obtained earlier from somewhere the chain's writer cannot change (your own log, a separate append-only store) -- e.g. `{ index: chain.length - 1, entryHash: chain[chain.length - 1].entryHash }` saved right after you last trusted this chain. The chain is invalid unless it has an entry at `anchor.index` with exactly this entryHash. This is what actually catches truncate-and-re-append or a full rewrite from genesis -- `expectedMinLength` alone does not. Entries after the anchor are only checked for internal consistency, not against anything external.
expectedMinLengthNoIf you know how long the chain should be, pass it to catch entries deleted from the end and NOT replaced. Does NOT catch truncate-and-re-append (see `anchor`) -- a chain shortened and then extended back to this length with new entries passes this check.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly: it names the two truncation attacks, states that neither is caught by default, distinguishes what `expectedMinLength` versus `anchor` each catch, and explicitly bounds the guarantee ('tamper-evident, not tamper-proof', post-anchor entries only internally consistent). No behavioral trait is left implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core operation, then the option/attack analysis. It is long for a tool description and repeats the truncate-and-re-append distinction in several places, but the density is largely justified by the subtle security semantics being conveyed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers what is verified, what attacks are and are not detected, and the limits of the guarantee. The one minor gap is that it never states the exact response shape (verdict field, per-entry error reporting), though there is no output schema to lean on for that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; the description goes beyond the schema by explaining the security semantics of the optional parameters (what each attack each one catches, and the consequence of omitting both). That is real added meaning over the structured fields, though the mechanics of `anchor` are also restated in the schema itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource ('Independently re-verifies a hash chain from genesis') and immediately defines the exact operations performed (recomputes each entryHash, checks prevHash linkage). It is clearly distinguishable from the writer sibling append_audit_entry, which produces rather than checks chains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('Use this to audit a chain you did not produce yourself before trusting it') and, more importantly, tells the agent exactly which optional parameters to supply for which threat model and what silently passes if neither is given. This is unusually complete routing guidance for a verifier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 14 tool updatesv0.1.0
    • First observedappend_audit_entry
    • First observedassess_anonymous_authenticity
    • First observedcheck_claims_registry
    • First observedcheck_grounding
    • First observedcheck_mutation_invariance
    • First observedcheck_payout_invariance
    • First observedcheck_provenance_claims
    • First observedcompute_divergence
    • First observedcorroborate_evidence
    • First observedgrade_decision
    • First observedscaffold_agent_receipts
    • First observedscaffold_cost_governor
    • First observedscore_trust_identified
    • First observedverify_audit_chain

TDQS

A4.2/5.0

Scored across 14 tools

Disambiguation4/5

Most tools target clearly distinct concerns (citation grounding, evidence confidence, invariance, trust scoring, provenance, audit chain, decision grading), and each description explicitly names when to use a sibling tool instead. The one real soft spot is check_payout_invariance vs check_mutation_invariance, where the former is described as a special case of the latter, plus some conceptual overlap between check_provenance_claims and check_claims_registry, though the descriptions disambiguate both.

Naming Consistency4/5

All names follow a uniform snake_case verb_noun pattern (check_grounding, corroborate_evidence, append_audit_entry, verify_audit_chain, grade_decision, etc.). Verb choice varies (check/score/assess/grade/compute/scaffold) to reflect genuinely different actions, but the structural convention is consistent throughout.

Tool Count4/5

14 tools sits comfortably in the healthy 3-15 range, and each has a distinct, non-trivial purpose rather than being padding. The set is broad (citation checks, invariance tests, trust scoring, audit chains, decision grading, and two scaffold helpers), so it functions more as a curated honesty toolkit than a single tightly-scoped domain, but no tool feels redundant.

Completeness4/5

Coverage is broad for an 'honesty/integrity' toolkit, with natural pairs such as append_audit_entry/verify_audit_chain and score_trust_identified/assess_anonymous_authenticity covering both sides of each concern. Minor gaps exist (e.g. no dedicated truth-of-evidence checker, and grounding explicitly defers semantic entailment to an external kit), but the surface covers the stated intent without obvious dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers