honesty-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@honesty-mcpcheck this paragraph for ungrounded or forged citations"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
honesty-mcp
An MCP (Model Context Protocol) server that exposes eleven already-built, zero-runtime-dependency "honesty SDK" TypeScript libraries as tools any MCP-compatible coding agent (Claude Code, Claude Desktop, or any other MCP client) can call while building or auditing a product — grounded-citation checking, evidence corroboration grading, payout/mutation-invariance checks (scenario-tested, not a formal proof for every possible input), trust/ authenticity scoring, provenance-claim validation, a claims registry, a tamper-evident audit chain, and recommendation/decision grading plus engine-vs-human divergence, plus scaffolding for two runtime-library kits that don't fit the "check this content" shape.
This server is thin by design: every tool is a Zod input schema plus a handler that imports and calls the real, unmodified export from the wrapped kit. It does not reimplement any of the underlying logic.
Install
honesty-mcp is published on npm and wraps its eleven sibling "honesty SDK"
kits as ordinary registry dependencies (semver ranges, not local file:
paths) — npm install (or a plain npx) resolves everything from the
public registry, with nothing else to check out first.
The fastest way to run it, with no install step at all:
npx honesty-mcpOr add it to a project:
npm install honesty-mcpAdding it to an MCP client
Claude Code:
claude mcp add honesty-mcp -- npx -y honesty-mcpAny other MCP client that reads an mcpServers config (Claude Desktop's
claude_desktop_config.json uses the same shape):
{
"mcpServers": {
"honesty-mcp": {
"command": "npx",
"args": ["-y", "honesty-mcp"]
}
}
}The server runs on the stdio transport — it reads JSON-RPC requests from stdin and writes responses to stdout, logging only a one-line startup banner to stderr so stdout stays a clean JSON-RPC channel. This is the transport Claude Code, Claude Desktop, and most local/CLI MCP client integrations expect for a locally-run server; it is not an HTTP server and has no port to browse to.
Listing in the MCP registry
honesty-mcp also ships a server.json (validated against the official
schema) so it can be listed on the
official MCP registry,
under the name io.github.lkopietz3-byte/honesty-mcp. Listing is a manual,
one-time step the maintainer runs after this package is on npm — it is
not run as part of this repo's CI or npm publish. Exact commands,
from the registry's own docs
(publishing guide,
CLI reference):
# 1. Install the publisher CLI (macOS/Linux)
curl -L "https://github.com/modelcontextprotocol/registry/releases/latest/download/mcp-publisher_$(uname -s | tr '[:upper:]' '[:lower:]')_$(uname -m | sed 's/x86_64/amd64/;s/aarch64/arm64/').tar.gz" | tar xz mcp-publisher
sudo mv mcp-publisher /usr/local/bin/
# 2. Log in with GitHub (proves ownership of the io.github.lkopietz3-byte namespace)
mcp-publisher login github
# 3. Publish (validates server.json against the schema, then submits it)
mcp-publisher publishpackage.json's mcpName field (io.github.lkopietz3-byte/honesty-mcp)
must match server.json's name field exactly — that's how the registry
verifies the npm package and the registry listing belong to the same
publisher (see
package-types.mdx, "Ownership Verification").
Related MCP server: Agent Trust Stack MCP Server
Tools
9 content-checking tools' worth of kits, 12 tools (kits 1–9 — hand one of these real content or data, get back a structured verdict):
Tool | Wraps | One-line purpose |
| grounding-kit | Detects ungrounded or forged citations in AI-generated text, sentence by sentence. |
| corroboration-kit | Grades confidence in a claim from independent evidence signals — not a vote count. |
| payout-invariance-kit | Checks, for the scenarios you supply, whether a ranking engine's output changes with who pays more ( |
| mutation-invariance-kit | Checks, for the scenarios you supply, whether a decision/score/ranking function's output depends on a variable it claims not to (protected attribute, geography, price, or anything you name). Not a formal proof for every possible input. |
| trust-core ( | Scores an entity from known, identified contributors (reviewer accounts, raters, inspectors). |
| trust-core ( | Scores unattributed/scraped sentiment on how organic it looks, with a heuristic discount for two specific manipulation patterns (source concentration, suspiciously uniform sentiment) — not a fraud detector. |
| provenance-kit | Flags certainty-implying language ("(verified)", "guaranteed") not backed by an appropriate provenance tier. |
| claims-registry-kit | Buckets public-facing claims as current / stale / unverified against their linked evidence and last-verified date. |
| audit-chain-kit | Appends one entry to a hash-chained, tamper-evident audit log. |
| audit-chain-kit | Independently re-verifies a hash chain from genesis; detects mutation and severed links unconditionally, plain tail deletion with |
| advice-ledger-kit | Grades one recommendation-and-decision pair against a before/after observation log — exposure-aligned, with separate floors per window and machine-readable refusal codes. |
| advice-ledger-kit | Measures how often an engine and a human disagreed, and where a later outcome exists, reports engine-right/human-right as separate, never-blended counts. |
2 scaffolding tools (agent-receipt-kit, cost-governor-kit — these are
runtime libraries a project installs and imports into its own running
code, not something checked on-demand the same way; see the code comments
in src/tools/agentReceiptScaffold.ts / costGovernorScaffold.ts for why a
"check this content" shape would misrepresent them):
Tool | Wraps | One-line purpose |
| agent-receipt-kit | Explains the issue-a-packet / verify-the-claim authorization pattern for AI agents, with an install step, a starter snippet, and (optionally) a live accepted/rejected worked example run against the real kit. |
| cost-governor-kit | Explains the estimated pre-call spend check + cache-aware pricing math (default 0.1x cache-read rate, overridable per model via |
That's 9 kits, 12 content-checking tool slots (payout-invariance-kit, audit-chain-kit, and advice-ledger-kit each got 2 tools instead of 1, since each bundles two genuinely distinct operations — runtime check vs. static grep; append vs. verify; grading one decision vs. measuring divergence across a whole population — that read better as separate, single-purpose MCP tools than as one tool with a mode switch) plus 2 scaffolding tools = 14 tools total.
A note on the code-execution tools
check_payout_invariance (runtime mode) and check_mutation_invariance
both wrap kit functions whose real signature takes an actual JavaScript
function (a ranking function, a mutation closure) — there is no way to
represent a function as MCP JSON arguments. Both tools accept that function
as a source-code string and build the real function via the Function
constructor (src/lib/buildFunction.ts) before calling the unmodified kit
export with it. compute_divergence's optional config.groupBySource uses
the same mechanism for advice-ledger-kit's groupBy config function, but
only when a caller actually supplies it — omitting it (the common case)
needs no code execution at all, since the library's own default groups by
each pair's group field. This is equivalent to eval for that one
string. It is appropriate here because this is a local, stdio-only dev tool
a coding agent runs against its own project's code — the same trust model
as that agent running node -e, vitest run, or any other local
code-execution tool — and it is not designed to be exposed to
untrusted, remote, or adversarial input. Don't wire this server up to
accept tool arguments from anyone other than the trusted local agent
driving it.
All three of these run inside a worker_threads Worker with a bounded
timeout (src/lib/runFunctionJob.ts), not on the server's own main
thread. Earlier versions called the built function synchronously in-process,
so a while (true) {} source string would block the entire stdio server
forever — every other in-flight or future tool call along with it. Now, if
the worker doesn't finish within the timeout (default 10 seconds,
override with the HONESTY_MCP_WORKER_TIMEOUT_MS environment variable, in
milliseconds), it is forcibly terminate()d and the tool call returns a
clear timeout error instead of hanging. This is not a sandbox — the
worker has this process's full OS-level privileges (filesystem, network,
environment variables) — it only bounds time. It does not stop caller
code from reading your filesystem, making network requests, or doing
anything else this Node process itself can do. The trust model above still
applies in full: only pass code you wrote or trust, and don't expose this
server to untrusted, remote, or adversarial input.
Honest limits
This server adds no guarantee beyond what the kit it wraps already documents. Every tool description was checked against its kit's own README "Honest limits"/"Limits" section as of the sibling-kit commits this branch was built against — read that kit's README for the full detail behind any one-line tool description or table row here. If a kit's own limits change later, this server's wording can drift out of sync again; nothing here re-checks that automatically.
check_payout_invariance,check_mutation_invariance, andcompute_divergence's worker-thread timeout bounds time only, not behavior — see "A note on the code-execution tools" above. It is not a sandbox.SERVER_VERSION(src/server.ts) andpackage.json'sversionare two separate values with no automated sync. They happen to agree today; a future release could forget to bump one.
Development
To work on this server itself (rather than just use it), clone the repo and build from source:
git clone https://github.com/lkopietz3-byte/honesty-mcp.git
cd honesty-mcp
npm install # resolves the 11 wrapped kits from the npm registry
npm run build # compiles src/ -> dist/
npm start # runs dist/index.js on stdionpm run lint # eslint . --max-warnings=0
npm run typecheck # tsc --noEmit over src/ + test/
npm test # vitest — spins up a real MCP Client/Server pair
# over the SDK's InMemoryTransport and calls tools
# end-to-end (see test/*.test.ts)
npm run build # tsc -p tsconfig.build.json -- compiles src/ -> dist/
npm run smoke # npm run build first, then node dist/smoke.js --
# a standalone script that does the same real
# client/server handshake, lists all tools, and
# calls one end-to-end
npm run verify # lint + typecheck + test + build + smoke, in order
npm run audit:dependencies # npm audit --package-lock-only --include=dev
# --ignore-scripts --audit-level=lowProject layout
src/
index.ts entrypoint: connects the server to stdio
server.ts createServer() -- builds the McpServer and registers every tool
smoke.ts standalone smoke-test script (see npm run smoke)
lib/
result.ts CallToolResult helpers (jsonResult / errorResult)
buildFunction.ts turns a JS source string into a callable function
runFunctionJob.ts runs that function (or the divergence groupBySource case)
inside a worker_threads Worker with a bounded timeout
tools/
grounding.ts
corroboration.ts
payoutInvariance.ts
mutationInvariance.ts
trustIdentified.ts
trustAnonymous.ts
provenance.ts
claimsRegistry.ts
auditChain.ts (append_audit_entry + verify_audit_chain)
adviceLedger.ts (grade_decision + compute_divergence)
agentReceiptScaffold.ts
costGovernorScaffold.ts
test/
server.test.ts end-to-end tests over a real in-memory MCP client/server pair,
covering every tool's known-good and known-bad input
buildFunction.test.ts unit tests for the source-string -> function builder
runFunctionJob.test.ts unit tests for the worker-thread job runner (normal jobs,
error propagation, timeout enforcement, env var parsing)
workerTimeout.test.ts end-to-end proof that the three code-execution tools stay
bounded by the worker-thread timeout instead of hangingAvailable Tools
14 toolsappend_audit_entryAppend a tamper-evident audit-chain entryA
Appends one entry to an append-only, hash-chained audit log and returns the new (longer) chain. Each entry's hash binds its payload, index, createdAt, a fixed format tag (formatVersion, currently "audit-chain-kit/v1"), and the previous entry's hash together, so any later tampering with an earlier entry -- or an entry produced by an incompatible version of audit-chain-kit -- breaks the chain in a way verify_audit_chain will detect. Use this wherever you need a tamper-evident record of events (agent actions, approvals, state transitions) that a skeptical third party can later verify independently. This function does not persist anything itself -- store the returned chain (or just the new entry) in whatever your app already uses.
| Name | Required | Description | Default |
|---|---|---|---|
| chain | Yes | The existing chain, in order, exactly as previously returned/stored. Pass [] to start a new chain. | |
| payload | Yes | Caller data for the new entry -- any JSON-serializable value. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: it discloses the append-only hash-chain mechanism, what each entry hash binds, the fixed format tag, that tampering or an incompatible format version will break the chain in a detectable way, and crucially that the function does not persist anything itself. This is exactly the behavioral context an agent needs before invoking it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and return value, then mechanism, then usage guidance, then the important non-persistence caveat. It is somewhat long, but nearly every sentence carries useful technical information for an annotation-free tool handling a hash chain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a tamper-evident hash chain, the absence of annotations, and the lack of an output schema, the description is complete enough: it explains the return value, persistence expectations, verification path, and versioning constraints. An agent has what it needs to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents both parameters and their nested fields thoroughly. The description adds thematic explanation of how the chain and payload are bound by hashes, but does not add parameter-level syntax or format details beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb and resource: appending one entry to an append-only, hash-chained audit log and returning the new longer chain. It also names verify_audit_chain and explains how tampering or version mismatch will be detected there, which clearly distinguishes it from the verification sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit usage context: 'Use this wherever you need a tamper-evident record of events (agent actions, approvals, state transitions) that a skeptical third party can later verify independently.' It also clarifies the non-persistence behavior and what to do with the return value, though it does not explicitly state when not to use it or name a direct alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
assess_anonymous_authenticityScore how organic unattributed signals lookA
Scores how organic a corpus of UNATTRIBUTED, scraped sentiment signals (crawled mentions, imported reviews with no verifiable identity, aggregator feeds) looks, weighing a positive composite of consensus/diversity/volume/recency against a heuristic penalty for two specific, cheap manipulation patterns: evidence concentrated in a single source, and suspiciously uniform sentiment (near-maximal with near-zero variance -- the fingerprint of copy-pasted or purchased praise). This is NOT a fraud or astroturf detector: it cannot show that sentiment is fabricated or that any reviewer is fake, and a campaign that varies its wording/sentiment and spreads across several sources isn't caught by these two checks. Treat a low score as 'looks statistically unusual in a specific way worth a human look', not as a fraud finding. Use this for reviews/mentions/buzz with no identity behind them. For signals from known, identified contributors, use score_trust_identified instead.
| Name | Required | Description | Default |
|---|---|---|---|
| now | Yes | ISO 'now' timestamp recency decay is computed against. Pass a fixed value for determinism. | |
| config | No | Partial override merged over the library's illustrative EXAMPLE_ANONYMOUS_CONFIG -- omit to use the example config as-is. | |
| signals | Yes | The unattributed signals to assess. May be empty. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it explains the positive composite, the two specific penalty patterns, and the tool's limits (it cannot show fabrication and misses varied, multi-source campaigns). It does not explicitly state side-effect purity or deterministic behavior beyond the schema's now parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded, and the description is dense but largely earns its length by covering scope, interpretation, and limitations. Some sentences are long, but the complexity of the tool justifies most of the space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives rich context for a complex scoring tool, including scope, interpretation, and caveats. Because there is no output schema, it would be stronger if it specified the return score's range or shape, but it is still sufficient to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters and nested config properties thoroughly. The description adds conceptual meaning about consensus/diversity/volume/recency and the two manipulation checks, but no additional parameter-level semantics beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: it scores how organic a corpus of unattributed, scraped sentiment signals looks. It clearly distinguishes itself from the sibling score_trust_identified and explicitly states what kind of signals it handles.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance (reviews/mentions/buzz with no identity behind them) and names the alternative tool for identified contributors. It also clarifies when not to over-interpret the result, saying a low score is not a fraud finding.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_claims_registryCheck public claims for stale or missing evidenceA
Keeps public-facing product claims honest over time by checking each one against its own claimed evidence -- the SOC2-control-evidence pattern applied to marketing/product copy instead of compliance controls. Buckets every claim into 'current' (evidence present, review within policy), 'stale' (evidence present but verifiedAt is older than maxAgeDays, or unparseable), or 'unverified' (no evidenceRef at all -- this always wins over staleness, since a fresh date next to an empty reference proves nothing). Use this as a periodic 'Monday-morning' review or a CI gate on a claims registry, to catch marketing copy that drifted out of sync with what the product actually does after a refactor. Note: this only checks that a reference EXISTS and is fresh, not that the thing it points to still actually supports the claim's text -- pair with check_grounding for that.
| Name | Required | Description | Default |
|---|---|---|---|
| now | No | ISO 'now' timestamp to evaluate against. Defaults to the current time. | |
| claims | Yes | The claims to evaluate. | |
| maxAgeDays | Yes | Staleness policy: evidence older than this many days is flagged stale. 0 is valid (every claim must have been verified today or it's stale) -- the kit requires a finite number >= 0. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full burden and does so well: it defines the three buckets, the precedence rule ('unverified always wins over staleness'), and the scope limit (existence + freshness only, not grounding). It stops short of stating side effects or that it is a pure read, but the semantic disclosure is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then usage, then the caveat, and every section is relevant. The SOC2-controls analogy costs a few words but meaningfully conveys the pattern, so the size is justified for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description supplies the missing return semantics by naming the three classification outcomes and their tie-breaking. Nothing essential for correct invocation or interpretation appears to be absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and every parameter is already documented in the schema (including the empty/whitespace evidenceRef rule and the maxAgeDays semantics). The description reinforces the maxAgeDays/verifiedAt relationship in prose but adds no format or syntax detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (checking) applied to a specific resource (public-facing product claims registry) and immediately defines the three output buckets. It is clearly distinguishable from siblings, naming check_grounding as the adjacent tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames when to use it ('periodic Monday-morning review or a CI gate') and states the boundary condition and the alternative tool to pair with ('not that the thing it points to still actually supports the claim's text -- pair with check_grounding'). This is exactly the when/when-not/alternative guidance expected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_groundingCheck citation groundingA
Detects ungrounded or forged citations in AI-generated text. Splits text into sentences and classifies each one against evidence: 'grounded' (cites a marker whose evidence plausibly supports it), 'placeholder' (an honest 'TBD'/unknown gap, no fake citation), 'ungrounded' (a claim with no citation at all), or 'invalid' (cites a marker id that is missing from evidence, or whose evidence doesn't plausibly support the sentence under the default matcher -- i.e. a forged or hallucinated citation; this outranks every other status). This is a mechanical/structural check, not a truth checker: the default support check is naive substring/word-overlap matching, not semantic entailment -- it can pass a coincidental word match and can fail a genuine paraphrase, and it cannot verify that the evidence itself is true. Use this before shipping any AI-written report, summary, or answer that cites sources, to catch a model inventing or misattributing a citation. Treat any 'invalid' sentence as a hard stop; treat 'ungrounded' sentences as claims that should probably cite something but currently don't. For higher-stakes content, use grounding-kit directly with a custom supports() function (embedding-similarity or NLI-based) instead of the default matcher.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The AI-generated text to check. Citation markers use the kit's default convention `[[cite:id]]` (e.g. "The bridge opened in 1932 [[cite:source-a]]."). This tool uses the default marker/placeholder patterns; if your generator emits a different citation syntax (e.g. "[1]"), rewrite markers to `[[cite:1]]` before calling, or use grounding-kit directly with a custom markerPattern. | |
| evidence | Yes | Map of citation marker id -> the evidence text/span it claims to support. This is the closed world: a marker cited in `text` whose id is NOT a key here is flagged invalid, and a marker whose evidence text doesn't plausibly support the sentence is also flagged invalid. Pass {} if there is no evidence at all (every citation will then be invalid, and uncited claims will be 'ungrounded'). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that this is a mechanical/structural check, that the default matcher is naive substring/word-overlap rather than semantic entailment, that it can both false-positive (coincidental word match) and false-negative (genuine paraphrase), and that it cannot verify the truth of the evidence itself. It also states the priority rule ('invalid outranks every other status').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, then classification definitions, matcher caveat, and escalation path in a logical order where each sentence carries distinct information. It is long, but the length is justified by the four-status taxonomy and the non-obvious matcher limitations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description fully enumerates the return classifications and their meanings, and covers the matcher's failure modes and parameter interplay. For a 2-parameter tool with a nested evidence map and no annotations, nothing an agent needs to invoke and interpret it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema by framing `evidence` as a closed world and explaining the parameter interaction (a marker cited in `text` missing from `evidence` is invalid; empty evidence makes all citations invalid). It reinforces the `[[cite:id]]` marker convention and notes what to do with alternate syntax.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource ('Detects ungrounded or forged citations in AI-generated text') and the body enumerates the four classification outcomes. It is clearly distinguishable from siblings like corroborate_evidence or check_provenance_claims by naming the exact artifact it inspects (sentence-level citation markers vs. evidence map).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('before shipping any AI-written report, summary, or answer that cites sources') and when to escalate ('for higher-stakes content, use grounding-kit directly with a custom supports() function'). It also names the alternative tool and the condition that selects it, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_mutation_invarianceCheck whether a decision function ignores a given inputA
Checks -- for the scenarios you supply, not a formal proof for every possible input -- whether a decision, score, or ranking function's output changes depending on a variable it claims not to depend on: a protected attribute (name, inferred ethnicity/gender/age signal), geography, price, or any axis you name. Re-runs your actual function once per named mutation scenario and confirms the output is byte-identical to the unmutated baseline; a scenario whose mutation didn't actually change the input is flagged 'vacuous' rather than silently counting as a pass. This is the general form of check_payout_invariance -- use this one for hiring/lending/insurance/housing-style fairness claims or any other 'should not depend on X' claim; use check_payout_invariance specifically for the payout/commission axis (it also has a static-import-grep mode this tool doesn't need). Pass JS source for the function under test and each mutation -- this runs in a worker thread with a bounded timeout, not a sandbox, so only pass code you wrote or trust. A pass covers only the mutations you ran: it says nothing about values you didn't try, fields changed one at a time but never together, or a proxy field you never touched (a ZIP code standing in for race, a graduation year for age). Re-run this in CI whenever the function changes.
| Name | Required | Description | Default |
|---|---|---|---|
| fnSource | Yes | JS source for the pure function `(input) => output` under test, e.g. "(applicant) => scoreApplicant(applicant)". Built and run in a worker thread with a bounded timeout (default 10s, see README). | |
| baseInput | Yes | The baseline input to fn. | |
| scenarios | Yes | Named mutation scenarios to re-run fn under. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden and does so thoroughly: it re-runs the function once per scenario, confirms byte-identical output, flags vacuous mutations, runs in a worker thread with bounded timeout that is NOT a sandbox, and states the trust constraint on passed code. It also discloses coverage limitations (untried values, one-at-a-time fields, proxy fields).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and the sibling routing, and almost every sentence earns its place given the tool's complexity. The downside is a single dense paragraph stitched with em dashes that is harder to skim than a structured list would be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no output schema, it explains result semantics (byte-identical = pass, vacuous flag) and the exact scope of what a pass covers. It stops short of describing the concrete result object shape, but the caveats and constraints an agent needs to call it correctly are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds meaning beyond the schema, notably the 'vacuous' scenario flagging and the notion that a mutation must actually change the input to count. It reinforces that scenario names appear in failure output and that mutations rewrite one axis on a copy.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('checks whether ... output changes') and resource (a decision/score/ranking function's dependence on a named variable), and explicitly differentiates itself from sibling check_payout_invariance. An agent can tell this apart from siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('hiring/lending/insurance/housing-style fairness claims or any other should-not-depend-on-X claim') versus when to prefer the alternative ('use check_payout_invariance specifically for the payout/commission axis'). Also tells the agent to re-run in CI whenever the function changes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_payout_invarianceCheck whether a ranking engine ignores payoutA
Checks -- for the specific scenarios you supply, not a formal proof for every possible payout configuration -- whether a ranking/recommendation/comparison engine's output ordering changes depending on which option pays the operator more (affiliate commission, sponsored placement, referral fee). Two modes: runtime re-runs your actual ranking function under adversarial payout-mutation scenarios you name and confirms the result is byte-identical to the unmutated baseline (pass JS source for the ranking function and each mutation -- this runs in a worker thread with a bounded timeout, not a sandbox, so only pass code you wrote or trust; a scenario whose mutation didn't actually change the input is flagged 'vacuous' rather than silently counting as a pass). static-imports instead greps a set of source files for any reference to payout-related identifiers, to assert the ranking engine's code never even has payout data in scope -- no code execution needed for this mode, and it's a best-effort text/regex grep, not a real parser (it won't catch a dynamically-built import specifier or a re-export under an aliased name). Use runtime when you can call the ranking function directly; use static-imports as a cheaper, complementary check on the engine's source. A passing result means no difference in the scenarios tested, not that the function is payout-neutral in general -- write adversarial and boundary scenarios, not one easy case, and re-run this in CI whenever the ranking logic changes.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | Which check to run. | |
| files | No | [static-imports mode, required] Source files to scan for payout references. | |
| baseInput | No | [runtime mode, required] The baseline input to rankFn -- e.g. an array of candidate objects each carrying a payout/commission field. | |
| mutations | No | [runtime mode, required, non-empty] Named adversarial payout-mutation scenarios. | |
| rankFnSource | No | [runtime mode, required] JS source for a pure ranking function `(input) => result`, e.g. "(candidates) => candidates.slice().sort((a, b) => b.score - a.score)". Built and run in a worker thread with a bounded timeout (default 10s, see README). | |
| stripComments | No | [static-imports mode] Strip comments before matching, so a mention in a comment doesn't count. Default true. | |
| payoutIdentifiers | No | [static-imports mode, required, non-empty] Identifiers that must never appear in the ranking engine's source, e.g. "commission", "payout", "affiliateRate". | |
| caseInsensitiveMatch | No | [static-imports mode] Case-insensitive identifier matching. Default true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: it discloses that runtime code executes in a worker thread with a bounded timeout and is NOT a sandbox (only pass trusted code), that static-imports is a best-effort grep rather than a real parser with named blind spots, and that vacuous mutations are flagged rather than counted as passes. It also scopes what a 'pass' actually means.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded, but the body is one very dense multi-clause sentence stuffed with parentheticals and em-dash asides, which makes it hard to scan for the mode-selection rules. The information mostly earns its place, but the structure hurts retrieval rather than helping it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter, two-mode tool with no output schema, the description covers execution model, trust boundaries, mode tradeoffs, and pass semantics well. It does not describe the shape of the returned result beyond 'shown in failure output', which is the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 8 parameters (baseline 3). The description adds genuine meaning beyond it: an example shape for baseInput, guidance that mutateSource must be pure and adversarial with concrete examples, an example rankFnSource signature, and the security caveat on passing code. It stops short of enumerating defaults/mode-requirements per parameter, which the schema does cover.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (checks) and a precisely scoped resource (whether a ranking/recommendation/comparison engine's output ordering depends on operator payout), and immediately bounds the claim ('for the specific scenarios you supply, not a formal proof'). It is distinguishable from the sibling check_mutation_invariance by its explicit payout/commission focus.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes mode selection: 'Use `runtime` when you can call the ranking function directly; use `static-imports` as a cheaper, complementary check on the engine's source.' It also states the ongoing workflow expectation ('re-run this in CI whenever the ranking logic changes') and warns to write adversarial/boundary scenarios rather than one easy case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_provenance_claimsCheck claims for unbacked certainty languageA
Scans reader-facing claims for certainty-implying language ("(verified)", "independently verified", "guaranteed", "fact-checked", "100% accurate", ...) that isn't backed by an appropriate provenance tier, plus tiers that require a sourceRef but don't have one, and claims carrying an unrecognized tier. This is the exact pattern that caught ~150 false '(verified)' labels on a live site after they had already shipped -- run it on any copy, marketing page, or AI-drafted content that makes factual-sounding claims before it ships, not after. The default phrase list is a starting point drawn from that one incident, not a taxonomy -- it will miss phrases it doesn't know about (e.g. 'clinically proven', 'third-party tested'); extend certaintyPhrases for your domain. Negation detection is a fixed character window before a match, not a parser, so it can miss a negation in an earlier clause or over-suppress one further away. An empty result means every claim's certainty language (that this tool's phrase list and negation window caught) is backed by its tier.
| Name | Required | Description | Default |
|---|---|---|---|
| claims | Yes | The claims to check. | |
| caseSensitive | No | Case-sensitive phrase matching. Default false. | |
| negationWindow | No | Characters before a phrase match to scan for a negation word ('not', 'without', ...). Default 40. | |
| certaintyPhrases | No | Override the certainty-phrase vocabulary. Defaults to a generic starter list ('(verified)', 'independently verified', 'proprietary dataset', 'guaranteed', 'fact-checked', '100% accurate', ...). | |
| certaintyRequiresTier | No | Tiers strong enough to back certainty language. Default ['verified']. | |
| requireSourceRefForTiers | No | Tiers that must carry a non-empty sourceRef. Default ['verified']. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden and does so unusually well: it discloses the incident-derived provenance of the default phrase list, the phrase-list coverage gap ('will miss phrases it doesn't know about'), the negated-window limitation ('fixed character window... not a parser'), and the precise meaning of an empty result. These are exactly the caveats an agent needs to avoid over-trusting the output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core detection behavior, then limitations. Dense and every sentence carries signal, though the incident anecdote is longer than strictly required and could be tightened without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description does most of the work, covering scope, limits, and empty-result semantics. Its one gap is the shape of the returned offenses (it references 'offense messages' but never describes the result structure), which matters for a tool with no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds value beyond the schema by explaining that certaintyPhrases defaults are a starting point drawn from one incident and should be extended per domain, and by clarifying the tier semantics context for certaintyRequiresTier/requireSourceRefForTiers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('scans reader-facing claims for certainty-implying language') and enumerates the three detection classes (unbacked certainty phrases, tiers requiring a sourceRef, unrecognized tiers). This is clearly distinct from siblings like check_grounding or check_claims_registry without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong workflow context ('run it on any copy, marketing page, or AI-drafted content that makes factual-sounding claims before it ships, not after') and a domain-extension hint for certaintyPhrases. It stops short of naming a sibling alternative or the condition that would route to one, so it falls just below the top bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compute_divergenceMeasure where two judges disagree, and who was rightA
Computes how often two independent judges (e.g. an automated engine and a human override) rated the same object differently, and -- where a later outcome exists to settle it -- reports engineRight and humanRight as two SEPARATE counts, never combined into one blended accuracy number (a system can be right most of the time overall and still wrong every single time a human bothers to overrule it, and that second fact is the one worth acting on). Divergence needs BOTH a count floor (minDivergentCount) and a rate floor (minDivergentRate) before it's 'reportable' -- either alone lets noise through. Calibration (who was right) carries its own separate floor (minResolvedDivergent) and its own status, so a population can legitimately have reportable divergence and refused calibration at the same time: plenty of disagreements, not enough of them settled by a later outcome yet. Supply group on each pair (or a custom config.groupBySource) to also get a sorted per-group breakdown alongside the overall report. Use this whenever a system has both an automated judgment and a human override/correction recorded for the same objects, to find out whether overruling the system is actually earning its keep. For grading one recommendation's own before/after outcome window instead of a whole population of engine-vs-human calls, use grade_decision.
| Name | Required | Description | Default |
|---|---|---|---|
| pairs | Yes | Every occasion both judges could have spoken on. Pairs missing one judgment are still counted, just not comparable. | |
| config | No | Floors, thresholds, and optional custom grouping, every field optional with a library default (see advice-ledger-kit's DEFAULT_DIVERGENCE_CONFIG). Omit entirely to use the library's defaults. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: it discloses that engineRight/humanRight are never blended, that divergence requires BOTH a count and rate floor, that calibration has its own separate floor and status (so divergence can be reportable while calibration is refused), and that groupBySource runs in a worker thread with a bounded timeout and a code-trust caveat.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the core computation, but the body is dense with nested parentheticals and run-on sentences that carry multiple ideas each. For a complex statistical tool the length is partly justified, yet several clauses could be trimmed without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description must carry the load; it explains the computed counts, the floor semantics, the separate calibration status, and the grouping behavior. It stops short of describing the report's full shape (e.g. examples, overall structure), which is the one remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds conceptual meaning beyond the per-field text: why two separate counts exist (system can be right overall yet wrong on every overrule) and why both floors must clear together. Some of this overlaps the schema's own phrasing, keeping it below a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('computes how often two independent judges rated the same object differently') and even names the two output counts (engineRight/humanRight). It explicitly distinguishes itself from the sibling grade_decision by scope (population of engine-vs-human calls vs one recommendation's window).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger ('use this whenever a system has both an automated judgment and a human override/correction recorded for the same objects') and an explicit alternative with its own condition ('for grading one recommendation's own before/after outcome window ... use grade_decision'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
corroborate_evidenceGrade evidence corroboration for a claimA
Grades confidence in a claim from a set of independent evidence signals -- NOT a vote count. 'confirmed' requires 2+ distinct supporting sources AND at least one non-textual signal; disagreement among sources surfaces as 'mixed' rather than being averaged away; a null result is 'not-found' only under adequate coverage, otherwise 'inconclusive' (a thin sample can't prove a negative); and thin coverage caps the verdict below 'confirmed' no matter how clean the signals look. Use this whenever several pieces of evidence were gathered for a claim (by you, another tool, or a research/verification pass) and you need an honest, non-inflated verdict instead of eyeballing how many checks 'passed'.
| Name | Required | Description | Default |
|---|---|---|---|
| signals | Yes | The evidence signals gathered about the claim. May be empty. | |
| coverage | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the verdict logic (2+ distinct sources plus a non-textual signal for 'confirmed', disagreement surfacing as 'mixed', null results gated by coverage, thin coverage capping the verdict). It does not describe error behavior or the shape of the returned verdict object, leaving a modest gap for a scoring tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The key constraint ('NOT a vote count') is front-loaded and the grading rules follow in a tight, information-dense sentence. It is longer than average but nearly every clause carries a distinct rule; minor density without wasted sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description conveys the decision rules an agent needs to invoke it correctly and anticipate verdict semantics. It stops short of stating return-value structure, but the embedded verdict names give sufficient orientation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, but the description compensates by explaining how the parameters drive outcomes -- distinct `source` counting for independence, non-textual `kind` unlocking 'confirmed', and coverage capping the verdict. This adds real semantic meaning beyond the field-level schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('grades confidence in a claim from a set of independent evidence signals') and immediately differentiates itself from a naive alternative with 'NOT a vote count'. An agent can distinguish this from siblings like check_grounding or score_trust_identified without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it: 'whenever several pieces of evidence were gathered for a claim (by you, another tool, or a research/verification pass) and you need an honest, non-inflated verdict'. It gives a clear triggering context but does not name or contrast against specific sibling tools, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
grade_decisionGrade a decision against what happened after itA
Grades one recommendation-and-decision pair against a before/after observation log. Compares a pre-decision BASELINE window to a post-decision, exposure-aligned RESULT window on the same subjectId + checkKey, and returns 'holding' (the exposed post-decision window stayed under refuteThreshold bad observations AND its bad rate is not higher than the baseline's), 'not-holding' (EITHER the exposed post-decision window hit refuteThreshold bad observations, OR its bad rate is higher than the baseline's, even below that count -- 1 bad of 10 before, 1 bad of 3 exposed since is 'not-holding', not 'holding', despite the count staying under the default bar of 2), or 'refused' (the evidence didn't clear a floor -- see refusalCodes for exactly which one, never a vague 'unproven'). The rate comparison is exact -- it cross-multiplies the raw counts rather than comparing the rounded badRate fields or the sign of badRateDelta, so a not-holding can occur even when both windows display the same 3-place badRate. Always show badRateDelta and the raw baseline/result counts next to the verdict, don't quote 'holding' on its own. The verdict grades the DECISION, not just whether advice was taken: a DISMISSED recommendation whose problem later surfaced also grades 'not-holding', because the evidence sided with the advice either way. Two things this tool refuses to let slide: (1) 'nothing has gone wrong since' only counts as a result if the baseline shows the problem existed before -- otherwise it's 'baseline_lacks_negative_signal'; (2) only observations where the advice could actually have applied (exposed === true) can create or reverse the headline verdict -- everything else is reported separately as secondary (whose own wouldBeVerdict can itself be 'refused'), labelled non-headline, and can never become the headline. Use this to close the loop on any recommender/advisor system that logs recommendations, human accept/reject calls, and later outcomes, instead of letting advice quality go unmeasured forever. For grading a whole population of engine-vs-human calls at once (rather than one decision against its own before/after window), use compute_divergence instead.
| Name | Required | Description | Default |
|---|---|---|---|
| config | No | Floors and thresholds, every field optional with a library default (see advice-ledger-kit's DEFAULT_GRADE_CONFIG). Omit entirely to use the library's defaults. | |
| decision | Yes | The human's call on the recommendation. | |
| observations | Yes | Every observation available. This tool does the filtering by subjectId + checkKey itself, so passing the whole ledger (observations for other subjects/checks too) is fine. | |
| recommendation | Yes | The advice being graded. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and discharges it: it enumerates the three verdicts, the exact cross-multiplication semantics, the exposure (exposed === true) rule, the baseline_negative_signal refusal, the secondary/non-headline handling, and the downgrade of a dismissed-but-correct recommendation. This is far beyond what any structured field provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose and routing are front-loaded in the first sentence, and every sentence is load-bearing (verdict logic, exposure rule, secondary handling, alternative tool). It is dense and long, however, with the parenthetical worked example pushing it toward the verbose side, so it is efficient but not maximally tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex nested-object tool with no annotations and no output schema, the description supplies the missing return semantics: the three verdict strings, the fields to surface (badRateDelta, baseline/result counts), secondary.wouldBeVerdict, and refusalCodes. An agent has enough to interpret results without an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description genuinely adds meaning: it explains how refuteThreshold drives 'holding' vs 'not-holding', that the rate comparison cross-multiplies raw counts rather than badRate fields, and how exposure gates the headline verdict. It stops short of restating per-field defaults, which is fine since the schema already covers them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb (grades) plus resource (a recommendation-and-decision pair against a before/after observation log) and immediately defines the scope: same subjectId + checkKey, baseline vs result window. It is trivially distinguishable from siblings like compute_divergence, which it names as the population-level alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use this to close the loop on any recommender/advisor system that logs recommendations, human accept/reject calls, and later outcomes' and names when NOT to use it: 'For grading a whole population of engine-vs-human calls at once ... use compute_divergence instead.' Both the when and the alternative are stated outright.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scaffold_agent_receiptsScaffold the agent-receipt-kit authorization patternA
agent-receipt-kit is a RUNTIME LIBRARY your own agent-orchestration code imports and calls at two specific moments (issue a WorkPacket before an agent runs, verify its AgentClaim after) -- it is not something this MCP server can 'check' on demand the way it checks a document's citations, because the packet only exists inside your application's own runtime. Call this tool to get the pattern explained, an install step, a copy-pasteable starter snippet for wiring issuePacket/verifyReceipt into your own agent loop, and (optionally) a live worked example run against the real kit -- one accepted claim, one rejected claim -- so you can see actual output before wiring it in. Use this when you're building or reviewing anything that lets an AI agent report back what it did (a coding agent, a browser-automation agent, a data-processing agent) and you don't yet trust that report by construction.
| Name | Required | Description | Default |
|---|---|---|---|
| includeWorkedExample | No | Also run a live accepted/rejected example against the real issuePacket/verifyReceipt, not just show a snippet. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full burden, and it does substantial work: it clarifies this is documentation/scaffold output rather than a live check, that the worked example runs against the real kit producing one accepted and one rejected claim, and that the snippet is copy-pasteable. It stops short of saying whether the worked example has side effects or resource cost, which is the one remaining behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a long single paragraph, but it is front-loaded with the most important framing (runtime library, not a server-side check) before the deliverables. A few phrases restate ('so you can see actual output before wiring it in'), which keeps it just under a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-parameter scaffolding tool with no output schema, the description fully covers what the agent receives (explanation, install step, snippet, optional live example), when to invoke it, and how it differs from the verification siblings. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single boolean, and the schema already states it runs a live accepted/rejected example. The description's '(optionally) a live worked example ... one accepted claim, one rejected claim' largely restates that, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific deliverable (pattern explanation, install step, copy-pasteable starter snippet, optional worked example) and explicitly contrasts it with siblings: 'it is not something this MCP server can check on demand the way it checks a document's citations.' An agent can distinguish this scaffolding tool from check_grounding / corroborate_evidence without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use: 'Use this when you're building or reviewing anything that lets an AI agent report back what it did ... and you don't yet trust that report by construction,' with concrete example domains. It also implicitly excludes the runtime-check interpretation by explaining the packet only exists in the caller's own runtime.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scaffold_cost_governorScaffold the cost-governor-kit spend-safety patternA
cost-governor-kit is a RUNTIME LIBRARY your app installs and calls at the moment it's about to make an AI API call -- it is not something this MCP server can 'check' on demand the way it checks a document's citations. Its advisory check-then-commit piece (withReserveConfirm) needs a live UsageLedger backed by YOUR database and an async callback making the real call, neither of which exist as content to hand this tool. Call this to get all three pieces explained (estimated pre-call spend check, cache-aware pricing math, advisory successful-call usage counting), an install step, a copy-pasteable starter snippet, and (optionally) a live worked example: real checkPreCallCeiling calls (one allowed, one blocked), a real withReserveConfirm run against an in-memory demo ledger (one call under the limit, one over, one whose commit fails after a successful call), and a real estimateCostUsd comparison of the default 0.1x cache-read rate against a caller-supplied cacheReadPerMillion override. Use this when you're building or reviewing anything that calls a paid AI API and want a pre-call spend estimate plus usage counting that skips a call that threw -- NOT a strict concurrent limit and not proof a timed-out call was never billed by the provider (see concurrency_and_recovery in the output, and use withCapacityReservation instead if you need a real reservation).
| Name | Required | Description | Default |
|---|---|---|---|
| includeWorkedExample | No | Also run live examples against the real checkPreCallCeiling/withReserveConfirm, not just show a snippet. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: it explains that the withReserveConfirm piece needs a live UsageLedger and async callback that cannot be handed to this tool, and that the worked example runs against an in-memory demo ledger (so no external data dependency). It also scopes the tool's limits (advisory only, not a hard concurrency limit, not proof of provider billing). It stops short of stating side effects of running the live example (latency/cost/credentials), keeping it just under a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The content is substantive and mostly earns its place, but it is delivered as one dense, meandering paragraph with heavy parenthetical nesting for a tool that takes one optional boolean. It is not front-loaded on the single most important fact, and readability suffers; it would be far stronger as a short lead sentence plus structured bullets.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description must carry the full load, and it does enumerate what is returned (three explained pieces, install step, snippet, optional live examples with specific cases) and names the concurrency_and_recovery section of the output. It is nearly complete, only leaving unclear whether the live examples incur latency or require credentials.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% description coverage for the single includeWorkedExample parameter, so the baseline is 3. The description restates the same idea ('optionally... a live worked example') and enumerates what the examples contain, which adds content detail but no syntax or format meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (explain the three cost-governor-kit pieces, provide an install step, a starter snippet, and optionally live worked examples) for a specific resource. It also distinguishes itself from the sibling check_* tools by stressing it is a runtime library, not an on-demand checker. The purpose is clear, though it is buried in a long first sentence rather than stated crisply up front.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when (building or reviewing anything that calls a paid AI API and wanting a pre-call spend estimate plus usage counting that skips a call that threw). It names exclusions (NOT a strict concurrent limit, not proof a timed-out call was never billed) and routes to an alternative explicitly: use withCapacityReservation instead if you need a real reservation. This is model usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
score_trust_identifiedScore an entity from known contributorsA
Scores one entity (or one dimension of one entity -- quality, reliability, communication, ...) from signals contributed by KNOWN, identified sources: reviewer accounts, raters, inspectors, verified buyers. Weighs each signal by tier x source x proof-strength x reputation x recency decay, sums to an effective (credibility-weighted) sample size, and shrinks the result toward a domain baseline ('prior') by a configurable dial -- thin evidence stays close to the prior, deep evidence overrides it. Use this for trust/reputation scores backed by attributable evidence. For unattributed/scraped signals with no identity behind them, use assess_anonymous_authenticity instead.
| Name | Required | Description | Default |
|---|---|---|---|
| now | Yes | ISO 'now' timestamp recency decay is computed against. Pass a fixed value for determinism. | |
| dial | No | Shrinkage strength: a named preset, or a raw phantom-prior-signal count. Defaults to 'balanced'. | |
| prior | Yes | The domain/category baseline the score shrinks toward when evidence is thin. | |
| config | No | Partial override merged over the library's illustrative EXAMPLE_IDENTIFIED_CONFIG (tiers new/standard/verified/expert). Real callers should supply their own tier/source/proof vocabulary for their domain -- omit to use the example config as-is. | |
| signals | Yes | The identified signals to score from. May be empty (yields the prior). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it does substantial work: it discloses the weighting model (tier x source x proof-strength x reputation x recency decay), the effective-sample-size computation, and the shrinkage-toward-prior behavior with thin vs. deep evidence. It also notes the empty-signals edge case yields the prior. It stops short of stating what the response actually contains or any determinism caveat (the `now` param hint lives in the schema, not here).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the verb and scope, then the mechanism, then the routing rule — a sensible order with no filler sentences. It is dense and somewhat long-winded in the mechanism sentence, but each clause carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, nested-schema tool with no output schema and no annotations, the description explains the computation well but never says what the tool returns (score, confidence band, effective sample size?). The config's confidence thresholds imply a confidence output, yet the return shape is left entirely to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds conceptual meaning the per-param descriptions do not: it explains that `dial` is the shrinkage strength and that `prior` is the baseline thin evidence shrinks toward, framing how the two required params interact. It does not add anything for `signals` item fields beyond the schema's own vocabulary guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('scores one entity or one dimension') and scopes it to signals from KNOWN, identified sources. It explicitly distinguishes itself from the sibling assess_anonymous_authenticity by naming the unattributed/scraped case, so an agent can route without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use case ('trust/reputation scores backed by attributable evidence') and an explicit exclusion with the alternative tool named ('for unattributed/scraped signals ... use assess_anonymous_authenticity instead'). The condition that selects each tool is stated, not implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_audit_chainVerify a tamper-evident audit chainA
Independently re-verifies a hash chain from genesis: recomputes every entry's hash and confirms each entry's prevHash matches the preceding entry's entryHash. Detects a mutated payload (stored entryHash no longer matches recomputed hash) and a severed link or spliced-out middle entry (prevHash mismatch) with no options at all. Two DIFFERENT truncation attacks need two DIFFERENT options, and neither is on by default: (1) entries deleted from the end and NOT replaced -- pass expectedMinLength to catch a chain that's simply shorter than it should be (a tail deletion alone leaves no broken pointer for the walk to find); this does NOT catch truncate-and-re-append (deleting tail entries and then appending new, self-consistent ones to reach the same length again -- expectedMinLength sees the 'right' length and passes). (2) to catch truncate-and-re-append, pass anchor: an { index, entryHash } checkpoint you saved earlier from a source the chain's writer cannot edit (your own log, a separate append-only store). The chain is invalid unless the entry at anchor.index has exactly that entryHash -- this also catches a full rewrite from genesis, but entries after the anchor are only checked for internal consistency, not against anything external. Passing neither option means a truncate-and-re-append (or a full rewrite from genesis) can pass as 'valid'. Proves the chain is internally consistent and unaltered since it was hashed (and, with anchor, unaltered up to the anchor point); it does NOT prove nobody with write access ever rewrote the whole chain from genesis when no anchor is supplied (tamper-evident, not tamper-proof). Use this to audit a chain you did not produce yourself before trusting it.
| Name | Required | Description | Default |
|---|---|---|---|
| chain | Yes | The chain to verify, exactly as stored/received. | |
| anchor | No | A checkpoint obtained earlier from somewhere the chain's writer cannot change (your own log, a separate append-only store) -- e.g. `{ index: chain.length - 1, entryHash: chain[chain.length - 1].entryHash }` saved right after you last trusted this chain. The chain is invalid unless it has an entry at `anchor.index` with exactly this entryHash. This is what actually catches truncate-and-re-append or a full rewrite from genesis -- `expectedMinLength` alone does not. Entries after the anchor are only checked for internal consistency, not against anything external. | |
| expectedMinLength | No | If you know how long the chain should be, pass it to catch entries deleted from the end and NOT replaced. Does NOT catch truncate-and-re-append (see `anchor`) -- a chain shortened and then extended back to this length with new entries passes this check. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it names the two truncation attacks, states that neither is caught by default, distinguishes what `expectedMinLength` versus `anchor` each catch, and explicitly bounds the guarantee ('tamper-evident, not tamper-proof', post-anchor entries only internally consistent). No behavioral trait is left implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core operation, then the option/attack analysis. It is long for a tool description and repeats the truncate-and-re-append distinction in several places, but the density is largely justified by the subtle security semantics being conveyed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers what is verified, what attacks are and are not detected, and the limits of the guarantee. The one minor gap is that it never states the exact response shape (verdict field, per-entry error reporting), though there is no output schema to lean on for that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description goes beyond the schema by explaining the security semantics of the optional parameters (what each attack each one catches, and the consequence of omitting both). That is real added meaning over the structured fields, though the mechanics of `anchor` are also restated in the schema itself.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('Independently re-verifies a hash chain from genesis') and immediately defines the exact operations performed (recomputes each entryHash, checks prevHash linkage). It is clearly distinguishable from the writer sibling append_audit_entry, which produces rather than checks chains.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('Use this to audit a chain you did not produce yourself before trusting it') and, more importantly, tells the agent exactly which optional parameters to supply for which threat model and what silently passes if neither is given. This is unusually complete routing guidance for a verifier.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
14 tool updates
v0.1.0- First observed
append_audit_entry - First observed
assess_anonymous_authenticity - First observed
check_claims_registry - First observed
check_grounding - First observed
check_mutation_invariance - First observed
check_payout_invariance - First observed
check_provenance_claims - First observed
compute_divergence - First observed
corroborate_evidence - First observed
grade_decision - First observed
scaffold_agent_receipts - First observed
scaffold_cost_governor - First observed
score_trust_identified - First observed
verify_audit_chain
TDQS
Scored across 14 tools
Most tools target clearly distinct concerns (citation grounding, evidence confidence, invariance, trust scoring, provenance, audit chain, decision grading), and each description explicitly names when to use a sibling tool instead. The one real soft spot is check_payout_invariance vs check_mutation_invariance, where the former is described as a special case of the latter, plus some conceptual overlap between check_provenance_claims and check_claims_registry, though the descriptions disambiguate both.
All names follow a uniform snake_case verb_noun pattern (check_grounding, corroborate_evidence, append_audit_entry, verify_audit_chain, grade_decision, etc.). Verb choice varies (check/score/assess/grade/compute/scaffold) to reflect genuinely different actions, but the structural convention is consistent throughout.
14 tools sits comfortably in the healthy 3-15 range, and each has a distinct, non-trivial purpose rather than being padding. The set is broad (citation checks, invariance tests, trust scoring, audit chains, decision grading, and two scaffold helpers), so it functions more as a curated honesty toolkit than a single tightly-scoped domain, but no tool feels redundant.
Coverage is broad for an 'honesty/integrity' toolkit, with natural pairs such as append_audit_entry/verify_audit_chain and score_trust_identified/assess_anonymous_authenticity covering both sides of each concern. Minor gaps exist (e.g. no dedicated truth-of-evidence checker, and grounding explicitly defers semantic entailment to an external kit), but the surface covers the stated intent without obvious dead ends.
Maintenance
Related MCP Connectors
Tamper-evident proof creation and verification for AI agents via MCP, A2A, and REST.
MCP-first toolbox for agents: KV storage, auth, queue, and utility tools. Free in early access.
Hash-chained HMAC-signed audit log MCP for A2A (agent-to-agent) calls. Every tool-call, agent-ha...
Governed MCP: agent audit, provenance, deterministic checks, and receipt-backed FragGate execution.
Related MCP Servers
- AlicenseAqualityDmaintenanceTrust intelligence MCP server for AI agents. 19 tools for identity stamps, reputation scoring (0-100), agent registry, forensic audit trails, ERC-8004 bridge, and A2A passports via x402 USDC micropayments.191Apache 2.0
- AlicenseAqualityCmaintenanceProvides tools for Chain of Consciousness provenance logging and Agent Rating Protocol reputation scoring to establish trust and accountability for AI agents. It enables tamper-evident activity tracking, integrity verification, and bilateral reputation management.111Apache 2.0
- AlicenseAqualityCmaintenanceCryptographic accountability for AI agents. Ed25519-signed receipts for every MCP tool call. Constraints, chains, AI judgment, invoicing, and local dashboard included.243 npm1MIT

ejentum-mcpofficial
AlicenseAqualityDmaintenanceExposes the four Ejentum cognitive harnesses (reasoning, code, anti-deception, memory) as MCP tools any agentic client can call. Drop-in scaffolding that catches LLM failure modes like sycophancy, hallucination, and reasoning shortcuts.478 npm16MIT