Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}

Tools

Functions exposed to the LLM to take actions

NameDescription
check_groundingA

Detects ungrounded or forged citations in AI-generated text. Splits text into sentences and classifies each one against evidence: 'grounded' (cites a marker whose evidence plausibly supports it), 'placeholder' (an honest 'TBD'/unknown gap, no fake citation), 'ungrounded' (a claim with no citation at all), or 'invalid' (cites a marker id that is missing from evidence, or whose evidence doesn't plausibly support the sentence under the default matcher -- i.e. a forged or hallucinated citation; this outranks every other status). This is a mechanical/structural check, not a truth checker: the default support check is naive substring/word-overlap matching, not semantic entailment -- it can pass a coincidental word match and can fail a genuine paraphrase, and it cannot verify that the evidence itself is true. Use this before shipping any AI-written report, summary, or answer that cites sources, to catch a model inventing or misattributing a citation. Treat any 'invalid' sentence as a hard stop; treat 'ungrounded' sentences as claims that should probably cite something but currently don't. For higher-stakes content, use grounding-kit directly with a custom supports() function (embedding-similarity or NLI-based) instead of the default matcher.

corroborate_evidenceA

Grades confidence in a claim from a set of independent evidence signals -- NOT a vote count. 'confirmed' requires 2+ distinct supporting sources AND at least one non-textual signal; disagreement among sources surfaces as 'mixed' rather than being averaged away; a null result is 'not-found' only under adequate coverage, otherwise 'inconclusive' (a thin sample can't prove a negative); and thin coverage caps the verdict below 'confirmed' no matter how clean the signals look. Use this whenever several pieces of evidence were gathered for a claim (by you, another tool, or a research/verification pass) and you need an honest, non-inflated verdict instead of eyeballing how many checks 'passed'.

check_payout_invarianceA

Checks -- for the specific scenarios you supply, not a formal proof for every possible payout configuration -- whether a ranking/recommendation/comparison engine's output ordering changes depending on which option pays the operator more (affiliate commission, sponsored placement, referral fee). Two modes: runtime re-runs your actual ranking function under adversarial payout-mutation scenarios you name and confirms the result is byte-identical to the unmutated baseline (pass JS source for the ranking function and each mutation -- this runs in a worker thread with a bounded timeout, not a sandbox, so only pass code you wrote or trust; a scenario whose mutation didn't actually change the input is flagged 'vacuous' rather than silently counting as a pass). static-imports instead greps a set of source files for any reference to payout-related identifiers, to assert the ranking engine's code never even has payout data in scope -- no code execution needed for this mode, and it's a best-effort text/regex grep, not a real parser (it won't catch a dynamically-built import specifier or a re-export under an aliased name). Use runtime when you can call the ranking function directly; use static-imports as a cheaper, complementary check on the engine's source. A passing result means no difference in the scenarios tested, not that the function is payout-neutral in general -- write adversarial and boundary scenarios, not one easy case, and re-run this in CI whenever the ranking logic changes.

check_mutation_invarianceA

Checks -- for the scenarios you supply, not a formal proof for every possible input -- whether a decision, score, or ranking function's output changes depending on a variable it claims not to depend on: a protected attribute (name, inferred ethnicity/gender/age signal), geography, price, or any axis you name. Re-runs your actual function once per named mutation scenario and confirms the output is byte-identical to the unmutated baseline; a scenario whose mutation didn't actually change the input is flagged 'vacuous' rather than silently counting as a pass. This is the general form of check_payout_invariance -- use this one for hiring/lending/insurance/housing-style fairness claims or any other 'should not depend on X' claim; use check_payout_invariance specifically for the payout/commission axis (it also has a static-import-grep mode this tool doesn't need). Pass JS source for the function under test and each mutation -- this runs in a worker thread with a bounded timeout, not a sandbox, so only pass code you wrote or trust. A pass covers only the mutations you ran: it says nothing about values you didn't try, fields changed one at a time but never together, or a proxy field you never touched (a ZIP code standing in for race, a graduation year for age). Re-run this in CI whenever the function changes.

score_trust_identifiedA

Scores one entity (or one dimension of one entity -- quality, reliability, communication, ...) from signals contributed by KNOWN, identified sources: reviewer accounts, raters, inspectors, verified buyers. Weighs each signal by tier x source x proof-strength x reputation x recency decay, sums to an effective (credibility-weighted) sample size, and shrinks the result toward a domain baseline ('prior') by a configurable dial -- thin evidence stays close to the prior, deep evidence overrides it. Use this for trust/reputation scores backed by attributable evidence. For unattributed/scraped signals with no identity behind them, use assess_anonymous_authenticity instead.

assess_anonymous_authenticityA

Scores how organic a corpus of UNATTRIBUTED, scraped sentiment signals (crawled mentions, imported reviews with no verifiable identity, aggregator feeds) looks, weighing a positive composite of consensus/diversity/volume/recency against a heuristic penalty for two specific, cheap manipulation patterns: evidence concentrated in a single source, and suspiciously uniform sentiment (near-maximal with near-zero variance -- the fingerprint of copy-pasted or purchased praise). This is NOT a fraud or astroturf detector: it cannot show that sentiment is fabricated or that any reviewer is fake, and a campaign that varies its wording/sentiment and spreads across several sources isn't caught by these two checks. Treat a low score as 'looks statistically unusual in a specific way worth a human look', not as a fraud finding. Use this for reviews/mentions/buzz with no identity behind them. For signals from known, identified contributors, use score_trust_identified instead.

check_provenance_claimsA

Scans reader-facing claims for certainty-implying language ("(verified)", "independently verified", "guaranteed", "fact-checked", "100% accurate", ...) that isn't backed by an appropriate provenance tier, plus tiers that require a sourceRef but don't have one, and claims carrying an unrecognized tier. This is the exact pattern that caught ~150 false '(verified)' labels on a live site after they had already shipped -- run it on any copy, marketing page, or AI-drafted content that makes factual-sounding claims before it ships, not after. The default phrase list is a starting point drawn from that one incident, not a taxonomy -- it will miss phrases it doesn't know about (e.g. 'clinically proven', 'third-party tested'); extend certaintyPhrases for your domain. Negation detection is a fixed character window before a match, not a parser, so it can miss a negation in an earlier clause or over-suppress one further away. An empty result means every claim's certainty language (that this tool's phrase list and negation window caught) is backed by its tier.

check_claims_registryA

Keeps public-facing product claims honest over time by checking each one against its own claimed evidence -- the SOC2-control-evidence pattern applied to marketing/product copy instead of compliance controls. Buckets every claim into 'current' (evidence present, review within policy), 'stale' (evidence present but verifiedAt is older than maxAgeDays, or unparseable), or 'unverified' (no evidenceRef at all -- this always wins over staleness, since a fresh date next to an empty reference proves nothing). Use this as a periodic 'Monday-morning' review or a CI gate on a claims registry, to catch marketing copy that drifted out of sync with what the product actually does after a refactor. Note: this only checks that a reference EXISTS and is fresh, not that the thing it points to still actually supports the claim's text -- pair with check_grounding for that.

append_audit_entryA

Appends one entry to an append-only, hash-chained audit log and returns the new (longer) chain. Each entry's hash binds its payload, index, createdAt, a fixed format tag (formatVersion, currently "audit-chain-kit/v1"), and the previous entry's hash together, so any later tampering with an earlier entry -- or an entry produced by an incompatible version of audit-chain-kit -- breaks the chain in a way verify_audit_chain will detect. Use this wherever you need a tamper-evident record of events (agent actions, approvals, state transitions) that a skeptical third party can later verify independently. This function does not persist anything itself -- store the returned chain (or just the new entry) in whatever your app already uses.

verify_audit_chainA

Independently re-verifies a hash chain from genesis: recomputes every entry's hash and confirms each entry's prevHash matches the preceding entry's entryHash. Detects a mutated payload (stored entryHash no longer matches recomputed hash) and a severed link or spliced-out middle entry (prevHash mismatch) with no options at all. Two DIFFERENT truncation attacks need two DIFFERENT options, and neither is on by default: (1) entries deleted from the end and NOT replaced -- pass expectedMinLength to catch a chain that's simply shorter than it should be (a tail deletion alone leaves no broken pointer for the walk to find); this does NOT catch truncate-and-re-append (deleting tail entries and then appending new, self-consistent ones to reach the same length again -- expectedMinLength sees the 'right' length and passes). (2) to catch truncate-and-re-append, pass anchor: an { index, entryHash } checkpoint you saved earlier from a source the chain's writer cannot edit (your own log, a separate append-only store). The chain is invalid unless the entry at anchor.index has exactly that entryHash -- this also catches a full rewrite from genesis, but entries after the anchor are only checked for internal consistency, not against anything external. Passing neither option means a truncate-and-re-append (or a full rewrite from genesis) can pass as 'valid'. Proves the chain is internally consistent and unaltered since it was hashed (and, with anchor, unaltered up to the anchor point); it does NOT prove nobody with write access ever rewrote the whole chain from genesis when no anchor is supplied (tamper-evident, not tamper-proof). Use this to audit a chain you did not produce yourself before trusting it.

grade_decisionA

Grades one recommendation-and-decision pair against a before/after observation log. Compares a pre-decision BASELINE window to a post-decision, exposure-aligned RESULT window on the same subjectId + checkKey, and returns 'holding' (the exposed post-decision window stayed under refuteThreshold bad observations AND its bad rate is not higher than the baseline's), 'not-holding' (EITHER the exposed post-decision window hit refuteThreshold bad observations, OR its bad rate is higher than the baseline's, even below that count -- 1 bad of 10 before, 1 bad of 3 exposed since is 'not-holding', not 'holding', despite the count staying under the default bar of 2), or 'refused' (the evidence didn't clear a floor -- see refusalCodes for exactly which one, never a vague 'unproven'). The rate comparison is exact -- it cross-multiplies the raw counts rather than comparing the rounded badRate fields or the sign of badRateDelta, so a not-holding can occur even when both windows display the same 3-place badRate. Always show badRateDelta and the raw baseline/result counts next to the verdict, don't quote 'holding' on its own. The verdict grades the DECISION, not just whether advice was taken: a DISMISSED recommendation whose problem later surfaced also grades 'not-holding', because the evidence sided with the advice either way. Two things this tool refuses to let slide: (1) 'nothing has gone wrong since' only counts as a result if the baseline shows the problem existed before -- otherwise it's 'baseline_lacks_negative_signal'; (2) only observations where the advice could actually have applied (exposed === true) can create or reverse the headline verdict -- everything else is reported separately as secondary (whose own wouldBeVerdict can itself be 'refused'), labelled non-headline, and can never become the headline. Use this to close the loop on any recommender/advisor system that logs recommendations, human accept/reject calls, and later outcomes, instead of letting advice quality go unmeasured forever. For grading a whole population of engine-vs-human calls at once (rather than one decision against its own before/after window), use compute_divergence instead.

compute_divergenceA

Computes how often two independent judges (e.g. an automated engine and a human override) rated the same object differently, and -- where a later outcome exists to settle it -- reports engineRight and humanRight as two SEPARATE counts, never combined into one blended accuracy number (a system can be right most of the time overall and still wrong every single time a human bothers to overrule it, and that second fact is the one worth acting on). Divergence needs BOTH a count floor (minDivergentCount) and a rate floor (minDivergentRate) before it's 'reportable' -- either alone lets noise through. Calibration (who was right) carries its own separate floor (minResolvedDivergent) and its own status, so a population can legitimately have reportable divergence and refused calibration at the same time: plenty of disagreements, not enough of them settled by a later outcome yet. Supply group on each pair (or a custom config.groupBySource) to also get a sorted per-group breakdown alongside the overall report. Use this whenever a system has both an automated judgment and a human override/correction recorded for the same objects, to find out whether overruling the system is actually earning its keep. For grading one recommendation's own before/after outcome window instead of a whole population of engine-vs-human calls, use grade_decision.

scaffold_agent_receiptsA

agent-receipt-kit is a RUNTIME LIBRARY your own agent-orchestration code imports and calls at two specific moments (issue a WorkPacket before an agent runs, verify its AgentClaim after) -- it is not something this MCP server can 'check' on demand the way it checks a document's citations, because the packet only exists inside your application's own runtime. Call this tool to get the pattern explained, an install step, a copy-pasteable starter snippet for wiring issuePacket/verifyReceipt into your own agent loop, and (optionally) a live worked example run against the real kit -- one accepted claim, one rejected claim -- so you can see actual output before wiring it in. Use this when you're building or reviewing anything that lets an AI agent report back what it did (a coding agent, a browser-automation agent, a data-processing agent) and you don't yet trust that report by construction.

scaffold_cost_governorA

cost-governor-kit is a RUNTIME LIBRARY your app installs and calls at the moment it's about to make an AI API call -- it is not something this MCP server can 'check' on demand the way it checks a document's citations. Its advisory check-then-commit piece (withReserveConfirm) needs a live UsageLedger backed by YOUR database and an async callback making the real call, neither of which exist as content to hand this tool. Call this to get all three pieces explained (estimated pre-call spend check, cache-aware pricing math, advisory successful-call usage counting), an install step, a copy-pasteable starter snippet, and (optionally) a live worked example: real checkPreCallCeiling calls (one allowed, one blocked), a real withReserveConfirm run against an in-memory demo ledger (one call under the limit, one over, one whose commit fails after a successful call), and a real estimateCostUsd comparison of the default 0.1x cache-read rate against a caller-supplied cacheReadPerMillion override. Use this when you're building or reviewing anything that calls a paid AI API and want a pre-call spend estimate plus usage counting that skips a call that threw -- NOT a strict concurrent limit and not proof a timed-out call was never billed by the provider (see concurrency_and_recovery in the output, and use withCapacityReservation instead if you need a real reservation).

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A4.2/5.0

Scored across 14 tools

Disambiguation4/5

Most tools target clearly distinct concerns (citation grounding, evidence confidence, invariance, trust scoring, provenance, audit chain, decision grading), and each description explicitly names when to use a sibling tool instead. The one real soft spot is check_payout_invariance vs check_mutation_invariance, where the former is described as a special case of the latter, plus some conceptual overlap between check_provenance_claims and check_claims_registry, though the descriptions disambiguate both.

Naming Consistency4/5

All names follow a uniform snake_case verb_noun pattern (check_grounding, corroborate_evidence, append_audit_entry, verify_audit_chain, grade_decision, etc.). Verb choice varies (check/score/assess/grade/compute/scaffold) to reflect genuinely different actions, but the structural convention is consistent throughout.

Tool Count4/5

14 tools sits comfortably in the healthy 3-15 range, and each has a distinct, non-trivial purpose rather than being padding. The set is broad (citation checks, invariance tests, trust scoring, audit chains, decision grading, and two scaffold helpers), so it functions more as a curated honesty toolkit than a single tightly-scoped domain, but no tool feels redundant.

Completeness4/5

Coverage is broad for an 'honesty/integrity' toolkit, with natural pairs such as append_audit_entry/verify_audit_chain and score_trust_identified/assess_anonymous_authenticity covering both sides of each concern. Minor gaps exist (e.g. no dedicated truth-of-evidence checker, and grounding explicitly defers semantic entailment to an external kit), but the surface covers the stated intent without obvious dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues