Skip to main content
Glama

GroundLens

Execution verification runtime for AI systems and agents

Exact

Explicit scope

Reproducible

Signed evidence

Offline

Deterministic checks

No hidden truth claims

Pinned artifacts

Hash-chained records

No network required

License

PyPI Docs Rust Python OpenSSF Best Practices OpenSSF Scorecard REUSE status SLSA

What it is · Quick start · Architecture · How it works · Engine · Runtime · Records · MCP server · Determinism · Examples · Docs · FAQ · Roadmap

What GroundLens is

Runs locally

No telemetry

No SaaS dependency

Signed records

Reproducible

Local

No data sent

Fully self-hosted

Ed25519 + chain

Pinned

GroundLens is an execution verification runtime for AI systems and agents. It turns observable AI execution into deterministic, policy-governed evidence that can be independently verified. It provides a vendor-neutral runtime and evidence protocol for observing AI executions, evaluating claims, tool calls, actions and outcomes against composable verifiers and policies, and producing signed, reproducible evidence records.

You don't have to trust GroundLens. You can verify the record

The unit is the execution: an ordered sequence of steps an AI system or agent takes, from a model call and a retrieval to a tool call, an action with side effects and a human approval. GroundLens records each step, checks it, decides, and seals the run into a signed record anyone can verify offline. Verifying a single answer is the smallest case, a run with one claim.

  • an answer, and the claims inside it → PASS, REVIEW or FAIL

  • a tool call or an action → ALLOW, REVIEW or DENY

Where a guardrail blocks or scores an output in the moment and leaves nothing behind, GroundLens leaves signed, hash-chained evidence a third party can check without trusting you. It runs locally, needs no access to your weights, prompts or architecture, and is built for teams shipping AI answers and agents into regulated or high-stakes workflows who need proof, not a score.

Related MCP server: Sentry MCP

Verification Suite

You can verify the record yourself. The suite is how: checks anyone can run, in four parts, none of them a single "trust score".

  • Conformance — does GroundLens do what it says? Seven contracts (evidence, policy, determinism, integrity, tamper detection, offline verification, scope), PASS or FAIL each: python suite/conformance.py

  • Performance — latency, record size and throughput, as per-property percentiles: python suite/performance.py

  • Interoperability — the same record verified by the Python API, the command line and an independent from-scratch verifier; a tampered record rejected by all three: python suite/interoperability.py

  • Verifier evaluation — detection metrics for the individual verifiers, which are swappable. This measures the verifiers, not the infrastructure.

Every check traces to a public standard, RFC or regulation (Ed25519, SHA-256, canonical JSON, append-only logs, SLSA, C2PA, W3C Verifiable Credentials, EU AI Act Art. 12 and 15). See suite/STANDARDS.md. Full suite in suite/.

$ python suite/conformance.py
PASS  Evidence generation      PASS  Record integrity       PASS  Offline verification
PASS  Policy semantics         PASS  Tamper detection       PASS  Scope boundaries
PASS  Decision determinism
ALL CONTRACTS PASS

$ python suite/interoperability.py
Python API, command line and an independent verifier all accept the genuine record
all three reject a tampered record
PORTABLE AND INDEPENDENTLY VERIFIABLE

Quick start

pip install groundlens

The package installs the engine, the groundlens and glv commands, and needs no other dependency. Numbers and rules are checked out of the box; the model-based verifiers need one optional download, shown at the end.

Verify an answer

Check a model's answer against the sources it was given, under a policy, and get a record you can keep.

from groundlens import verify

question = "What is the invoice total?"
source   = "The total amount due is 10,000 dollars, payable within 30 days of receipt."
answer   = "The invoice total is 1,000 dollars, due in 30 days."

record = verify(answer, [("invoice.pdf#p1", source)], question=question)

print(record.decision)     # 'FAIL'
print(record.report())
FAIL  policy=groundlens_default_v1  record=rec_350455f44e60_4dbfea8eb79c
  c2   groundlens.numeric   contradicted   0.00   nearest in invoice.pdf#p1: '10,000 dollars'

Ten thousand is not one thousand. A similarity score would rate the right answer and the wrong one alike; the numeric verifier compares the quantities exactly and points at the source number the answer lost to. The 30 days are supported in both, so they do not appear in the report: it shows only what a reviewer needs to look at.

Every verification is sealed:

record.content_hash            # 'sha256:…' — same input, policy and bundle → same hash, any machine
record.verify()                # recompute every hash and the Ed25519 signature, offline; raises if altered
record.regulatory_mapping      # the articles this decision concerns, under the policy

Verify a run

Give GroundLens an MCP execution trace and an execution policy. It records the run as a hash-linked event log, gates it, and seals a signed run record.

from groundlens import verify_run

record = verify_run(
    "examples/run/trace.jsonl",            # an MCP session, as JSON-RPC lines
    "examples/run/execution-policy.yaml",  # the rules for what the agent may do
    run_id="run_demo",
    system="invoice-agent",
)

print(record.gate)         # 'DENY'  — the run called shell.exec, which the policy forbids
print(record.breaches)     # ()      — nothing *ran* against the policy; the call was denied, not executed
print(record.record_hash)  # 'sha256:…'  — signed and chained, like an answer record

gate is the verdict over the whole run: ALLOW, REVIEW or DENY, rolled up from the strictest step. breaches is different and narrower: it lists actions that actually executed against the policy, an action the policy forbade or one that needed a human approval that never came. Here the forbidden tool was stopped, so the run is DENY with no breach. The same thing on the command line, with the real output:

glv run verify --trace examples/run/trace.jsonl --policy examples/run/execution-policy.yaml \
  --run-id run_demo --system invoice-agent --log runs.jsonl
# exit code 1  (0 ALLOW · 3 REVIEW · 1 DENY)

glv run check runs.jsonl
# ok  1 run records, chain intact, all signatures verify

A runnable version of both is under examples/run.

Enable the model-based verifiers

The numeric and rules verifiers need nothing. The lexical, semantic and NLI verifiers need the base bundle: the multilingual encoder and a multilingual entailment model, their tokenizers, and a manifest of hashes.

groundlens bundle pull base      # downloaded once; the only command that uses the network

With the bundle installed, the lexical verifier runs (one row per content word, each anchored to the source word it was scored against), and so do the semantic verifier (sentence similarity to the nearest source) and groundlens.nli (entailment, neutral or contradiction for each statement). The download is verified against a hash pinned in the engine.

from groundlens import verify

record = verify(
    "El importe de la factura es de 10.000 euros, pagaderos en 30 días.",
    [("factura", "El importe total asciende a 10.000 euros, pagaderos en un plazo de 30 días.")],
    locale="es",
)

for e in record.evidence:
    if e.verifier_id == "groundlens.lexical":
        print(e.result, round(e.score, 2), e.source_text)
# for example — scores are a 0–1 contextual support from the frozen encoder
supported  0.93  importe
supported  0.90  pagaderos
supported  0.88  factura

The score is a contextual similarity, so the same word used differently scores lower, and the weakest anchor is what a reviewer reads first. In an isolated environment, copy the bundle directory by hand and point GROUNDLENS_BUNDLE_DIR at it.

The command line

Everything except the model-based verifiers works with the base install alone.

groundlens verify --answer answer.txt --question question.txt \
  --source "invoice.pdf#p1=invoice.txt" --policy eu_ai_act_high_risk_v1 --log records.jsonl
groundlens record verify records.jsonl        # every hash, every link, every signature
groundlens report records.jsonl --out report  # report.md, report.json, README-auditor.md
groundlens policy lint policies/eu_ai_act_high_risk_v1.yaml
groundlens bundle status                      # is the base bundle installed, where, which hash

Exit codes: 0 PASS, 1 FAIL, 2 error, 3 REVIEW. The Rust binary glv exposes the same commands and adds execution verification: glv run verify seals an agent run (exit 0 / 3 / 1 on ALLOW / REVIEW / DENY) and glv run check verifies a log of run records offline.

Full documentation, including the API reference and concept guides, is at groundlens.readthedocs.io.

Architecture

Verifying a single answer is the smallest case, a run with one claim, so one contract covers both ends of the range:

  • an answer, and the claims inside it, gets PASS, REVIEW or FAIL from verifiers and a policy;

  • a tool call or an action gets ALLOW, REVIEW or DENY from an execution policy.

Either way the run is sealed into a signed, chained record. What the record keeps of the world is hashes, not content, so it is safe to hold in a regulated place while staying independently verifiable.

GroundLens sits beside your AI system, not inside it. It observes what the system produces and does, and never sees your weights, your prompts or your internal architecture, so independent verification is possible even in a bank or a sensitive deployment.

It reads a run from what an agent already emits. An agent driving its tools speaks the Model Context Protocol (MCP); GroundLens ingests those JSON-RPC messages and turns them into a run, recording hashes of the arguments and results, never the content itself. Recording a run needs no change to how the agent is built.

The engine and runtime are a Rust workspace, wrapped for Python, with no runtime dependencies; glv is the same code as a binary. No engine or runtime crate depends on an HTTP or TLS library, and a CI job fails the build if one ever appears. The only network operation in the project is one explicit command, bundle pull, which fetches the optional model bundle. Verification never reaches the network.

For the full design, the crate-by-crate layout, the core contracts (verifier, evidence, claim, policy, record, run) and the data flow, see ARCHITECTURE.md.

How it works

A verifier produces evidence, not truth: it reports what it measured and how sure it is, and none of them decides. A policy interprets the evidence and reaches the decision. The whole chain becomes a record: the input hashes, the verifiers and model hashes that ran, the evidence, the policy and its hash, the decision, the regulatory mapping, and the hash of the previous record, sealed with an Ed25519 signature. A log of records is an audit trail you can hand over as a file.

Engine

The engine verifies an answer and the claims inside it. A verifier produces evidence; a policy turns it into PASS, REVIEW or FAIL. What an agent did, its tool calls and actions, is decided by the execution policy in the Runtime.

verifier

what it does

example

groundlens.numeric

numbers, currencies, percentages and physical units, compared exactly in base units

answer 1,000 vs source 10,000 → contradicted; 1.2 km = 1200 m; 212 °F = 100 °C; $37.35 billion = a cell 37,350 under "in millions"

groundlens.rules

your own symbolic rules, run as a verifier

rule "an APR must be a percentage" → an APR written as a bare number is contradicted

groundlens.lexical

whether each word of the answer is anchored in the sources, by contextual token similarity, reported as the weakest anchor

word pagaderos anchored to the source and scored 0.90; a word with no support scores low and surfaces first

groundlens.nli

whether a source entails, contradicts or is neutral to each statement in the answer

statement "the fee is 0.75%" against a source saying 0.50% → contradiction at high confidence

semantic.cosine

how close each statement in the answer is, in the encoder's meaning space, to the nearest source, as cosine similarity

statement paraphrasing a source scores near 1.0; one on an unrelated topic scores low and surfaces for review

groundlens.numeric and groundlens.rules are exact (bit-identical on any machine) and need no download. groundlens.lexical, groundlens.nli and semantic.cosine are reproducible (a pinned model, scores within a declared tolerance across machines) and run from the base bundle: lexical and semantic on its encoder, nli on its entailment model. groundlens bundle pull base installs all three. Similarity is not entailment, so semantic.cosine reports support but never a contradiction. Geometric (SGI, DGI) and LLM-judge verifiers are planned; every one plugs into the same contract.

Locales matter for numbers: 1.234 is one thousand in Spanish and one and a bit in English. GroundLens reads en, es, ca, de, fr, it, pt, nl and Swiss formats, knows short and long scale words, and keeps every legitimate reading of an ambiguous numeral instead of guessing. The base bundle's encoder covers about a hundred languages.

A policy is a short YAML file you control. Two policies over the same evidence can reach different decisions, and both are correct: that is where your risk appetite lives, not in the engine. The bundled eu_ai_act_high_risk_v1 maps outcomes to Art. 15(1) (accuracy and robustness), Art. 14(4)(a) (human oversight) and Art. 12(1) (record keeping) of Regulation (EU) 2024/1689; every policy has a version and a hash, and the hash goes into every record it decides. Scores from statistical verifiers drift slightly between machines, so each threshold carries a guard band, and a score inside it is REVIEW everywhere.

record = verify(answer, sources, policy="eu_ai_act_high_risk_v1")
record.decision              # 'FAIL'
record.regulatory_mapping    # [{'article': 'Art. 15(1)', ...}, {'article': 'Art. 12(1)', ...}]

Runtime

The runtime verifies an execution. It records each step of a run as an event in a hash-linked log, a model call, a retrieval, a tool request and its result, an action, a human approval, and an execution policy decides what the agent may do.

An execution policy is a short, ordered list of rules. Each rule matches a tool call or an action and carries an effect: DENY stops the step, REVIEW holds it for a human, ALLOW lets it proceed. The first rule that matches decides; when none does, the default applies, so a conservative deployment denies anything it did not explicitly allow. The gate is pure rule matching, with the same exact guarantee as the numeric verifier.

id: eu_high_risk_v1
rules:
  - id: no-shell          # a shell tool is never allowed, from any server
    match: tool
    name: shell.exec
    effect: DENY
  - id: high-risk         # any action at or above high risk needs a human
    match: risk_at_least
    risk: high
    effect: REVIEW
default: ALLOW

After a run, GroundLens audits the whole log against the policy, rolls it up to a single verdict, and flags any action that ran against it: one the policy forbade, or one that needed a human approval that never came. The verdict and the breaches go into the signed record, so an auditor can replay a run and see whether the policy was honoured.

Evidence records

Whether GroundLens checked one answer or a whole run, the result is the same kind of artefact: a signed record, chained to the one before it, that anyone can verify offline.

record.content_hash     # same input, policy and bundle → same hash, on any machine
record.verify()         # recompute every hash and the Ed25519 signature, offline
Record.verify_chain(Record.read_log("records.jsonl"))

Change one byte anywhere in a record and verification fails. Append records to a JSON Lines log and each one carries the hash of the previous one. groundlens report turns a log into a human-readable report with a one-page guide for auditors.

MCP server

The same verification is available as an MCP server, so an agent (or any Model Context Protocol client, including Claude) can call GroundLens as a tool: check an answer, gate an execution, or verify a log of records. It is a thin layer over the engine and runs over stdio.

It is an optional extra, so the base package keeps its zero dependencies:

pip install "groundlens[mcp]"
groundlens-mcp                     # runs the server over stdio

Three tools:

tool

what it does

verify_answer

verify an answer against its sources under a policy; returns the decision, the evidence and the signed record

verify_run

gate an MCP execution trace under an execution policy; returns ALLOW / REVIEW / DENY, any breaches and the run record

verify_records

verify a log of records offline: every hash, every link, every signature

Point an MCP client at the groundlens-mcp command. See the docs for a client configuration example.

Determinism

GroundLens is deterministic where it can be, and reproducible where it cannot.

exact verifiers and the execution gate use no floating point: the same input gives the same result, bit for bit, on any machine. reproducible verifiers run a pinned model in f32 on a pure-Rust inference engine, and their scores stay within a declared tolerance across machines. Anything non_deterministic, such as an LLM judge, is recorded with its model, prompt hash and settings, and decides only if the policy allows it.

This is tested, not asserted: CI runs the invoice example, with and without the model bundle, on Linux, macOS and Windows under a Turkish locale and a Pacific timezone, and compares the record hash with a committed value.

Examples

Two notebooks under examples/notebooks run in Google Colab:

  • Verify an AI answer against its sources: one example in English, German, French, Spanish and Italian, from pip install to a signed record, with a wrong number, a paraphrase and a policy change.

  • Evidence records for auditors: a log of verifications, chain verification, tamper detection, the EU AI Act mapping and the report an auditor receives.

And a shell example of a whole agent run under examples/run: a trace, an execution policy and a signed run record.

Contributions are welcome; see CONTRIBUTING.md and SECURITY.md.

groundlens.dev · Javier Marín, 2026 (javier@groundlens.dev)

Available Tools

3 tools
verify_answerB

Verify an answer against its sources under a policy and return the sealed record.

    sources: (id, text) pairs, {"id","text"} dicts, or bare strings.
    policy: a built-in name (e.g. "eu_ai_act_high_risk_v1"), a path, or YAML.
    Returns the decision (PASS/REVIEW/FAIL), the evidence, the regulatory
    mapping and the record with its content hash.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
answerYes
localeNound
policyNo
sourcesYes
questionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavior. It does describe the return value (decision, evidence, regulatory mapping, record with content hash), which is helpful. However, it does not state whether the operation is read-only, whether it stores or modifies any data, or what side effects might occur. For a verification tool, this is a notable gap, especially since the action of returning a 'sealed record' implies some immutability but not explicitly a non-destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded, stating the core action in the first sentence. It then efficiently lists input format variants and the return contents. The multi-line formatting with indentation is slightly unconventional but does not harm readability. There is minimal redundancy, and every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters (2 required) and an output schema exists, the description is moderately complete. It covers the key inputs (sources, policy) and mentions the return structure. However, it omits explanation of 'locale' and 'question', and does not provide usage context relative to sibling tools or error scenarios. The presence of an output schema lightens the need to detail return fields, but the missing parameter semantics and lack of sibling differentiation reduce completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the semantics of 'sources' (formats) and 'policy' (built-in, path, YAML). The 'answer' parameter is implicitly clear from the first sentence. However, 'locale' and 'question' are not described at all. Thus, the description covers only a portion of the parameters, leaving two parameters with no guidance beyond their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description begins with a clear, specific verb and resource: 'Verify an answer against its sources under a policy and return the sealed record.' This distinguishes it from siblings (verify_run, verify_records) by focusing on answer verification, which is a distinct operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides usage details such as acceptable formats for sources (id/text pairs, dicts, strings) and policy (built-in name, path, YAML), which implicitly guides the caller. However, it does not explicitly state when to use this tool versus the sibling tools verify_run or verify_records, nor does it mention any exclusions or alternative conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_recordsA

Verify a log of records offline: every hash, every link, every signature.

    records: the JSON Lines text of an answer-record or run-record log.
    Returns {"ok", "verified", "kind"}; fails if any record or link was altered.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
recordsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does meaningful work: it discloses the return shape ('Returns {"ok", "verified", "kind"}'), the failure mode ('fails if any record or link was altered'), and that the operation happens offline. It stops short of explicitly stating verification is non-destructive, a minor gap given 'verify' implies it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact — purpose is front-loaded in the first sentence, followed by the parameter and then the return/failure behavior. Every clause carries information an agent needs; there is no filler or restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter verification tool with an output schema present, the description covers purpose, input format, return shape, and failure behavior — nearly everything needed to call it correctly. Minor gaps like the possible values of 'kind' are left to the output schema, which is acceptable per the rubric.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it documents 'records' as 'the JSON Lines text of an answer-record or run-record log,' adding format and content meaning the schema lacks. It doesn't specify the exact structure of a valid record, but for a single string parameter the added semantics are substantial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Verify a log of records offline') with concrete scope ('every hash, every link, every signature'), so an agent can tell exactly what operation this performs. It also distinguishes this from the siblings verify_run and verify_answer by clarifying that it accepts both answer-record and run-record logs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by noting the tool accepts 'an answer-record or run-record log,' which hints it covers the domains of both siblings. However, it never names verify_run or verify_answer or gives an explicit when-to-use vs. when-not-to-use rule, leaving the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_runA

Verify an MCP execution trace under an execution policy and return the run record.

    trace: the MCP session as JSON-RPC messages (JSON Lines).
    policy: the execution policy, as YAML/JSON text or a path.
    Returns the gate (ALLOW/REVIEW/DENY), any breaches, and the signed run record.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
traceYes
policyYes
run_idYes
systemYes
started_atNo
system_versionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the return values (gate, breaches, signed run record) but does not mention potential side effects (e.g., whether it writes or stores anything), permission requirements, or error behavior. This is some behavioral context but incomplete for a tool with no annotation safety net.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is reasonably concise, with the purpose front-loaded and parameters broken into clear lines. It avoids redundant wording and communicates the key return values efficiently, though it could be tightened slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return format details are not strictly required, and the description already provides a high-level return summary. However, given the six-parameter complexity and lack of annotations, the description should explain all parameters and ideally differentiate usage from siblings. It covers the core purpose but leaves several parameters and usage guidance gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains trace (format: JSON-RPC messages as JSON Lines) and policy (format: YAML/JSON text or path), which is useful. However, it does not explain run_id, system, started_at, or system_version, leaving 4 of 6 parameters undocumented in both schema and description. This is a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool verifies an MCP execution trace against an execution policy and returns the run record with gate, breaches, and signed record. This specific verb+resource distinguishes it from sibling tools verify_answer and verify_records, which target different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by specifying it is for verifying execution traces, which gives clear context. However, it does not explicitly mention when not to use it or point to alternatives like verify_answer or verify_records, so it lacks explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv3.0.6
    • Removedfind_unsupported_words
    • Addedverify_answer
    • Addedverify_records
    • Addedverify_run
  2. 1 tool updatev0.1.0
    • First observedfind_unsupported_words

TDQS

A4/5.0

Scored across 3 tools

Disambiguation5/5

The three tools address clearly different verification targets: execution traces, answer-source pairs, and record logs. No two tools accept the same kind of input or produce the same kind of output, so an agent can select among them without ambiguity.

Naming Consistency5/5

All tool names follow the same verify_<noun> pattern with snake_case, matching the verb-object convention. The naming makes the input type immediately predictable from the tool name.

Tool Count5/5

At three tools, the surface is tightly scoped to the verification domain: run traces, answers, and record-chain integrity. Each tool covers a distinct workflow and none feels redundant.

Completeness5/5

The toolkit covers the full observed verification lifecycle: generating verified run records, generating answer records, and validating logs of those records. Policies are provided as parameters rather than requiring separate management tools, so there are no obvious dead ends.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers