groundlens
Official
GroundLens: The verification and evidence layer for AI.
What it is · What you install · Quick start · Verifiers · Policies · Evidence records · Command line · Notebooks · Roadmap
What GroundLens is
GroundLens is AI verification infrastructure.
AI systems increasingly produce factual claims, recommendations and decisions that organisations need to trust. Tools exist to evaluate models, observe applications, detect hallucinations or apply individual guardrails. What they do not provide is a common verification layer in which several verification methods can be combined, governed by explicit policies, and turned into durable evidence. GroundLens provides that layer.
In plain words: you give GroundLens an AI answer and the documents it was
supposed to be based on. GroundLens runs a set of independent checks on it,
applies the rules your organisation wrote, gives you PASS, REVIEW or
FAIL, and writes the whole check into a signed record that anyone can
verify later, offline.
flowchart LR
A[answer + sources] --> B[claims] --> C[verifiers] --> D[evidence] --> E[policy] --> F[decision]
E --> G[signed evidence record]Three ideas carry the design:
A verifier produces evidence, not truth. Exact numeric checks, lexical grounding, semantic similarity, NLI, the geometric SGI and DGI indices, symbolic rules, your own verifiers and, if you allow it, an LLM judge: each one reports what it measured and how sure it is. None of them decides.
A policy interprets the evidence. A short YAML file you control says which verifiers are required, recommended, optional or forbidden, what thresholds apply, and how evidence becomes a decision. It can map each outcome to the governance or regulatory control it concerns, such as an article of the EU AI Act.
The whole chain becomes a record. Input hashes, the verifiers and model hashes that ran, the evidence, the policy and its hash, the decision, the regulatory mapping, and the hash of the previous record, sealed with an Ed25519 signature. A log of records is an audit trail you can hand over as a file.
GroundLens does not need to know how your AI system is built. It works on outputs and evidence, locally, with no network access, so independent verification is possible even in sensitive environments.
GroundLens does not try to be the best hallucination detector. It aims to be the infrastructure through which AI verification is performed, governed and evidenced.
What pip install groundlens gives you
It is worth being exact about this.
pip install groundlens installs the GroundLens engine (GLV, for
GroundLens Verification): a Rust library wrapped for Python, with no
runtime dependencies and no network access of any kind. It contains the claim extractor, the exact numeric verifier
(numbers, currencies, percentages, physical units, in several locales), the
symbolic rules verifier, the policy engine with two bundled
policies, the signed evidence records, and the groundlens command
line. Everything in this README except the lexical verifier works with that
install alone. Nothing leaves your machine.
groundlens bundle pull base is a separate, explicit step. It downloads the
base bundle (about 470 MB: the multilingual-e5-small encoder in f32,
its tokenizer and a manifest of hashes) from this repository's releases
into a per-user directory, checks it against a hash pinned in the engine,
and refuses anything else. It is the only command in the package that
opens a network connection. With the bundle installed, the lexical
verifier runs and every record names the bundle by hash. In an isolated
environment, copy the bundle directory by hand and point
GROUNDLENS_BUNDLE_DIR at it.
What it is not: not a hosted service, not a model, not a wrapper around an LLM API, and not a hallucination score you compare with a magic number.
Quick start
pip install groundlens
groundlens bundle pull base # optional: enables the lexical verifier (≈470 MB, once)from groundlens import verify
question = "What is the invoice total?"
source = "...the total amount due is 10,000 dollars, payable within 30 days..."
answer = "The invoice total is 1,000 dollars, due in 30 days."
record = verify(answer, [("invoice.pdf#p1", source)], question=question)
print(record.report())FAIL policy=groundlens_default_v1 record=rec_350455f44e60_4dbfea8eb79c
c2 groundlens.numeric contradicted 0.00 nearest in invoice.pdf#p1: '10,000 dollars'Ten is not a hundred. A similarity score would rate the right answer and the wrong one at 0.99; the numeric verifier compares the quantities exactly and points at the source passage the number lost to.
Python 3.10 or later, on Linux, macOS and Windows.
Verifiers
verifier | what it does | guarantee | in |
| numbers, currencies, percentages and physical units, compared exactly in base units: | exact, bit-identical everywhere | yes |
| your own symbolic rules (an APR must be a percentage, a date must fall inside the contract term) | exact | yes |
| whether each word of the answer is anchored in the sources, by contextual token similarity on a frozen multilingual encoder, reported as the weakest anchor rather than an average | reproducible: pinned model hash, scores within 1e-6 across machines | with the base bundle |
NLI, semantic, SGI, DGI, LLM judge | entailment, meaning, geometric grounding and model-based judgement | optional verifiers, see the roadmap | later releases |
Locales matter for numbers: 1.234 is one thousand in Spanish and one and a
bit in English. GroundLens reads en, es, ca, de, fr, it, pt,
nl and Swiss formats, knows short and long scale words, and keeps every
legitimate reading of an ambiguous numeral instead of guessing. The base
bundle's encoder covers about a hundred languages.
Policies decide, the engine only measures
A policy is a short YAML file. Two policies over the same evidence can reach different decisions, and both are correct: that is where your risk appetite lives, not in the engine.
record = verify(answer, sources, policy="eu_ai_act_high_risk_v1")
record.decision # 'FAIL'
record.regulatory_mapping # [{'article': 'Art. 15(1)', ...}, {'article': 'Art. 12(1)', ...}]The bundled eu_ai_act_high_risk_v1 policy maps outcomes to Art. 15(1)
(accuracy and robustness) and Art. 12(1) (record keeping) of Regulation
(EU) 2024/1689. Write your own with Policy.from_yaml(); every policy has a
version and a hash, and the hash goes into every record it decides.
id: acme_rag_v1
version: 1.0.0
verifiers:
required: [groundlens.numeric, groundlens.lexical]
forbidden: [llm_judge.*]
thresholds:
groundlens.lexical: { support_min: 0.60, guard_band: 0.02 }
decision:
any_contradiction_from: [groundlens.numeric, groundlens.rules.*]
unresolved_claims: REVIEWScores from statistical verifiers drift slightly between machines, so every
threshold carries a guard band: a score inside the band is REVIEW
everywhere, never PASS on one laptop and FAIL on another.
groundlens policy lint refuses a band narrower than the verifier's
declared tolerance.
Every check leaves a record
record.content_hash # same input, policy and bundle → same hash, on any machine
record.verify() # recompute every hash and the Ed25519 signature, offline
Record.verify_chain(Record.read_log("records.jsonl"))Change one byte anywhere in a record and verification fails. Append records
to a JSON Lines log and each one carries the hash of the previous one.
groundlens report turns a log into a human-readable report with a
one-page guide for auditors.
Command line
groundlens verify --answer answer.txt --question question.txt \
--source "invoice.pdf#p1=invoice.txt" --policy eu_ai_act_high_risk_v1 --log records.jsonl
groundlens record verify records.jsonl # every hash, every link, every signature
groundlens report records.jsonl --out report # report.md, report.json, README-auditor.md
groundlens policy lint policies/eu_ai_act_high_risk_v1.yaml
groundlens bundle status # is the base bundle installed, where, which hashExit codes: 0 PASS, 1 FAIL, 2 error, 3 REVIEW. The Rust binary
glv exposes the same commands.
Same input, same answer, on any machine
Each verifier declares what it guarantees. exact verifiers use no
floating point at all. reproducible verifiers run a pinned model, in
f32, on a pure-Rust inference engine, and their scores stay within a
declared tolerance. Anything non_deterministic, such as an LLM judge, is
recorded with its model, prompt hash and settings, and only decides if the
policy says so.
This is tested rather than promised: the CI runs the invoice example, with and without the lexical channel, on Linux, macOS and Windows under a Turkish locale and a Pacific timezone, and compares the record hash with a committed value.
Built in Rust, used from Python
The engine is a Rust workspace under crates/: contracts and hashing
(gl-core), text normalisation (gl-text), numerals and units
(gl-numeric), the verifiers, the policy engine, records, bundles, the
model host (gl-onnx, on tract, no
native library) and the one pipeline everything calls (gl-engine). The
Python package is a thin binding over it; glv is the same engine as a
binary. No engine crate depends on an HTTP or TLS library, and a CI job
fails the build if one ever does.
cargo build --release # engine and glv
cd python && maturin build --release # Python wheelNotebooks
Two notebooks under examples/notebooks run in
Google Colab against the published package:
Verify an AI answer against its sources: one example in English, German, French, Spanish and Italian, from
pip installto a signed record, with a wrong number, a paraphrase and a policy change.Evidence records for auditors: a log of verifications, chain verification, tamper detection, the EU AI Act mapping and the report an auditor receives.
Roadmap
The open-source engine is the adoption and trust layer. Next, in order: entailment (NLI) and semantic verifiers on the same model host, the geometric SGI and DGI verifiers, calibration tooling, and a public benchmark reporting false positive rate at 95 % recall.
The commercial product is verification at production scale: calibration, evidence packages, policies, governance, private deployment, specialised verifiers and regulatory mappings. It is built around the evidence and the policies, on top of this engine, not instead of it.
Contributions are welcome; see CONTRIBUTING.md and SECURITY.md.
groundlens.dev · PyPI · Apache-2.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/groundlens-dev/groundlens'
If you have feedback or need assistance with the MCP directory API, please join our Discord server