Skip to main content
Glama

Groundlens: ein Korrekturleser für RAG-Antworten

Groundlens

PyPI Python License Runtime dependencies groundlens MCP server

CI OpenSSF Best Practices OpenSSF Scorecard Determinism

groundlens.dev

So funktioniert es · Installation · Schnellstart · MCP-Server · Einschränkungen · Reproduzierbarkeit

Groundlens ist ein Korrekturleser für das, was Ihr Modell schreibt. Es markiert die Wörter, die Ihre Quellen nicht stützen – und zeigt Ihnen, was jedes einzelne hätte sagen sollen. Es prüft RAG-Antworten auf Grounding und Faithfulness gegenüber den abgerufenen Quellen – die Aufgabe, für die Menschen zu Halluzinationserkennung, Zitierprüfung oder RAG-Evaluierung greifen – und unterscheidet sich darin, dass es einer prüfenden Person Markierungen und Belege liefert statt eines Urteils oder eines Scores, den man mit einem Schwellwert abgleicht.

QUESTION    What is the invoice total?
SOURCE      ...the total amount due is 10,000 dollars, payable within 30 days...
ANSWER      The invoice total is 1,000 dollars, due in 30 days.

GROUNDLENS  1,000   nothing supports this.   Closest in invoice.pdf#p1: '10,000'

Es sagt Ihnen nie, dass die Antwort falsch ist. Es sagt Ihnen, welches Wort Sie sich ansehen und welches Dokument Sie öffnen sollen. Dreißig Sekunden menschlicher Aufmerksamkeit statt fünf Minuten.

Funktionsweise

Wie Groundlens Wörter und Zahlen prüft

Groundlens vergleicht Wörter und Zahlen auf zwei verschiedene Arten:

Wörter

Zahlen

Wörter werden über die Bedeutung verankert. Die Unterstützung eines Wortes ist die höchste Kosinus-Ähnlichkeit, die es gegen irgendein Wort der Quellen erreicht, mit einem eingefrorenen Standard-Encoder – derselben Art, die Ihr Retrieval bereits verwendet.

Zahlen werden über die Arithmetik verankert. Die Zahl wird mit normalisierter Formatierung zu einem Wert geparst – 10.000, 10000, $10.000, 10 000 und (unter einer deklarierten Locale) 10.000 sind eine Zahl – und dann gegen jeden Wert in den Quellen geprüft. Die Unterstützung ist exakt 1.0 oder exakt 0.0. Ähnlichkeit darf nicht mitstimmen.

Groundlens liefert den niedrigsten Wert als Ausgabe, nicht den Durchschnitt. Jede Token-Ähnlichkeitsmetrik aggregiert über den Mittelwert, und der Mittelwert ist der Ort, an dem Einzel-Token-Fehler sterben.

Ein praktisches Beispiel: zehn ist nicht hundert

Ein abgerufenes Dokument besagt, dass der Gesamtbetrag 10.000 Dollar beträgt. Die Antwort sagt 1.000 Dollar. Ein Mensch erkennt das sofort, ohne Finanzstudium.

Embedding-Ähnlichkeit erkennt es nicht. Der Kosinus zwischen der richtigen und der falschen Antwort liegt bei etwa 0,99 – der Fehler löst sich im Vektor auf, wie ein Tintentropfen in einem Becken. Ein LLM-Judge auch nicht: Er liest auf Plausibilität, und „der Gesamtbetrag beträgt 1.000 Dollar“ ist ein völlig plausibler Satz über eine Rechnung. Ein trainierter Span-Detektor auch nicht, weil Einzelziffern-Substitutionen in seinen Trainingslabels selten sind.

Satz-Encoder organisieren Text nach Vokabular, Thema und Struktur. Nie nach Wahrheit. Eine falsche Zahl in einem korrekten Satz ist für einen Paraphrase-kollabierenden Encoder beinahe eine Paraphrase.

Auf dieser Rechnung beträgt die mittlere Unterstützung der falschen Antwort 0,79 – was gut aussieht. Der schwächste Anker ist 0,00 – das ist eine Markierung am Rand.

Betrieblicher Schwellwert

Diese Bibliothek hat keinen Standard-Schwellwert. Ein Schwellwert ist eine Eigenschaft einer Bereitstellung, nicht einer Methode. Er hängt vom Encoder, von Ihren Daten und davon ab, was ein falsch Positives im Vergleich zu einem falsch Negativen kostet. Nichts davon ist hier bekannt.

Hinter der Regel steht eine Messung. Über das von uns durchlaufene Betriebspunkt-Raster lag die beste Falsch-Positiv-Rate bei 95 Prozent Recall bei 0,65, für jeden von uns getesteten Single-Pass-Detektor, einschließlich dieses. Bei dem Recall, den eine regulierte Prüfung tatsächlich braucht, ist kein fester Schnitt in diesem Raster brauchbar. Einen auszuliefern hieße, eine Zahl auszuliefern, von der wir bereits wissen, dass sie nicht hält.

Unterstützungswerte und der schwächste Anker

Was groundlens bietet:

  • Einen Unterstützungswert pro Wort, wobei niedriger bedeutet, dass die Quellen es weniger stützen.

  • Markierungen mit Beleg: das Wort, seine Spanne, seine Unterstützung und der nächste Belegsatz, damit eine prüfende Person jede Markierung in Sekunden überprüfen kann.

  • Eine Funktion calibrate(), die einen Schnitt auf Ihren eigenen gelabelten Daten anpasst. Sie weigert sich, mit weniger als 200 gelabelten Beispielen zu laufen, weil der Schnitt darunter Rauschen ist.

Wenn Sie einen Schwellwert in Ihrer Pipeline benötigen, führen Sie calibrate() auf Ihren gelabelten Daten aus:

from groundlens import calibrate

point = calibrate(labelled, target_recall=0.95)
print(point.threshold, point.fpr, point.fpr_ci95)   # read the fpr first

calibrate() benötigt mindestens 200 gelabelte Beispiele, weil darunter ein 95%-Recall-Schwellwert aus einer Handvoll Punkten geschätzt wird.

Related MCP server: Sentry MCP

Installation

pip install groundlens              # zero runtime dependencies. Not numpy, not torch
pip install "groundlens[encoder]"   # + the reference sentence encoder
pip install "groundlens[encoder,mcp]"   # + the MCP server, for Claude Desktop and friends

Die Kerninstallation zieht überhaupt kein Paket nach sich, und ein CI-Job lässt den Build fehlschlagen, falls sich das jemals ändert. Die vorherige Version installierte etwa zwei Gigabyte Deep-Learning-Stack, bevor man irgendetwas getan hatte.

Schnellstart

from groundlens import proofread, SentenceTransformerEncoder

answer = "The invoice total is 4.75% payable within 45 days."
sources = [("policy.pdf#p3", "The rate stated in the policy is 3.90% and the term is 30 days.")]

marks = proofread(answer, sources, encoder=SentenceTransformerEncoder(), k=2)

print(marks.report())
#  4.75%   support 0.00    nearest in policy.pdf#p3: '3.90%'
#  45      support 0.00    nearest in policy.pdf#p3: '30'

Jede Markierung trägt ihren Beleg:

for anchor in marks.weakest:
    anchor.text            # '4.75%'          the word in the answer
    anchor.span            # (21, 26)         where it sits
    anchor.kind            # 'numeral'        checked by arithmetic, not meaning
    anchor.support         # 0.0              absent from the sources
    anchor.evidence_id     # 'policy.pdf#p3'  which document to open
    anchor.evidence_text   # '3.90%'          what it should have matched

Aus der Shell:

groundlens read --answer answer.txt --context policy.pdf#p3=policy.txt

MCP-Server

Derselbe Korrekturleser, in Ihrem Assistenten. Groundlens bringt einen MCP-Server mit, sodass Claude Desktop, Claude Code, Cursor, VS Code oder jeder andere MCP-Client eine Antwort gegen ihre Quellen prüfen kann, ohne das Gespräch zu verlassen. Er läuft lokal über stdio. Kein Text verlässt das System.

pip install "groundlens[encoder,mcp]"
python -m groundlens.mcp

Dann richten Sie Ihren Client darauf aus. In claude_desktop_config.json – oder der entsprechenden mcp.json in Cursor und VS Code:

{
  "mcpServers": {
    "groundlens": {
      "command": "python",
      "args": ["-m", "groundlens.mcp"]
    }
  }
}

Verwenden Sie den absoluten Pfad zu dem Python, auf dem Groundlens installiert ist, falls es nicht das auf Ihrem PATH ist: /path/to/venv/bin/python.

Das eine Tool

find_unsupported_words(answer, sources, k=4, locale="und")

answer

die zu prüfende Modellausgabe

sources

[{"id": "policy.pdf#p3", "text": "..."}]. Die id kommt in den Ergebnissen zurück, damit die Leserin weiß, welches Dokument sie öffnen soll

k

wie viele der schwächsten Anker zurückgegeben werden sollen

locale

wie diese Dokumente Zahlen schreiben. es liest 1.234 als 1234, en als 1.234, und behält beide Lesarten

Es gibt die schwächsten Anker mit ihren Belegen zurück, den Boden, die Encoder-ID und einen sha256 des Befunds:

{
  "weakest_anchors": [
    {
      "word": "4.75%",
      "support": 0.0,
      "checked_by": "arithmetic",
      "closest_in_sources": "3.90%",
      "source_id": "policy.pdf#p3",
      "notes": []
    }
  ],
  "floor": 0.0,
  "n_marked": 12,
  "encoder_id": "all-mpnet-base-v2@<revision-sha>",
  "sha256": "..."
}

Ein Tool, absichtlich. Der vorherige Server bewarb drei, und so wird aus einem Produkt drei Geschichten, bevor es jemand installiert hat.

Es gibt kein Urteil und keinen Schwellwert, hier wie überall sonst in dieser Bibliothek. Ein support von 0,00 bei einer Zahl bedeutet, dass dieser Wert in den Quellen fehlt. Bei einem Wort bedeutet es, dass kein lexikalischer Anker gefunden wurde, was in einer treuen Paraphrase normal ist. Der Server meldet die Markierungen; die Leserin entscheidet.

Der Encoder lädt beim ersten Aufruf, nicht beim Start, und das Modell wird einmal heruntergeladen (etwa 420 MB), wenn es zum ersten Mal verwendet wird.

Einschränkungen

  • Es kann berechnete Werte nicht verifizieren – „der Umsatz hat sich verdreifacht“ gegen eine Quelle, die sagt „der Umsatz stieg von 5 Mio. auf 15 Mio.“.

  • Der Wortkanal prüft, ob ein Wort von den Quellen unterstützt wird. Er prüft nicht, ob es am richtigen Ort hängt. Wenn eine Antwort „zahlbar in 30 Tagen“ über Rechnung A sagt und die 30 Tage woanders im selben Kontext zu Rechnung B gehören, ist das Wort unterstützt und es erscheint keine Markierung.

  • Es kann keine Schlussfolgerungen prüfen. Das gehört zu Entailment-Modellen.

  • Es erbt Ihr Retrieval. Wenn die Passage falsch ist, ist auch das Grounding der Antwort falsch.

  • Die Segmentierung setzt Leerzeichen-getrennte Schriften voraus und warnt eher, als dass sie so tut, als ob der Text überwiegend CJK oder Thai wäre.

Reproduzierbarkeit

  • Der Zahlenkanal ist exakt. Dezimalvergleich, fester Arithmetik-Kontext, Locale aus einem Argument und nie aus LC_ALL. Byte-für-Byte identisch auf jeder Maschine – CI beweist es auf zehn OS × Python-Kombinationen unter PYTHONHASHSEED=random und einer türkischen Locale.

  • Der lexikalische Kanal ist ein float32-Kosinus aus einer festgepinnten Encoder-Revision – nicht einem Modellnamen, weil ein stilles Neu-Upload jede Zahl ändern würde, die Sie je veröffentlicht haben. Er reproduziert auf 1e-6 über Plattformen hinweg, und die Reihenfolge der schwächsten Anker ist stabil. Er ist nicht bit-identisch zwischen x86 und Apple Silicon, und wir behaupten das auch nicht.

  • marks.sha256 deckt die Struktur und die Zahlenunterstützungen exakt ab und rundet lexikalische Unterstützungen auf sechs Dezimalstellen. Den Hash zu reproduzieren reproduziert den Befund, nicht die letzten Bits der Arithmetik.

groundlens.dev · PyPI · Retractions · Contributing · Apache-2.0

Available Tools

3 tools
verify_answerB

Verify an answer against its sources under a policy and return the sealed record.

    sources: (id, text) pairs, {"id","text"} dicts, or bare strings.
    policy: a built-in name (e.g. "eu_ai_act_high_risk_v1"), a path, or YAML.
    Returns the decision (PASS/REVIEW/FAIL), the evidence, the regulatory
    mapping and the record with its content hash.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
answerYes
localeNound
policyNo
sourcesYes
questionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavior. It does describe the return value (decision, evidence, regulatory mapping, record with content hash), which is helpful. However, it does not state whether the operation is read-only, whether it stores or modifies any data, or what side effects might occur. For a verification tool, this is a notable gap, especially since the action of returning a 'sealed record' implies some immutability but not explicitly a non-destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded, stating the core action in the first sentence. It then efficiently lists input format variants and the return contents. The multi-line formatting with indentation is slightly unconventional but does not harm readability. There is minimal redundancy, and every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters (2 required) and an output schema exists, the description is moderately complete. It covers the key inputs (sources, policy) and mentions the return structure. However, it omits explanation of 'locale' and 'question', and does not provide usage context relative to sibling tools or error scenarios. The presence of an output schema lightens the need to detail return fields, but the missing parameter semantics and lack of sibling differentiation reduce completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the semantics of 'sources' (formats) and 'policy' (built-in, path, YAML). The 'answer' parameter is implicitly clear from the first sentence. However, 'locale' and 'question' are not described at all. Thus, the description covers only a portion of the parameters, leaving two parameters with no guidance beyond their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description begins with a clear, specific verb and resource: 'Verify an answer against its sources under a policy and return the sealed record.' This distinguishes it from siblings (verify_run, verify_records) by focusing on answer verification, which is a distinct operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides usage details such as acceptable formats for sources (id/text pairs, dicts, strings) and policy (built-in name, path, YAML), which implicitly guides the caller. However, it does not explicitly state when to use this tool versus the sibling tools verify_run or verify_records, nor does it mention any exclusions or alternative conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_recordsA

Verify a log of records offline: every hash, every link, every signature.

    records: the JSON Lines text of an answer-record or run-record log.
    Returns {"ok", "verified", "kind"}; fails if any record or link was altered.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
recordsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does meaningful work: it discloses the return shape ('Returns {"ok", "verified", "kind"}'), the failure mode ('fails if any record or link was altered'), and that the operation happens offline. It stops short of explicitly stating verification is non-destructive, a minor gap given 'verify' implies it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact — purpose is front-loaded in the first sentence, followed by the parameter and then the return/failure behavior. Every clause carries information an agent needs; there is no filler or restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter verification tool with an output schema present, the description covers purpose, input format, return shape, and failure behavior — nearly everything needed to call it correctly. Minor gaps like the possible values of 'kind' are left to the output schema, which is acceptable per the rubric.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it documents 'records' as 'the JSON Lines text of an answer-record or run-record log,' adding format and content meaning the schema lacks. It doesn't specify the exact structure of a valid record, but for a single string parameter the added semantics are substantial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Verify a log of records offline') with concrete scope ('every hash, every link, every signature'), so an agent can tell exactly what operation this performs. It also distinguishes this from the siblings verify_run and verify_answer by clarifying that it accepts both answer-record and run-record logs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by noting the tool accepts 'an answer-record or run-record log,' which hints it covers the domains of both siblings. However, it never names verify_run or verify_answer or gives an explicit when-to-use vs. when-not-to-use rule, leaving the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_runA

Verify an MCP execution trace under an execution policy and return the run record.

    trace: the MCP session as JSON-RPC messages (JSON Lines).
    policy: the execution policy, as YAML/JSON text or a path.
    Returns the gate (ALLOW/REVIEW/DENY), any breaches, and the signed run record.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
traceYes
policyYes
run_idYes
systemYes
started_atNo
system_versionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the return values (gate, breaches, signed run record) but does not mention potential side effects (e.g., whether it writes or stores anything), permission requirements, or error behavior. This is some behavioral context but incomplete for a tool with no annotation safety net.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is reasonably concise, with the purpose front-loaded and parameters broken into clear lines. It avoids redundant wording and communicates the key return values efficiently, though it could be tightened slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return format details are not strictly required, and the description already provides a high-level return summary. However, given the six-parameter complexity and lack of annotations, the description should explain all parameters and ideally differentiate usage from siblings. It covers the core purpose but leaves several parameters and usage guidance gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains trace (format: JSON-RPC messages as JSON Lines) and policy (format: YAML/JSON text or path), which is useful. However, it does not explain run_id, system, started_at, or system_version, leaving 4 of 6 parameters undocumented in both schema and description. This is a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool verifies an MCP execution trace against an execution policy and returns the run record with gate, breaches, and signed record. This specific verb+resource distinguishes it from sibling tools verify_answer and verify_records, which target different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by specifying it is for verifying execution traces, which gives clear context. However, it does not explicitly mention when not to use it or point to alternatives like verify_answer or verify_records, so it lacks explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv3.0.6
    • Removedfind_unsupported_words
    • Addedverify_answer
    • Addedverify_records
    • Addedverify_run
  2. 1 tool updatev0.1.0
    • First observedfind_unsupported_words

TDQS

A4/5.0

Scored across 3 tools

Disambiguation5/5

The three tools address clearly different verification targets: execution traces, answer-source pairs, and record logs. No two tools accept the same kind of input or produce the same kind of output, so an agent can select among them without ambiguity.

Naming Consistency5/5

All tool names follow the same verify_<noun> pattern with snake_case, matching the verb-object convention. The naming makes the input type immediately predictable from the tool name.

Tool Count5/5

At three tools, the surface is tightly scoped to the verification domain: run traces, answers, and record-chain integrity. Each tool covers a distinct workflow and none feels redundant.

Completeness5/5

The toolkit covers the full observed verification lifecycle: generating verified run records, generating answer records, and validating logs of those records. Policies are provided as parameters rather than requiring separate management tools, so there are no obvious dead ends.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers