groundlens
OfficialGroundlens: корректор ответов RAG

Как это работает · Установка · Быстрый старт · MCP-сервер · Ограничения · Воспроизводимость
Groundlens — это корректор того, что пишет ваша модель. Он отмечает слова, которые не подтверждаются вашими источниками, и показывает, что каждое из них должно было сказать. Он проверяет ответы RAG на обоснованность и точность относительно извлечённых источников — ту задачу, для которой люди обращаются к обнаружению галлюцинаций, проверке цитат или оценке RAG, — и отличается тем, что возвращает пометки и доказательства для рецензента, а не вердикт или оценку для порога.
QUESTION What is the invoice total?
SOURCE ...the total amount due is 10,000 dollars, payable within 30 days...
ANSWER The invoice total is 1,000 dollars, due in 30 days.
GROUNDLENS 1,000 nothing supports this. Closest in invoice.pdf#p1: '10,000'Он никогда не говорит, что ответ неверен. Он говорит, на какое слово посмотреть и какой документ открыть. Тридцать секунд внимания человека вместо пяти минут.
Как это работает

Groundlens подходит к сравнению слов и чисел двумя разными способами:
Слова | Числа |
Слова привязаны к смыслу. Поддержка слова — это наибольшее косинусное сходство, которое оно достигает с любым словом источников, с использованием замороженного готового энкодера — того же типа, который уже использует ваша система поиска. | Числа привязаны к арифметике. Числительное разбирается в значение с нормализованным форматированием — |
Groundlens выдаёт наименьшую оценку, а не среднюю. Каждая метрика сходства токенов агрегируется по среднему, и именно в среднем одиночные ошибки токенов умирают.
Практический пример: десять — это не сто
Извлечённый документ говорит, что общая сумма к оплате составляет 10 000 долларов. Ответ говорит 1 000 долларов. Человек замечает это мгновенно, без финансового образования.
Сходство эмбеддингов — нет. Косинус между правильным и неправильным ответом составляет около 0,99 — ошибка растворяется в векторе, как капля чернил в бассейне. Судья LLM тоже не замечает: он читает на правдоподобие, и «общая сумма составляет 1 000 долларов» — вполне правдоподобное предложение о счёте. Обученный детектор фрагментов тоже не замечает, потому что замены одной цифры редки в его обучающих метках.
Кодировщики предложений организуют текст по словарю, теме и структуре. Никогда по истине. Неправильное число внутри правильного предложения для энкодера, схлопывающего перефразирования, почти перефразирование.
В этом счёте средняя поддержка неправильного ответа составляет 0,79 — что выглядит нормально. Самое слабое якорное значение — 0,00 — это пометка на полях.
Эксплуатационный порог
В этой библиотеке нет порога по умолчанию. Порог — это свойство развёртывания, а не метода. Он зависит от энкодера, от ваших данных и от того, сколько вам стоит ложноположительный результат по сравнению с ложноотрицательным. Ничего из этого здесь неизвестно.
За правилом стоит измерение. По сетке рабочих точек, которую мы прогнали, лучшая частота ложноположительных результатов при 95% полноте составила 0,65 для каждого протестированного нами однопроходного детектора, включая этот. При полноте, которая действительно нужна регулируемой проверке, ни одно фиксированное отсечение в этой сетке не применимо. Выпустить его означало бы выпустить число, которое, как мы уже знаем, не выполняется.

Что предоставляет groundlens:
Оценка поддержки для каждого слова, где меньшее значение означает меньшую поддержку источниками.
Пометки с подтверждениями: слово, его диапазон, его поддержка и ближайшее предложение-доказательство, чтобы рецензент мог проверить любой вызов за секунды.
Функция
calibrate(), которая подбирает отсечение на ваших собственных размеченных данных. Она отказывается работать на менее чем 200 размеченных примерах, потому что ниже этого отсечение — это шум.
Если вам нужен порог в вашем конвейере, запустите calibrate() на ваших размеченных данных:
from groundlens import calibrate
point = calibrate(labelled, target_recall=0.95)
print(point.threshold, point.fpr, point.fpr_ci95) # read the fpr first
calibrate()требует как минимум 200 размеченных примеров, потому что ниже этого порог 95% полноты оценивается по нескольким точкам.
Related MCP server: Sentry MCP
Установка
pip install groundlens # zero runtime dependencies. Not numpy, not torch
pip install "groundlens[encoder]" # + the reference sentence encoder
pip install "groundlens[encoder,mcp]" # + the MCP server, for Claude Desktop and friendsОсновная установка не тянет ни одного пакета, и задание CI ломает сборку, если это когда-либо изменится. Предыдущая версия устанавливала примерно два гигабайта стека глубокого обучения, прежде чем вы что-либо сделали.
Быстрый старт
from groundlens import proofread, SentenceTransformerEncoder
answer = "The invoice total is 4.75% payable within 45 days."
sources = [("policy.pdf#p3", "The rate stated in the policy is 3.90% and the term is 30 days.")]
marks = proofread(answer, sources, encoder=SentenceTransformerEncoder(), k=2)
print(marks.report())
# 4.75% support 0.00 nearest in policy.pdf#p3: '3.90%'
# 45 support 0.00 nearest in policy.pdf#p3: '30'Каждая пометка несёт своё подтверждение:
for anchor in marks.weakest:
anchor.text # '4.75%' the word in the answer
anchor.span # (21, 26) where it sits
anchor.kind # 'numeral' checked by arithmetic, not meaning
anchor.support # 0.0 absent from the sources
anchor.evidence_id # 'policy.pdf#p3' which document to open
anchor.evidence_text # '3.90%' what it should have matchedИз оболочки:
groundlens read --answer answer.txt --context policy.pdf#p3=policy.txtMCP-сервер
Тот же корректор внутри вашего ассистента. Groundlens поставляется с MCP-сервером, поэтому Claude Desktop, Claude Code, Cursor, VS Code или любой другой MCP-клиент может проверить ответ по его источникам, не покидая разговор. Он работает локально через stdio. Никакой текст никуда не отправляется.
pip install "groundlens[encoder,mcp]"
python -m groundlens.mcpЗатем укажите вашему клиенту на него. В claude_desktop_config.json — или эквивалентном mcp.json в Cursor и VS Code:
{
"mcpServers": {
"groundlens": {
"command": "python",
"args": ["-m", "groundlens.mcp"]
}
}
}Используйте абсолютный путь к Python, на котором установлен Groundlens, если это не тот, что в вашем PATH: /path/to/venv/bin/python.
Единственный инструмент
find_unsupported_words(answer, sources, k=4, locale="und")
| вывод модели для проверки |
|
|
| сколько самых слабых якорей вернуть |
| как эти документы записывают числа. |
Он возвращает самые слабые якоря с их подтверждениями, нижнюю границу, идентификатор энкодера и sha256 результата:
{
"weakest_anchors": [
{
"word": "4.75%",
"support": 0.0,
"checked_by": "arithmetic",
"closest_in_sources": "3.90%",
"source_id": "policy.pdf#p3",
"notes": []
}
],
"floor": 0.0,
"n_marked": 12,
"encoder_id": "all-mpnet-base-v2@<revision-sha>",
"sha256": "..."
}Один инструмент, намеренно. Предыдущий сервер рекламировал три, и именно так один продукт превращается в три истории, прежде чем кто-либо его установил.
Здесь нет вердикта и порога, как и везде в этой библиотеке. support 0.00 для числа означает, что это значение отсутствует в источниках. Для слова это означает, что лексический якорь не найден, что обычно для точного перефразирования. Сервер сообщает пометки; читатель решает.
Энкодер загружается при первом вызове, а не при запуске, и модель загружается один раз (около 420 МБ) при первом использовании.
Ограничения
Он не может проверить вычисленные значения — «выручка утроилась» против источника, говорящего «выручка выросла с 5M до 15M».
Канал слов проверяет, поддерживается ли слово источниками. Он не проверяет, привязано ли оно к правильному объекту. Если ответ говорит «оплата в течение 30 дней» о счёте A, а 30 дней относятся к счёту B в другом месте того же контекста, слово поддерживается, и пометка не появляется.
Он не может проверить рассуждения. Это относится к моделям энтайлмента.
Он наследует ваш поиск. Если фрагмент неверен, то и обоснованность ответа неверна.
Сегментация предполагает скрипты с разделителями-пробелами и предупреждает, а не притворяется, когда текст в основном на CJK или тайском.
Воспроизводимость
Числовой канал точен. Десятичное сравнение, фиксированный арифметический контекст, локаль из аргумента и никогда из
LC_ALL. Побайтово идентичен на любой машине — CI доказывает это на десяти комбинациях ОС × Python подPYTHONHASHSEED=randomи турецкой локалью.Лексический канал — это косинус float32 от закреплённой ревизии энкодера — не имени модели, потому что тихая повторная загрузка изменила бы каждое число, которое вы когда-либо публиковали. Он воспроизводится до 1e-6 на разных платформах, и порядок самых слабых якорей стабилен. Он не побитово идентичен между x86 и Apple Silicon, и мы не утверждаем, что это так.
marks.sha256покрывает структуру и числовые поддержки точно, а лексические поддержки округляет до шести десятичных знаков. Воспроизведение хэша воспроизводит результат, а не последние биты арифметики.
groundlens.dev · PyPI · Опровержения · Вклад · Apache-2.0
Available Tools
3 toolsverify_answerB
Verify an answer against its sources under a policy and return the sealed record.
sources: (id, text) pairs, {"id","text"} dicts, or bare strings.
policy: a built-in name (e.g. "eu_ai_act_high_risk_v1"), a path, or YAML.
Returns the decision (PASS/REVIEW/FAIL), the evidence, the regulatory
mapping and the record with its content hash.
| Name | Required | Description | Default |
|---|---|---|---|
| answer | Yes | ||
| locale | No | und | |
| policy | No | ||
| sources | Yes | ||
| question | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavior. It does describe the return value (decision, evidence, regulatory mapping, record with content hash), which is helpful. However, it does not state whether the operation is read-only, whether it stores or modifies any data, or what side effects might occur. For a verification tool, this is a notable gap, especially since the action of returning a 'sealed record' implies some immutability but not explicitly a non-destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded, stating the core action in the first sentence. It then efficiently lists input format variants and the return contents. The multi-line formatting with indentation is slightly unconventional but does not harm readability. There is minimal redundancy, and every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 5 parameters (2 required) and an output schema exists, the description is moderately complete. It covers the key inputs (sources, policy) and mentions the return structure. However, it omits explanation of 'locale' and 'question', and does not provide usage context relative to sibling tools or error scenarios. The presence of an output schema lightens the need to detail return fields, but the missing parameter semantics and lack of sibling differentiation reduce completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the semantics of 'sources' (formats) and 'policy' (built-in, path, YAML). The 'answer' parameter is implicitly clear from the first sentence. However, 'locale' and 'question' are not described at all. Thus, the description covers only a portion of the parameters, leaving two parameters with no guidance beyond their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a clear, specific verb and resource: 'Verify an answer against its sources under a policy and return the sealed record.' This distinguishes it from siblings (verify_run, verify_records) by focusing on answer verification, which is a distinct operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides usage details such as acceptable formats for sources (id/text pairs, dicts, strings) and policy (built-in name, path, YAML), which implicitly guides the caller. However, it does not explicitly state when to use this tool versus the sibling tools verify_run or verify_records, nor does it mention any exclusions or alternative conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_recordsA
Verify a log of records offline: every hash, every link, every signature.
records: the JSON Lines text of an answer-record or run-record log.
Returns {"ok", "verified", "kind"}; fails if any record or link was altered.
| Name | Required | Description | Default |
|---|---|---|---|
| records | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does meaningful work: it discloses the return shape ('Returns {"ok", "verified", "kind"}'), the failure mode ('fails if any record or link was altered'), and that the operation happens offline. It stops short of explicitly stating verification is non-destructive, a minor gap given 'verify' implies it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact — purpose is front-loaded in the first sentence, followed by the parameter and then the return/failure behavior. Every clause carries information an agent needs; there is no filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter verification tool with an output schema present, the description covers purpose, input format, return shape, and failure behavior — nearly everything needed to call it correctly. Minor gaps like the possible values of 'kind' are left to the output schema, which is acceptable per the rubric.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: it documents 'records' as 'the JSON Lines text of an answer-record or run-record log,' adding format and content meaning the schema lacks. It doesn't specify the exact structure of a valid record, but for a single string parameter the added semantics are substantial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Verify a log of records offline') with concrete scope ('every hash, every link, every signature'), so an agent can tell exactly what operation this performs. It also distinguishes this from the siblings verify_run and verify_answer by clarifying that it accepts both answer-record and run-record logs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by noting the tool accepts 'an answer-record or run-record log,' which hints it covers the domains of both siblings. However, it never names verify_run or verify_answer or gives an explicit when-to-use vs. when-not-to-use rule, leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_runA
Verify an MCP execution trace under an execution policy and return the run record.
trace: the MCP session as JSON-RPC messages (JSON Lines).
policy: the execution policy, as YAML/JSON text or a path.
Returns the gate (ALLOW/REVIEW/DENY), any breaches, and the signed run record.
| Name | Required | Description | Default |
|---|---|---|---|
| trace | Yes | ||
| policy | Yes | ||
| run_id | Yes | ||
| system | Yes | ||
| started_at | No | ||
| system_version | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return values (gate, breaches, signed run record) but does not mention potential side effects (e.g., whether it writes or stores anything), permission requirements, or error behavior. This is some behavioral context but incomplete for a tool with no annotation safety net.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is reasonably concise, with the purpose front-loaded and parameters broken into clear lines. It avoids redundant wording and communicates the key return values efficiently, though it could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return format details are not strictly required, and the description already provides a high-level return summary. However, given the six-parameter complexity and lack of annotations, the description should explain all parameters and ideally differentiate usage from siblings. It covers the core purpose but leaves several parameters and usage guidance gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains trace (format: JSON-RPC messages as JSON Lines) and policy (format: YAML/JSON text or path), which is useful. However, it does not explain run_id, system, started_at, or system_version, leaving 4 of 6 parameters undocumented in both schema and description. This is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool verifies an MCP execution trace against an execution policy and returns the run record with gate, breaches, and signed record. This specific verb+resource distinguishes it from sibling tools verify_answer and verify_records, which target different resources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by specifying it is for verifying execution traces, which gives clear context. However, it does not explicitly mention when not to use it or point to alternatives like verify_answer or verify_records, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v3.0.6- Removed
find_unsupported_words - Added
verify_answer - Added
verify_records - Added
verify_run
1 tool update
v0.1.0- First observed
find_unsupported_words
TDQS
Scored across 3 tools
The three tools address clearly different verification targets: execution traces, answer-source pairs, and record logs. No two tools accept the same kind of input or produce the same kind of output, so an agent can select among them without ambiguity.
All tool names follow the same verify_<noun> pattern with snake_case, matching the verb-object convention. The naming makes the input type immediately predictable from the tool name.
At three tools, the surface is tightly scoped to the verification domain: run traces, answers, and record-chain integrity. Each tool covers a distinct workflow and none feels redundant.
The toolkit covers the full observed verification lifecycle: generating verified run records, generating answer records, and validating logs of those records. Policies are provided as parameters rather than requiring separate management tools, so there are no obvious dead ends.
Maintenance
Related MCP Connectors
Fact-checks generated content against your sources of truth showing what to trust, change, & verify.
- PortMemOAuthcom.portmem
Check AI-written drafts against your source documents, with the passage behind every verdict.
Real-time fact-check, citation verification, and source-freshness for AI agents.
DraftCheck: flags AI-writing patterns in prose with fix hints. Deterministic linter, not a detector.
Related MCP Servers
AlicenseBqualityFmaintenanceAn MCP server that provides a comprehensive interface to Semgrep, enabling users to scan code for security vulnerabilities, create custom rules, and analyze scan results through the Model Context Protocol.6701 PyPI687MIT
Sentry MCPofficial
AlicenseAqualityAmaintenanceA remote Model Context Protocol server acting as middleware to the Sentry API, allowing AI assistants like Claude to access Sentry data and functionality through natural language interfaces.738 npm917MIT
Vectara MCP serverofficial
AlicenseAqualityDmaintenanceOpen source MCP server for Vectara21,497 PyPI29Apache 2.0
Arkheia Hallucinationofficial
AlicenseNot gradedqualityDmaintenanceDetect fabrication and hallucination in any LLM output. Score responses from GPT-4o, Claude, Gemini, Llama and 30+ models. Free tier included.2MIT