groundlens
OfficialGroundlens: RAG 답변을 위한 교정기

작동 방식 · 설치 · 빠른 시작 · MCP 서버 · 한계 · 재현성
Groundlens는 모델이 작성한 내용을 위한 교정기입니다. 소스가 뒷받침하지 않는 단어에 표시를 남기고, 각 단어가 무엇이어야 했는지 보여줍니다. RAG 답변의 grounding과 충실성을 검색된 소스에 대해 검사합니다. 이는 사람들이 환각 탐지, 인용 확인, RAG 평가를 위해 사용하는 작업이지만, Groundlens는 점수나 임계값을 제시하는 판정 대신 검토자를 위한 표시와 증거를 반환한다는 점에서 다릅니다.
QUESTION What is the invoice total?
SOURCE ...the total amount due is 10,000 dollars, payable within 30 days...
ANSWER The invoice total is 1,000 dollars, due in 30 days.
GROUNDLENS 1,000 nothing supports this. Closest in invoice.pdf#p1: '10,000'답변이 틀렸다고 말하지 않습니다. 어떤 단어를 봐야 하는지, 어떤 문서를 열어야 하는지 알려줍니다. 5분 대신 30초의 인간의 주의만으로 충분합니다.
작동 방식

Groundlens는 단어와 숫자 비교를 두 가지 다른 방식으로 접근합니다:
단어 | 숫자 |
단어는 의미에 의해 고정됩니다. 단어의 지지도(support)는 소스의 모든 단어 중 가장 높은 코사인 유사도로, 동결된 기성 인코더를 사용합니다 — 검색에서 이미 사용하는 것과 같은 종류입니다. | 숫자는 산술에 의해 고정됩니다. 숫자는 형식이 정규화된 값으로 파싱됩니다 — |
Groundlens는 평균이 아닌 최저 점수를 출력으로 제공합니다. 모든 토큰 유사도 메트릭은 평균으로 집계되며, 평균은 단일 토큰 오류가 사라지는 곳입니다.
실용적인 예: 10은 100이 아니다
검색된 문서에 총 납부액이 10,000달러라고 나와 있습니다. 답변은 1,000달러라고 말합니다. 인간은 금융 학위 없이도 즉시 알아차립니다.
임베딩 유사도는 그렇지 않습니다. 올바른 답과 틀린 답 사이의 코사인은 약 0.99입니다 — 오류는 잉크 한 방울이 웅덩이에 녹아들듯 벡터 속으로 용해됩니다. LLM 판정자도 마찬가지입니다: 그럴듯함을 읽을 뿐이며, "총액은 1,000달러입니다"는 인보이스에 관한 완벽하게 그럴듯한 문장입니다. 훈련된 스팬 탐지기도 마찬가지입니다. 단일 숫자 치환은 훈련 레이블에서 드물기 때문입니다.
문장 인코더는 어휘, 주제, 구조에 따라 텍스트를 구성합니다. 결코 진실에 의해서가 아닙니다. 올바른 문장 안의 잘못된 숫자는, 의역을 압축하는 인코더에게는 거의 의역에 가깝습니다.
그 인보이스에서 틀린 답변의 평균 지지도는 0.79입니다 — 괜찮아 보입니다. 가장 약한 앵커는 0.00입니다 — 여백에 표시를 남길 만한 값입니다.
운영 임계값
이 라이브러리에는 기본 임계값이 없습니다. 임계값은 방법의 속성이 아니라 배포의 속성입니다. 인코더, 데이터, 그리고 거짓 양성이 거짓 음성보다 얼마나 비싼지에 따라 달라집니다. 여기서는 그 어떤 것도 알 수 없습니다.
규칙 뒤에는 측정이 있습니다. 우리가 실행한 운영 지점 그리드 전반에서, 95% 재현율에서의 최상의 거짓 양성률은 0.65였으며, 우리가 테스트한 모든 단일 패스 탐지기(이것 포함)에 해당했습니다. 규제된 검토가 실제로 필요로 하는 재현율에서는, 그 그리드의 어떤 고정 컷도 사용할 수 없습니다. 하나를 제공한다는 것은 이미 성립하지 않는다는 것을 알고 있는 숫자를 제공한다는 뜻입니다.

groundlens가 제공하는 것은 다음과 같습니다:
단어별 지지도 점수. 낮을수록 소스에 덜 뒷받침된다는 뜻입니다.
영수증이 있는 표시: 단어, 스팬, 지지도, 그리고 가장 가까운 증거 문장. 검토자는 몇 초 안에 어떤 판단이든 확인할 수 있습니다.
calibrate()함수. 자체 레이블링된 데이터에 컷을 맞춥니다. 200개 미만의 레이블링된 예제에서는 실행을 거부합니다. 그 이하에서는 컷이 노이즈이기 때문입니다.
파이프라인에 임계값이 필요하다면, 레이블링된 데이터에서 calibrate()를 실행하세요:
from groundlens import calibrate
point = calibrate(labelled, target_recall=0.95)
print(point.threshold, point.fpr, point.fpr_ci95) # read the fpr first
calibrate()는 최소 200개의 레이블링된 예제가 필요합니다. 그 이하에서는 95% 재현율 임계값이 소수의 지점에서 추정되기 때문입니다.
Related MCP server: Sentry MCP
설치
pip install groundlens # zero runtime dependencies. Not numpy, not torch
pip install "groundlens[encoder]" # + the reference sentence encoder
pip install "groundlens[encoder,mcp]" # + the MCP server, for Claude Desktop and friends핵심 설치는 어떤 패키지도 가져오지 않으며, CI 작업은 그런 일이 발생하면 빌드를 실패시킵니다. 이전 버전은 아무것도 하기 전에 약 2기가바이트의 딥러닝 스택을 설치했습니다.
빠른 시작
from groundlens import proofread, SentenceTransformerEncoder
answer = "The invoice total is 4.75% payable within 45 days."
sources = [("policy.pdf#p3", "The rate stated in the policy is 3.90% and the term is 30 days.")]
marks = proofread(answer, sources, encoder=SentenceTransformerEncoder(), k=2)
print(marks.report())
# 4.75% support 0.00 nearest in policy.pdf#p3: '3.90%'
# 45 support 0.00 nearest in policy.pdf#p3: '30'모든 표시는 영수증을 동반합니다:
for anchor in marks.weakest:
anchor.text # '4.75%' the word in the answer
anchor.span # (21, 26) where it sits
anchor.kind # 'numeral' checked by arithmetic, not meaning
anchor.support # 0.0 absent from the sources
anchor.evidence_id # 'policy.pdf#p3' which document to open
anchor.evidence_text # '3.90%' what it should have matched셸에서:
groundlens read --answer answer.txt --context policy.pdf#p3=policy.txtMCP 서버
동일한 교정기를 어시스턴트 안에서 사용할 수 있습니다. Groundlens는 MCP 서버를 제공하므로, Claude Desktop, Claude Code, Cursor, VS Code 또는 다른 MCP 클라이언트가 대화를 떠나지 않고도 답변을 소스에 대해 확인할 수 있습니다. stdio를 통해 로컬에서 실행됩니다. 어떤 텍스트도 외부로 나가지 않습니다.
pip install "groundlens[encoder,mcp]"
python -m groundlens.mcp그런 다음 클라이언트를 서버에 연결하세요. claude_desktop_config.json — 또는 Cursor와 VS Code의 해당 mcp.json에서:
{
"mcpServers": {
"groundlens": {
"command": "python",
"args": ["-m", "groundlens.mcp"]
}
}
}Groundlens가 설치된 Python이 PATH에 있는 것이 아니라면 절대 경로를 사용하세요: /path/to/venv/bin/python.
단 하나의 도구
find_unsupported_words(answer, sources, k=4, locale="und")
| 확인할 모델 출력 |
|
|
| 반환할 가장 약한 앵커의 수 |
| 이 문서들이 숫자를 쓰는 방식. |
가장 약한 앵커를 영수증, 바닥값, 인코더 id, 그리고 결과의 sha256과 함께 반환합니다:
{
"weakest_anchors": [
{
"word": "4.75%",
"support": 0.0,
"checked_by": "arithmetic",
"closest_in_sources": "3.90%",
"source_id": "policy.pdf#p3",
"notes": []
}
],
"floor": 0.0,
"n_marked": 12,
"encoder_id": "all-mpnet-base-v2@<revision-sha>",
"sha256": "..."
}의도적으로 도구 하나뿐입니다. 이전 서버는 세 개를 광고했고, 그것이 하나의 제품이 누구도 설치하기 전에 세 개의 이야기가 되는 방법입니다.
이 라이브러리의 다른 모든 곳과 마찬가지로 판정도 임계값도 없습니다. 숫자에 대한 support 0.00은 그 값이 소스에 없다는 뜻입니다. 단어에 대한 0.00은 어휘 앵커를 찾지 못했다는 뜻이며, 이는 충실한 의역에서 흔한 일입니다. 서버는 표시를 보고할 뿐, 판단은 독자가 합니다.
인코더는 시작 시가 아니라 첫 호출 시 로드되며, 모델은 처음 사용할 때 한 번 다운로드됩니다(약 420MB).
한계
계산된 값을 검증할 수 없습니다 — "수익이 3배가 되었다"는 "수익이 5M에서 15M으로 증가했다"는 소스에 대해 검증할 수 없습니다.
단어 채널은 단어가 소스에 의해 뒷받침되는지 확인합니다. 올바른 대상에 연결되어 있는지는 확인하지 않습니다. 답변이 인보이스 A에 대해 "30일 내 지불"이라고 말하고, 그 30일이 같은 맥락의 다른 곳에서 인보이스 B에 속한다면, 그 단어는 뒷받침되며 표시가 나타나지 않습니다.
추론을 확인할 수 없습니다. 그것은 함의 모델의 몫입니다.
검색의 한계를 그대로 물려받습니다. 구절이 틀리면 답변의 grounding도 틀립니다.
분할은 공백 구분 스크립트를 가정하며, 텍스트가 대부분 CJK 또는 태국어일 때는 속이는 대신 경고합니다.
재현성
숫자 채널은 정확합니다. 소수 비교, 고정 산술 컨텍스트,
LC_ALL이 아닌 인자에서 오는 로케일. 어떤 머신에서든 바이트 단위로 동일합니다 — CI가PYTHONHASHSEED=random과 터키어 로케일 하에서 10개의 OS × Python 조합으로 증명합니다.**어휘 채널은 고정된 인코더 리비전의 float32 코사인입니다.** 모델 이름이 아닙니다. 조용한 재업로드가 게시한 모든 숫자를 바꿀 수 있기 때문입니다. 플랫폼 간 1e-6까지 재현되며 가장 약한 앵커의 순서는 안정적입니다. x86과 Apple Silicon 사이에서 비트 단위로 동일하지는 않으며, 우리는 그렇게 주장하지 않습니다.
marks.sha256는 구조와 숫자 지지도를 정확히 포함하며, 어휘 지지도는 소수점 6자리로 반올림합니다. 해시를 재현하면 산술의 마지막 비트가 아닌 결과를 재현하는 것입니다.
groundlens.dev · PyPI · 철회 · 기여 · Apache-2.0
Available Tools
3 toolsverify_answerB
Verify an answer against its sources under a policy and return the sealed record.
sources: (id, text) pairs, {"id","text"} dicts, or bare strings.
policy: a built-in name (e.g. "eu_ai_act_high_risk_v1"), a path, or YAML.
Returns the decision (PASS/REVIEW/FAIL), the evidence, the regulatory
mapping and the record with its content hash.
| Name | Required | Description | Default |
|---|---|---|---|
| answer | Yes | ||
| locale | No | und | |
| policy | No | ||
| sources | Yes | ||
| question | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavior. It does describe the return value (decision, evidence, regulatory mapping, record with content hash), which is helpful. However, it does not state whether the operation is read-only, whether it stores or modifies any data, or what side effects might occur. For a verification tool, this is a notable gap, especially since the action of returning a 'sealed record' implies some immutability but not explicitly a non-destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded, stating the core action in the first sentence. It then efficiently lists input format variants and the return contents. The multi-line formatting with indentation is slightly unconventional but does not harm readability. There is minimal redundancy, and every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 5 parameters (2 required) and an output schema exists, the description is moderately complete. It covers the key inputs (sources, policy) and mentions the return structure. However, it omits explanation of 'locale' and 'question', and does not provide usage context relative to sibling tools or error scenarios. The presence of an output schema lightens the need to detail return fields, but the missing parameter semantics and lack of sibling differentiation reduce completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the semantics of 'sources' (formats) and 'policy' (built-in, path, YAML). The 'answer' parameter is implicitly clear from the first sentence. However, 'locale' and 'question' are not described at all. Thus, the description covers only a portion of the parameters, leaving two parameters with no guidance beyond their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a clear, specific verb and resource: 'Verify an answer against its sources under a policy and return the sealed record.' This distinguishes it from siblings (verify_run, verify_records) by focusing on answer verification, which is a distinct operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides usage details such as acceptable formats for sources (id/text pairs, dicts, strings) and policy (built-in name, path, YAML), which implicitly guides the caller. However, it does not explicitly state when to use this tool versus the sibling tools verify_run or verify_records, nor does it mention any exclusions or alternative conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_recordsA
Verify a log of records offline: every hash, every link, every signature.
records: the JSON Lines text of an answer-record or run-record log.
Returns {"ok", "verified", "kind"}; fails if any record or link was altered.
| Name | Required | Description | Default |
|---|---|---|---|
| records | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does meaningful work: it discloses the return shape ('Returns {"ok", "verified", "kind"}'), the failure mode ('fails if any record or link was altered'), and that the operation happens offline. It stops short of explicitly stating verification is non-destructive, a minor gap given 'verify' implies it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact — purpose is front-loaded in the first sentence, followed by the parameter and then the return/failure behavior. Every clause carries information an agent needs; there is no filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter verification tool with an output schema present, the description covers purpose, input format, return shape, and failure behavior — nearly everything needed to call it correctly. Minor gaps like the possible values of 'kind' are left to the output schema, which is acceptable per the rubric.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: it documents 'records' as 'the JSON Lines text of an answer-record or run-record log,' adding format and content meaning the schema lacks. It doesn't specify the exact structure of a valid record, but for a single string parameter the added semantics are substantial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Verify a log of records offline') with concrete scope ('every hash, every link, every signature'), so an agent can tell exactly what operation this performs. It also distinguishes this from the siblings verify_run and verify_answer by clarifying that it accepts both answer-record and run-record logs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by noting the tool accepts 'an answer-record or run-record log,' which hints it covers the domains of both siblings. However, it never names verify_run or verify_answer or gives an explicit when-to-use vs. when-not-to-use rule, leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_runA
Verify an MCP execution trace under an execution policy and return the run record.
trace: the MCP session as JSON-RPC messages (JSON Lines).
policy: the execution policy, as YAML/JSON text or a path.
Returns the gate (ALLOW/REVIEW/DENY), any breaches, and the signed run record.
| Name | Required | Description | Default |
|---|---|---|---|
| trace | Yes | ||
| policy | Yes | ||
| run_id | Yes | ||
| system | Yes | ||
| started_at | No | ||
| system_version | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return values (gate, breaches, signed run record) but does not mention potential side effects (e.g., whether it writes or stores anything), permission requirements, or error behavior. This is some behavioral context but incomplete for a tool with no annotation safety net.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is reasonably concise, with the purpose front-loaded and parameters broken into clear lines. It avoids redundant wording and communicates the key return values efficiently, though it could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return format details are not strictly required, and the description already provides a high-level return summary. However, given the six-parameter complexity and lack of annotations, the description should explain all parameters and ideally differentiate usage from siblings. It covers the core purpose but leaves several parameters and usage guidance gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains trace (format: JSON-RPC messages as JSON Lines) and policy (format: YAML/JSON text or path), which is useful. However, it does not explain run_id, system, started_at, or system_version, leaving 4 of 6 parameters undocumented in both schema and description. This is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool verifies an MCP execution trace against an execution policy and returns the run record with gate, breaches, and signed record. This specific verb+resource distinguishes it from sibling tools verify_answer and verify_records, which target different resources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by specifying it is for verifying execution traces, which gives clear context. However, it does not explicitly mention when not to use it or point to alternatives like verify_answer or verify_records, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v3.0.6- Removed
find_unsupported_words - Added
verify_answer - Added
verify_records - Added
verify_run
1 tool update
v0.1.0- First observed
find_unsupported_words
TDQS
Scored across 3 tools
The three tools address clearly different verification targets: execution traces, answer-source pairs, and record logs. No two tools accept the same kind of input or produce the same kind of output, so an agent can select among them without ambiguity.
All tool names follow the same verify_<noun> pattern with snake_case, matching the verb-object convention. The naming makes the input type immediately predictable from the tool name.
At three tools, the surface is tightly scoped to the verification domain: run traces, answers, and record-chain integrity. Each tool covers a distinct workflow and none feels redundant.
The toolkit covers the full observed verification lifecycle: generating verified run records, generating answer records, and validating logs of those records. Policies are provided as parameters rather than requiring separate management tools, so there are no obvious dead ends.
Maintenance
Related MCP Connectors
Fact-checks generated content against your sources of truth showing what to trust, change, & verify.
- PortMemOAuthcom.portmem
Check AI-written drafts against your source documents, with the passage behind every verdict.
Real-time fact-check, citation verification, and source-freshness for AI agents.
DraftCheck: flags AI-writing patterns in prose with fix hints. Deterministic linter, not a detector.
Related MCP Servers
AlicenseBqualityFmaintenanceAn MCP server that provides a comprehensive interface to Semgrep, enabling users to scan code for security vulnerabilities, create custom rules, and analyze scan results through the Model Context Protocol.6701 PyPI687MIT
Sentry MCPofficial
AlicenseAqualityAmaintenanceA remote Model Context Protocol server acting as middleware to the Sentry API, allowing AI assistants like Claude to access Sentry data and functionality through natural language interfaces.738 npm917MIT
Vectara MCP serverofficial
AlicenseAqualityDmaintenanceOpen source MCP server for Vectara21,497 PyPI29Apache 2.0
Arkheia Hallucinationofficial
AlicenseNot gradedqualityDmaintenanceDetect fabrication and hallucination in any LLM output. Score responses from GPT-4o, Claude, Gemini, Llama and 30+ models. Free tier included.2MIT