Skip to main content
Glama

Groundlens: RAG回答のための校正者

Groundlens

PyPI Python License Runtime dependencies groundlens MCP server

CI OpenSSF Best Practices OpenSSF Scorecard Determinism

groundlens.dev

仕組み · インストール · クイックスタート · MCPサーバー · 制限事項 · 再現性

Groundlensは、モデルが書いたもののための校正者です。ソースが裏付けない単語に印を付け、それぞれが何と言うべきだったかを示します。RAGの回答が、取得されたソースに対してグラウンディングされ忠実であるかをチェックします。これは、人々が幻覚検出、引用チェック、RAG評価に頼る仕事ですが、Groundlensは、しきい値を設定するための判定やスコアではなく、レビュアーのための印と証拠を返す点で異なります。

QUESTION    What is the invoice total?
SOURCE      ...the total amount due is 10,000 dollars, payable within 30 days...
ANSWER      The invoice total is 1,000 dollars, due in 30 days.

GROUNDLENS  1,000   nothing supports this.   Closest in invoice.pdf#p1: '10,000'

答えが間違っているとは決して言いません。どの単語を見るべきか、どのドキュメントを開くべきかを示します。5分の代わりに30秒の人間の注意で済みます。

仕組み

How Groundlens checks words and numbers

Groundlensは、単語と数値の比較を2つの異なる方法で行います。

単語

数値

単語は意味によって固定されます。単語のサポートは、凍結された既製のエンコーダー(検索ですでに使用しているのと同じ種類)を使用して、ソースの任意の単語に対して到達する最大のコサイン類似度です。

数値は算術によって固定されます。数字はフォーマットが正規化された値に解析され(10,000、10000、$10,000、10 000、および(宣言されたロケールでは)10.000は同じ数値)、ソース内のすべての値と照合されます。サポートは正確に1.0または正確に0.0です。類似度が投票することはありません。

Groundlensは平均ではなく、最低スコアを出力として提供します。トークン類似度の指標はすべて平均で集約されますが、平均は単一トークンのエラーが埋もれてしまう場所です。

実用的な例:10は100ではない

取得されたドキュメントには、請求額は10,000ドルと記載されています。回答は1,000ドルと述べています。人間なら、金融の学位がなくても即座に気づきます。

埋め込み類似度はそうではありません。正しい答えと間違った答えのコサイン類似度は約0.99で、エラーはインクの滴が池に溶けるようにベクトルに溶け込みます。LLM判定も同様です。もっともらしさを読むため、「合計は1,000ドル」という文は請求書について完全にもっともらしい文です。学習済みスパン検出器も同様です。1桁の置換はトレーニングラベルではまれだからです。

文エンコーダーは、語彙、トピック、構造によってテキストを整理します。真実によって整理することは決してありません。正しい文の中の間違った数字は、パラフレーズを同一視するエンコーダーにとっては、ほぼパラフレーズです。

その請求書では、間違った回答の平均サポートは0.79で、問題なさそうです。最も弱いアンカーは0.00で、余白の印です。

運用上のしきい値

このライブラリにはデフォルトのしきい値はありません。しきい値はメソッドのプロパティではなく、デプロイメントのプロパティです。エンコーダー、データ、そして偽陰性と比較した偽陽性のコストによって異なります。ここではそのどれも不明です。

このルールの背後には測定があります。実行した動作点グリッド全体で、95%の再現率での最良の偽陽性率は0.65で、テストしたすべてのシングルパス検出器(これを含む)で同じでした。規制されたレビューが実際に必要とする再現率では、そのグリッドのどの固定カットも使用できません。カットを出荷することは、すでに成立しないとわかっている数値を出荷することを意味します。

Support scores and the weakest anchor

groundlensが提供するものは次のとおりです。

  • 単語ごとのサポートスコア。低いほどソースによるサポートが少ないことを意味します。

  • 証拠付きのマーク:単語、そのスパン、サポート、および最も近いエビデンス文。レビューアは数秒で任意の呼び出しを確認できます。

  • 独自のラベル付きデータにカットを適合させる関数calibrate()。200未満のラベル付き例では実行を拒否します。それを下回るとカットがノイズになるためです。

パイプラインでしきい値が必要な場合は、ラベル付きデータでcalibrate()を実行します:

from groundlens import calibrate

point = calibrate(labelled, target_recall=0.95)
print(point.threshold, point.fpr, point.fpr_ci95)   # read the fpr first

calibrate()には少なくとも200のラベル付き例が必要です。それを下回ると、95%の再現率しきい値がわずかなポイントから推定されるためです。

Related MCP server: Sentry MCP

インストール

pip install groundlens              # zero runtime dependencies. Not numpy, not torch
pip install "groundlens[encoder]"   # + the reference sentence encoder
pip install "groundlens[encoder,mcp]"   # + the MCP server, for Claude Desktop and friends

コアインストールはパッケージを一切取り込みません。もし変更があればCIジョブがビルドを失敗させます。以前のバージョンは、何かをする前に約2ギガバイトのディープラーニングスタックをインストールしていました。

クイックスタート

from groundlens import proofread, SentenceTransformerEncoder

answer = "The invoice total is 4.75% payable within 45 days."
sources = [("policy.pdf#p3", "The rate stated in the policy is 3.90% and the term is 30 days.")]

marks = proofread(answer, sources, encoder=SentenceTransformerEncoder(), k=2)

print(marks.report())
#  4.75%   support 0.00    nearest in policy.pdf#p3: '3.90%'
#  45      support 0.00    nearest in policy.pdf#p3: '30'

すべてのマークには領収書が付いています:

for anchor in marks.weakest:
    anchor.text            # '4.75%'          the word in the answer
    anchor.span            # (21, 26)         where it sits
    anchor.kind            # 'numeral'        checked by arithmetic, not meaning
    anchor.support         # 0.0              absent from the sources
    anchor.evidence_id     # 'policy.pdf#p3'  which document to open
    anchor.evidence_text   # '3.90%'          what it should have matched

シェルから:

groundlens read --answer answer.txt --context policy.pdf#p3=policy.txt

MCPサーバー

同じ校正者がアシスタントの中にいます。GroundlensにはMCPサーバーが同梱されているため、Claude Desktop、Claude Code、Cursor、VS Code、その他のMCPクライアントは、会話を離れることなく回答をソースと照合できます。stdioを介してローカルで実行されます。テキストが外部に送信されることはありません。

pip install "groundlens[encoder,mcp]"
python -m groundlens.mcp

次に、クライアントでそれを指定します。claude_desktop_config.json、またはCursorとVS Codeの同等のmcp.jsonで:

{
  "mcpServers": {
    "groundlens": {
      "command": "python",
      "args": ["-m", "groundlens.mcp"]
    }
  }
}

PATHにあるPythonでない場合は、GroundlensがインストールされているPythonへの絶対パスを使用します:/path/to/venv/bin/python。

たった一つのツール

find_unsupported_words(answer, sources, k=4, locale="und")

answer

チェックするモデル出力

sources

[{"id": "policy.pdf#p3", "text": "..."}]。idは検出結果に含まれるため、読者はどのドキュメントを開くべきかわかります。

k

返す最も弱いアンカーの数

locale

これらのドキュメントでの数字の表記方法。esは1.234を1234として読み取り、enは1.234として読み取り、undは両方の読み取りを保持します。

最も弱いアンカーを、そのレシート、フロア、エンコーダーID、および検出結果のsha256とともに返します:

{
  "weakest_anchors": [
    {
      "word": "4.75%",
      "support": 0.0,
      "checked_by": "arithmetic",
      "closest_in_sources": "3.90%",
      "source_id": "policy.pdf#p3",
      "notes": []
    }
  ],
  "floor": 0.0,
  "n_marked": 12,
  "encoder_id": "all-mpnet-base-v2@<revision-sha>",
  "sha256": "..."
}

意図的に1つのツールです。以前のサーバーは3つを宣伝していましたが、それが、誰もインストールする前に1つの製品が3つのストーリーになる方法です。

このライブラリの他のすべての場所と同様に、ここには判定もしきい値もありません。数値のsupportが0.00の場合、その値がソースに存在しないことを意味します。単語の場合、語彙アンカーが見つからなかったことを意味し、これは忠実な言い換えでは普通のことです。サーバーはマークを報告し、読者が判断します。

エンコーダーは起動時ではなく最初の呼び出しでロードされ、モデルは初回使用時に一度だけダウンロードされます(約420 MB)。

制限事項

  • 計算された値を検証することはできません。「収益が5Mから15Mに増加した」というソースに対する「収益が3倍になった」など。

  • 単語チャネルは、単語がソースによってサポートされているかどうかをチェックします。正しいものに付随しているかどうかはチェックしません。回答が請求書Aについて「30日で支払可能」と述べており、同じコンテキストの別の場所で30日が請求書Bに属している場合、単語はサポートされており、マークは表示されません。

  • 推論をチェックすることはできません。それは含意モデルの役割です。

  • 取得を継承します。パッセージが間違っていれば、回答のグラウンディングも間違っています。

  • セグメンテーションはスペース区切りのスクリプトを想定しており、テキストが主にCJKまたはタイ語の場合は、偽装するのではなく警告を発します。

再現性

  • 数値チャネルは正確です。 10進比較、固定算術コンテキスト、ロケールは引数から取得され、LC_ALLからは決して取得されません。どのマシンでもバイト単位で同一です。CIはPYTHONHASHSEED=randomとトルコロケールの下で10のOS×Pythonの組み合わせでそれを証明しています。

  • 語彙チャネルは、固定されたエンコーダーのリビジョンからのfloat32コサインです。 モデル名ではありません。静かな再アップロードによって公開したすべての数値が変わるためです。プラットフォーム間で1e-6まで再現され、最も弱いアンカーの順序は安定しています。x86とApple Siliconの間でビット単位で同一ではなく、そのような主張はありません。

  • marks.sha256は構造と数値サポートを正確にカバーし、語彙サポートを小数第6位に丸めます。ハッシュを再現すると、算術の最後のビットではなく、検出結果が再現されます。

groundlens.dev · PyPI · Retractions · Contributing · Apache-2.0

Available Tools

3 tools
verify_answerB

Verify an answer against its sources under a policy and return the sealed record.

    sources: (id, text) pairs, {"id","text"} dicts, or bare strings.
    policy: a built-in name (e.g. "eu_ai_act_high_risk_v1"), a path, or YAML.
    Returns the decision (PASS/REVIEW/FAIL), the evidence, the regulatory
    mapping and the record with its content hash.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
answerYes
localeNound
policyNo
sourcesYes
questionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavior. It does describe the return value (decision, evidence, regulatory mapping, record with content hash), which is helpful. However, it does not state whether the operation is read-only, whether it stores or modifies any data, or what side effects might occur. For a verification tool, this is a notable gap, especially since the action of returning a 'sealed record' implies some immutability but not explicitly a non-destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded, stating the core action in the first sentence. It then efficiently lists input format variants and the return contents. The multi-line formatting with indentation is slightly unconventional but does not harm readability. There is minimal redundancy, and every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters (2 required) and an output schema exists, the description is moderately complete. It covers the key inputs (sources, policy) and mentions the return structure. However, it omits explanation of 'locale' and 'question', and does not provide usage context relative to sibling tools or error scenarios. The presence of an output schema lightens the need to detail return fields, but the missing parameter semantics and lack of sibling differentiation reduce completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the semantics of 'sources' (formats) and 'policy' (built-in, path, YAML). The 'answer' parameter is implicitly clear from the first sentence. However, 'locale' and 'question' are not described at all. Thus, the description covers only a portion of the parameters, leaving two parameters with no guidance beyond their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description begins with a clear, specific verb and resource: 'Verify an answer against its sources under a policy and return the sealed record.' This distinguishes it from siblings (verify_run, verify_records) by focusing on answer verification, which is a distinct operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides usage details such as acceptable formats for sources (id/text pairs, dicts, strings) and policy (built-in name, path, YAML), which implicitly guides the caller. However, it does not explicitly state when to use this tool versus the sibling tools verify_run or verify_records, nor does it mention any exclusions or alternative conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_recordsA

Verify a log of records offline: every hash, every link, every signature.

    records: the JSON Lines text of an answer-record or run-record log.
    Returns {"ok", "verified", "kind"}; fails if any record or link was altered.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
recordsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does meaningful work: it discloses the return shape ('Returns {"ok", "verified", "kind"}'), the failure mode ('fails if any record or link was altered'), and that the operation happens offline. It stops short of explicitly stating verification is non-destructive, a minor gap given 'verify' implies it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact — purpose is front-loaded in the first sentence, followed by the parameter and then the return/failure behavior. Every clause carries information an agent needs; there is no filler or restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter verification tool with an output schema present, the description covers purpose, input format, return shape, and failure behavior — nearly everything needed to call it correctly. Minor gaps like the possible values of 'kind' are left to the output schema, which is acceptable per the rubric.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it documents 'records' as 'the JSON Lines text of an answer-record or run-record log,' adding format and content meaning the schema lacks. It doesn't specify the exact structure of a valid record, but for a single string parameter the added semantics are substantial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Verify a log of records offline') with concrete scope ('every hash, every link, every signature'), so an agent can tell exactly what operation this performs. It also distinguishes this from the siblings verify_run and verify_answer by clarifying that it accepts both answer-record and run-record logs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by noting the tool accepts 'an answer-record or run-record log,' which hints it covers the domains of both siblings. However, it never names verify_run or verify_answer or gives an explicit when-to-use vs. when-not-to-use rule, leaving the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_runA

Verify an MCP execution trace under an execution policy and return the run record.

    trace: the MCP session as JSON-RPC messages (JSON Lines).
    policy: the execution policy, as YAML/JSON text or a path.
    Returns the gate (ALLOW/REVIEW/DENY), any breaches, and the signed run record.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
traceYes
policyYes
run_idYes
systemYes
started_atNo
system_versionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the return values (gate, breaches, signed run record) but does not mention potential side effects (e.g., whether it writes or stores anything), permission requirements, or error behavior. This is some behavioral context but incomplete for a tool with no annotation safety net.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is reasonably concise, with the purpose front-loaded and parameters broken into clear lines. It avoids redundant wording and communicates the key return values efficiently, though it could be tightened slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return format details are not strictly required, and the description already provides a high-level return summary. However, given the six-parameter complexity and lack of annotations, the description should explain all parameters and ideally differentiate usage from siblings. It covers the core purpose but leaves several parameters and usage guidance gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains trace (format: JSON-RPC messages as JSON Lines) and policy (format: YAML/JSON text or path), which is useful. However, it does not explain run_id, system, started_at, or system_version, leaving 4 of 6 parameters undocumented in both schema and description. This is a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool verifies an MCP execution trace against an execution policy and returns the run record with gate, breaches, and signed record. This specific verb+resource distinguishes it from sibling tools verify_answer and verify_records, which target different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by specifying it is for verifying execution traces, which gives clear context. However, it does not explicitly mention when not to use it or point to alternatives like verify_answer or verify_records, so it lacks explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv3.0.6
    • Removedfind_unsupported_words
    • Addedverify_answer
    • Addedverify_records
    • Addedverify_run
  2. 1 tool updatev0.1.0
    • First observedfind_unsupported_words

TDQS

A4/5.0

Scored across 3 tools

Disambiguation5/5

The three tools address clearly different verification targets: execution traces, answer-source pairs, and record logs. No two tools accept the same kind of input or produce the same kind of output, so an agent can select among them without ambiguity.

Naming Consistency5/5

All tool names follow the same verify_<noun> pattern with snake_case, matching the verb-object convention. The naming makes the input type immediately predictable from the tool name.

Tool Count5/5

At three tools, the surface is tightly scoped to the verification domain: run traces, answers, and record-chain integrity. Each tool covers a distinct workflow and none feels redundant.

Completeness5/5

The toolkit covers the full observed verification lifecycle: generating verified run records, generating answer records, and validating logs of those records. Policies are provided as parameters rather than requiring separate management tools, so there are no obvious dead ends.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers