Skip to main content
Glama

World Model MCP

面向 AI 编码代理的持久记忆 + 可选签名审计。

world-model-mcp 为 AI 编码代理提供持久记忆,外加一条后量子签名、可离线验证的审计轨迹(FIPS 205 混合 Ed25519 + SLH-DSA),MIT 许可,完全本地运行。

world-model-mcp 附带一个本地 SQLite 知识图谱,你的代理会在每一轮中查询它:幻觉变得可验证,修正会跨会话保留,回归在落地之前就会被捕获。开启审计链后,每个事件都会使用 FIPS 205 混合 Ed25519 + SLH-DSA 签名,可永久离线验证。MIT 许可,完全本地运行,支持 10+ 款 AI 编码代理,包括 Claude Code、Cursor、Codex、Continue、Cline、Windsurf、GitHub Copilot Chat、pi、OpenClaw 和 Hermes Agent。

最新版本:v0.16.2。 world-model demo 以一段实时公证演示收尾:三个示例决策被签名进一个真实的混合签名纪元(Ed25519 + SLH-DSA-SHA2-128f,FIPS 205),验证为 VALID,篡改一个字节以证明篡改检测实时触发(INVALID),然后恢复原状。零注册、零令牌、零网络;会在当前目录生成一个 world-model-demo-receipt.json,并输出一个可复制粘贴的 etch.systems/verify#<manifest> URL。v0.16.1 增加了在本地 liboqs 构建未编译 SLH-DSA 时优雅跳过(而不是让演示崩溃)的能力。完整版本历史见 CHANGELOG.md。

10 秒试用

pip install -U world-model-mcp && world-model demo

你会看到三个示例决策被签名进一个防篡改纪元,然后演示篡改一个字节并重新验证,以证明篡改会被立即检测到(INVALID),随后恢复并重新验证(再次 VALID)。离线、无账户、无网络。一个 world-model-demo-receipt.json 会落在你的当前目录中,并会打印一个可分享的 etch.systems/verify#… URL,供任何人在浏览器中检查同一份收据。

PyPI Downloads License: MIT Python 3.11+ world-model-mcp MCP server DOI

mcp-name: io.github.SaravananJaichandar/world-model-mcp


Related MCP server: Scrooge

更喜欢托管版本?

如果你想要同样的签名审计链,但又不想自己运行 OSS 服务器,etch.systems 提供了一个托管版本,使用本包作为其离线参考验证器。

  • 无需运行任何基础设施(30 秒内 POST 你的第一个事件,无需注册)

  • 合规框架映射到 SR 11-7、EU AI Act Article 12、ISO 42001、NIST AI RMF、SOC 2 CC7

  • 端到端提供外部锚定(Sigstore Rekor + Bitcoin OpenTimestamps)

  • 可移植的按代理身份,支持跨 MCP 客户端工具(Claude Code、Cursor、Continue、Cline、Codex)的密钥轮换和委派

  • 离线验证器保持不变:pip install world-model-mcp && etch-verify manifest.json 可针对固定的公钥验证托管链,而无需依赖托管服务在线

curl -X POST https://etch.systems/v1/your-project

返回一个 bearer 令牌和一个 MCP 端点,你可以将任何兼容 MCP 的客户端指向该端点。完整文档见 etch.systems。


常见问题

什么是 world-model-mcp? world-model-mcp 是面向 AI 编码代理的持久记忆,外加一条后量子签名、可离线验证的审计轨迹,以 MCP 服务器形式提供,代理在每一轮中都会查询它。它完全本地运行,以 MIT 许可的 Python 形式发布,并为 10+ 款 AI 编码代理提供可选适配器。

它与 Mem0、Letta 和其他代理记忆工具有什么不同? world-model-mcp 是唯一一款在记忆之外还附带混合后量子签名审计链(FIPS 205 Ed25519 + SLH-DSA-SHA2-128f)的代理记忆工具,并配有离线参考验证器(etch-verify),还可通过托管的 etch.systems 配套服务获得双重外部锚定(Sigstore Rekor + Bitcoin OpenTimestamps)。同类工具(Mem0、Letta、agentmemory)只提供记忆,没有签名审计链。逐机制对比见下方“对比”表。

审计轨迹是后量子安全的吗? 是的。每个已关闭的纪元都携带一个混合签名:Ed25519(经典)加上 SLH-DSA-SHA2-128f(FIPS 205 无状态基于哈希的后量子签名)。未来能够单独破解 Ed25519 的量子对手仍然要面对同一负载上的 SLH-DSA 签名;只有两个签名都被伪造,整条链才能被伪造。

如何在没有账户的情况下离线验证一条记录? 运行 pip install -U world-model-mcp && world-model demo。演示会将三个示例决策签名进一个真实的混合签名纪元,导出与 etch-verify CLI 读取的相同清单格式,验证为 VALID,篡改一个字节以证明篡改检测是实时的(INVALID),然后恢复并重新验证。一个 world-model-demo-receipt.json 会落在你的当前目录中,并会打印一个 etch.systems/verify#… URL,供任何人在浏览器中检查收据。零注册、零网络。

它支持哪些 AI 编码代理? Claude Code、Cursor、Codex、Continue、Cline、Windsurf、GitHub Copilot Chat、pi、OpenClaw 和 Hermes Agent。每个都有专门的入门仓库,位于 github.com/SaravananJaichandar/world-model-mcp-<agent>-starter 下。MCP 线格式是标准的,因此任何支持 MCP 的客户端都可以使用相同的工具。

我应该使用这个 OSS 包还是托管的 etch.systems 服务? 如果你想在自己的基础设施上运行审计链,或为内部工具添加签名记忆,请使用这个 OSS 包(MIT 许可,完全本地)。如果你想要以托管服务形式提供的同一条审计链,并附带合规框架映射(SR 11-7、EU AI Act Article 12、ISO 42001、NIST AI RMF、SOC 2 CC7)以及跨 MCP 客户端工具的可移植按代理身份,请使用 etch.systems。两种情况下离线参考验证器是相同的:pip install world-model-mcp && etch-verify manifest.json 可针对固定的公钥验证 OSS 生成或托管生成的清单。


与其他代理记忆 + 代理审计项目的对比

在审计 / 签名 / 锚定维度上与八个具名同类项目进行正面定位对比,这些维度正是代理记忆领域正在趋同的方向。world-model-mcp 的每个单元格都带有来源标记:own = 在发布产品中测量 / 观察到;cited = 取自竞争对手自己的公开落地页、仓库或新闻稿。

本表的权威来源位于托管服务仓库中:world-model-mcp-hosted/src/etch/competitor_matrix.py。如果你更改其中一处,请在两处都更新。

world-model-mcp(本仓库)+ Etch(托管)

Mem0

Letta

agentmemory

Unicity AOS

Repowise

Trinitite

Caura

FailproofAI

签名方案

Ed25519 + SLH-DSA-SHA2-128f(混合) [own]

无(仅静态加密)

无

无

BLAKE3 哈希链

确定性信号(无签名)

签名 + 哈希链(算法未公开)

无

未公开

后量子就绪

是(SLH-DSA,NIST FIPS 205) [own]

否

否

否

否(仅 BLAKE3 哈希)

否

未公开

否

否

离线参考验证器

etch-verify CLI,流式,随 PyPI 包发布 [own]

否

否

否

未公开

否(仅 SaaS)

基于浏览器的验证器

否

否

证明节奏强制

计划 systemd 定时器 + 按需验证 [own]

否

否

否

未公开

否

计划性证明运行

否

否

框架映射(条款 / 控制项 ID)

EU AI Act 第 12-15 条、SOC 2 CC6.6/7.2/7.3、ISO 27001 A.12.4/A.14.2 [own]

SOC 2 + HIPAA(徽章,无逐控制项映射)

未公开

否

未公开

EU AI Act(非适用性声明)

逐条引用(EU AI Act、SOC 2、SR 11-7)

SOC 2 进行中徽章

仅企业版提供 SOC 2

漂移检测

逐周期链完整性 + 证明趋势 [own]

否

成功率 + 错误跟踪

否

未公开

代码健康度分数变化

确定性重放 + 风险分数变化

否

评估器分数变化

注释支持(链内签名人工备注)

pin_annotation MCP 工具,签名到同一链中 [own]

否

否

否

未公开

否

未公开

否

否

浏览器验证器

是,/auditor/ 上的链完整性组件 [own]

否

否

否

未公开

否

是

否

否

共享链接(审计员访问,无需登录)

是,过期共享令牌(bs_ 前缀,最长 30 天) [own]

否

否

否

未公开

否

是,过期审计员访问

否

否

外部锚点(独立见证日志)

双重:Sigstore Rekor + Bitcoin OpenTimestamps(公开,可按项目选择退出) [own]

否

否

否

仅内部哈希链(无外部日志)

否

未公开

否

否

OSS 许可证

MIT(world-model-mcp) [own]

Apache 2.0

Apache 2.0

Apache 2.0

未公开

AGPL v3(核心)

无开源

Apache 2.0

无开源

GitHub 星标

通过 etch.systems/api/oss-stats 快照 [cited]

61.6k [cited]

23.9k [cited]

24.9k [cited]

7.1k [cited]

4.2k [cited]

无公开仓库

373 [cited]

无公开仓库

公开融资

自筹资金 [own]

$24.5M [cited]

$10M [cited]

未公开

$3M 种子轮 2026 年 2 月 [cited]

未公开

未公开

未公开

未公开

合规姿态(落地页所声称的)

SOC 2 Type I 进行中(目标 2026 年 8 月) [own]

SOC 2 + HIPAA(信任子域上的徽章)

未公开

无合规姿态

无合规姿态

无 SOC 2;EU AI Act 非适用性声明

逐条合规框架

SOC 2 进行中徽章

SOC 2 企业版

如何阅读此表:

  • 粗体单元格描述本仓库(OSS)或托管服务(etch.systems)中已交付的机制。

  • 同行单元格引自各竞品的公开落地页 / 仓库 / 新闻稿,非我们自行测量。

  • [own] 出现在 world-model-mcp 单元格中表示我们自行测量/观察了该机制;[cited] 表示该值来自第三方来源。

  • 此表仅涵盖审计 / 签名 / 锚定维度。通用记忆功能(检索准确率、适配器广度、LLM 支持)见下文 功能 部分。


数据

基准

分数

详情

SWE-bench Verified repeat-mistake

+10.2 分作为单次试验上限(49 个配对实例上 67.3% → 77.6%);多随机种子平均效应为 每实例 +0.24,95% CI [0, 0.47]

预注册,Claude Code 2.1.177 无头模式,Zenodo DOI 10.5281/zenodo.21076824。单次试验划分中,域内 +15.0 分,跨域 +6.9 分,零回归。

Contradiction-resolution

100.0%(auto 策略)

105 对 × 19 个类别,确定性(无 LLM)。自 v0.11.0 起提供。

Coach-Player verification

100.0% 精确匹配

12 个手工标注对(4 个有依据、4 个部分、4 个幻觉)。通过独立 Coach LLM 进行第 3 层对抗性验证。自 v0.12.12 起提供。

SWE-bench 数字是支撑性的实证声明。另外两个是已交付组件的内部正确性基准。可复现脚本位于各基准目录或链接的仓库中。


测试

1,494 个单元 + 集成 + 模糊测试覆盖已交付代码库。覆盖率下限门控在 71% / 66%(两档,取决于 CI 运行器的 liboqs 中是否包含 SLH-DSA)。

# Run tests
pytest -q

# With coverage
pytest --cov=world_model_server --cov-report=term-missing

# Fuzz targets (Atheris, requires Clang / libFuzzer)
python fuzz/fuzz_verify_manifest.py fuzz/corpus/ -max_total_time=60

# Non-Atheris smoke fuzz (runs in every CI pass, no system deps)
pytest tests/test_fuzz_smoke_verify.py -q

值得特别指出的测试套件:

  • FIPS 205 SLH-DSA 已知答案测试(tests/test_fips_205_slh_dsa_kat.py):锁定 SLH-DSA-SHA2-128f 的参数大小(公钥 = 32 字节,私钥 = 64 字节,签名 = 17,088 字节),签名/验证往返测试包含错误密钥 + 错误消息 + 变异签名拒绝,外加一个固定的 KAT 向量夹具(tests/fixtures/slh_dsa_kat_vectors.json),该夹具必须永远验证为真。

  • 流式验证器字节一致性(tests/test_etch_verify_streaming.py):锁定内存导出器与流式导出器之间的字节相同输出,从而无论操作者使用哪个导出器,审计员对作为记录工件的清单进行哈希时都会得到相同的 manifest_sha256。

  • 矛盾基准(benchmarks/contradictions-200/):105 对 × 19 个类别,确定性。在每次提交时将 auto 策略锁定为 100%。

CI 在每次推送时门控覆盖率下限 + 完整测试套件。参见 .github/workflows/pytest.yml。


认证审计链(v0.13+,选择加入)

对于审计跟踪必须可加密验证的部署(SOC 2、HIPAA、EU AI Act 或您自己的内部控制清单):

export WORLD_MODEL_AUDIT_LOG=on
# then start the world-model-mcp server as normal

在首次选择启用时,服务器会在现有的 audit.db 文件中创建两个新的 SQLite 表(tamper_evident_log 和 tamper_evident_epochs),并在第一个 epoch 关闭时生成一个新的混合密钥对。此后每个事件都会被追加到一条 SHA-256 Merkle 链中。当一个 epoch 关闭时(默认:1024 个事件),链根会使用混合 Ed25519 + SLH-DSA-SHA2-128f 信封签名:经典 + 后量子,因此即使椭圆曲线密码学被假设性攻破,审计跟踪仍然完好无损。

  • 可离线验证,永久有效。etch-verify CLI 随 PyPI 包一起发布;审计员在自己的笔记本电脑上运行它,初始下载后无需网络访问。

  • 签名证明作者身份,而不仅仅是顺序。哈希链替代方案可以证明没有任何内容被重新排序,但无法证明究竟是谁签名的。这条链两者都能回答。

  • 无仪表盘、无需注册、无外部服务。完全在你的进程中针对本地 SQLite 运行。

完整说明见 docs/AUDIT_LOG.md。

如需托管版本(增加 KMS 支持的密钥、公开透明度日志、对 Sigstore Rekor + Bitcoin OpenTimestamps 的外部锚定,以及面向合规性的操作员仪表盘),请参阅下方的 Etch 配套服务。


快速开始

三种最常见的安装方式。对于所有其他受支持的客户端(Cursor、Cline、Codex、Continue、Copilot、Windsurf、Goose、pi、OpenClaw、Hermes 等),请参阅 etch.systems/docs/install。

选项 1:Claude Desktop(一键安装)

从 Releases 下载最新的 .mcpb 并将其拖入 Claude Desktop。自动安装 hooks、MCP 服务器配置和依赖项。

选项 2:Claude Code / IDE 插件(pip install)

# 1. Install the package
pip install world-model-mcp

# 2. Set up in your project (auto-seeds the knowledge graph from existing code)
cd /path/to/your/project
python -m world_model_server.cli setup

# 3. Restart Claude Code
# Done. The world model is pre-populated and active.

下一步(可选): 使用托管的 Etch 公证服务将本地签名审计日志转变为审计员可验证的链——KMS 支持的密钥、公开透明度日志、外部锚定,以及供审计员使用的共享链接,无需运行任何基础设施。请参阅 托管配套服务:Etch。

选项 3:用于远程 / MCP 隧道部署的 HTTP 传输

pip install 'world-model-mcp[http]'
python -m world_model_server.server --transport http --port 8000

通过 Streamable HTTP 暴露 MCP,使远程代理可以通过网络连接。有关认证、CORS 和反向代理设置,请参阅 docs/http_transport.md。

其他客户端

通过 OSS CLI 安装。每条命令都会为该客户端写入正确的配置(默认使用 sys.executable 作为解释器路径,感知各客户端格式,通过 --force 和 --dry-run 标志防止覆盖):

python -m world_model_server.cli install-cursor     # Cursor
python -m world_model_server.cli install-cline      # Cline
python -m world_model_server.cli install-codex      # Codex
python -m world_model_server.cli install-continue   # Continue (also --global)
python -m world_model_server.cli install-copilot    # GitHub Copilot Chat (VS Code Insider)
python -m world_model_server.cli install-windsurf   # Windsurf
python -m world_model_server.cli install-pi         # pi
python -m world_model_server.cli install-openclaw   # OpenClaw
python -m world_model_server.cli install-hermes     # Hermes (MCP mode)
python -m world_model_server.cli install-hermes-provider  # Hermes (Elixir-native provider)

每个客户端的完整分步指南(含验证命令 + 故障排查)见 etch.systems/docs/install。


功能概述

world-model-mcp 是一个时间知识图谱,位于你的 AI 编码代理与其工作之间。它记录来自代码库的事实、实体和约束;在编辑边界根据已学习的约束验证每次代码更改;在上下文窗口压缩后重新注入相关上下文;通过置信度加权解析来跟踪矛盾;并通过独立的 Coach LLM 对抗性地验证检索结果。

特性

1. 防止幻觉。 每条记录的事实都带有溯源(asserted_by、confirmer、confirmation_state、evidence_type)。当代理查询某个事实时,它会连同答案一起获得置信度分数。当两个事实相互矛盾时,较新/置信度较高者胜出;落败者以 superseded_by 保留。代理永远不必猜测存储的事实是否仍然有效。

2. 从纠正中学习。 当你纠正代理时(record_correction),该纠正会作为一等事件与相关实体一起存储。下次代理查询任何涉及这些实体的内容时,纠正会首先浮现。纠正会跨会话、跨上下文压缩、跨代理重启持续生效。

3. 防止回归。 每个代码更改提案都会经过 validate_change 检查,该工具会遍历约束图并在编辑落地之前返回违规项。约束从你的代码库自动学习(seed_project),通过 PR 审查评论强化(ingest_pr_reviews),并可手动编写(record_event)。代理会看到违规项、建议以及约束的来源。

4. Coach-Player 对抗性验证。 Player 代理起草决策。独立的 Coach 代理独立地重新查询图谱以寻找先例,遍历 Merkle 证明,检查混合签名,然后才批准。Coach 从不信任 Player 的摘要;它每次都针对签名账本重新验证。在 12 对手工标注的配对中实现 100% 精确匹配。

完整技术架构见 docs/ARCHITECTURE.md。


MCP 工具

提供八个工具。以下为一行摘要;点击每个工具可查看 docs/mcp/ 中的签名 + 示例。

工具

功能

query_fact

获取存储的事实,包含溯源 + 置信度 + 来源

record_event

将事件追加到审计链(event_type 来自固定枚举)

validate_change

根据已学习的约束检查代码更改提案,返回违规项 + 建议

get_constraints

列出与实体或文件模式匹配的所有约束

record_correction

存储用户纠正,使其在下次相关查询时重新浮现

get_related_bugs

获取涉及给定文件的 bug + 每文件风险评分

seed_project

将现有代码库批量导入知识图谱

ingest_pr_reviews

将最近的 PR 审查评论转化为约束

完整工具文档见 docs/mcp/。


托管配套服务:Etch

在本地运行 world-model-mcp?Etch (etch.systems) 是基于同一 OSS 核心构建的托管治理平面,增加了合规团队在批准生产使用所需的功能:

  • KMS 加密的签名密钥(静态存储时永不为明文)

  • 带签名头部的公开透明度日志(防分裂视图)

  • 对 Sigstore Rekor + Bitcoin OpenTimestamps 的外部锚定(双重独立见证)

  • 操作员仪表盘,包含会话跟踪、链完整性视图、PII 扫描和客户端答案 PDF 导出

  • 独立审计员 CLI + 浏览器验证器,位于 /auditor/<slug>

  • 免费层,超出后按需付费

相同的密码学原语、相同的审计日志模式,对已发布的 OSS 零源代码更改。如果你独立运行,可跳过此项。

哪个适合我?

情况

路线

我希望为本地 AI 编码代理提供持久记忆

world-model-mcp(本仓库,MIT)

我希望六个月后能向监管机构证明代理决策

etch.systems

我希望获得法院可以自行验证的签名证据包

etch.systems

我希望实现签名审计历史的跨团队或跨供应商联合

etch.systems

我想两者都试试

先用 world-model-mcp 在本地开始;当你需要可让他人验证的签名证据时再添加 Etch


工作原理

world-model-mcp 是一个 MCP 服务器(stdio 或 Streamable HTTP),它暴露一个由 SQLite 支持的时间知识图谱。每次代理轮次都会查询图谱;每次代码更改或纠正都会写回。启动时,服务器会自动从你现有的代码库进行播种;关闭时,它会干净地刷新。

.claude/world-model/ 下的六个数据库:

  • entities.db:文件、函数、类、符号

  • facts.db:带溯源 + 置信度的语义事实

  • relationships.db:依赖、调用、导入

  • constraints.db:代理必须遵守的已学习规则

  • sessions.db:按会话的上下文跟踪

  • events.db:不可变事件日志(选择启用时为审计链)

包含图表的完整架构见 docs/ARCHITECTURE.md。


配置

环境变量(均为可选):

变量

用途

默认值

WORLD_MODEL_AUDIT_LOG

启用签名审计链

off

WORLD_MODEL_DB_PATH

覆盖 SQLite 数据库目录

.claude/world-model

WORLD_MODEL_TELEMETRY

选择加入匿名 OSS 遥测(请参阅隐私)

off

ANTHROPIC_API_KEY

仅在你启用可选的 LLM 支持功能时使用

none

完整配置参考见 docs/CONFIGURATION.md。


隐私与安全

  • 遥测默认关闭。 如果你使用 WORLD_MODEL_TELEMETRY=on 选择加入,聚合的安装级指标会发送到 etch.systems/api/telemetry/ingest(端点 URL,不包含源代码、提示词或 PII)。通过 DELETE /api/telemetry/install/{install_id} 支持被遗忘权。

  • 核心操作不需要 ANTHROPIC_API_KEY。 如果提供了密钥,某些可选功能(Coach-Player 第 3 层验证、LLM 支持的重新排序)会使用该 API。没有它,其他一切照常工作。

  • 报告安全问题: 发送邮件至 security@etch.systems。可按需提供 PGP 密钥。

完整说明见 docs/PRIVACY.md。


贡献

欢迎贡献。请参阅 CONTRIBUTING.md 了解开发环境搭建、编码标准、添加语言支持、编写测试和提交 PR。

特别需要帮助的领域:

  • 语言解析器(Go、Rust、Java、C++)

  • 额外的 MCP 客户端适配器

  • 框架集成(LangGraph、CrewAI、AutoGen、LlamaIndex;托管仓库中已有入门 shim)

  • benchmarks/ 中的基准测试贡献

请在提交第一个 PR 之前阅读 CLA.md。


许可证

MIT 许可证。免费用于商业和个人用途。


链接

Available Tools

31 tools
export_claude_mdB

Generate a CLAUDE.md document from the knowledge graph (top constraints, recent decisions, known bug regions, co-edit patterns).

ParametersJSON Schema
NameRequiredDescriptionDefault
max_constraintsNo

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavior. It states it generates a document but does not clarify whether it writes to a file, returns content, or has side effects. It also fails to mention any permissions, limits, or output format, leaving key behavioral aspects ambiguous.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with a clear verb and object, followed by a parenthetical list of content categories. It is concise and contains no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description should explain return values, side effects, and parameter behavior. It only covers purpose and content categories, leaving critical operational details absent for an export tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for the only parameter, max_constraints. The description broadly references 'top constraints' but does not explain that max_constraints limits that section, nor its units or effect on other sections. The parameter semantics are only indirectly inferable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Generate') and identifies the resource ('CLAUDE.md document from the knowledge graph'), and it lists the included content categories (top constraints, recent decisions, known bug regions, co-edit patterns), distinguishing it from sibling tools that query individual aspects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for producing an aggregated CLAUDE.md from knowledge graph data, but it does not explicitly state when to use this tool versus querying individual components via siblings like get_constraints or get_decision_log. It lacks explicit exclusions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_contradictionsC

Find pairs of facts that contradict each other based on similarity and status differences

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of disclosing behavior. It doesn't state whether the tool is read-only, what the output looks like, or any side effects. The hint about 'similarity and status differences' is the only behavioral clue, but it's insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise and front-loaded, but it sacrifices necessary detail. It's not bloated, but it's under-specified. The brevity doesn't earn its place because it lacks critical information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has two parameters, no output schema, and no annotations. The description provides no information about return values, parameter usage, or behavioral context. It is far from complete even for a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has two parameters (limit, query) with zero descriptions, and the description doesn't mention them at all. Since schema coverage is 0%, the description fails to compensate, leaving the agent with no understanding of what these parameters do.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool finds pairs of contradicting facts, using a specific verb ('Find') and resource. It adds a hint of the method ('similarity and status differences'), but doesn't explicitly differentiate from sibling tools like resolve_contradiction, though the distinction is evident from the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus alternatives. It doesn't mention any exclusions or prerequisites, leaving the agent to infer usage solely from the vague description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agents_md_constraintsA

Parse AGENTS.md / CLAUDE.md / GEMINI.md / .agents/skills/*.md in the project and return declarative constraints. Mixed into PreToolUse enforcement automatically; this tool exposes the same data for inspection.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathNo
project_dirNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden for behavioral disclosure. It explains the parsing action and its role in enforcement, but does not mention whether it performs a read-only operation, how missing files are handled, or any potential side effects. The context about automatic enforcement adds value but lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the main action and scoped to specific files. Every phrase contributes meaning, and the inspection/enforcement contrast adds valuable context without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's purpose and its relationship to PreToolUse enforcement, giving useful context. However, it omits parameter semantics and the exact return format (beyond 'declarative constraints'), which is especially important since there is no output schema. The lack of detail on expected inputs and outputs leaves gaps for an agent trying to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has two parameters (file_path and project_dir) with no descriptions and 0% coverage. The description does not mention either parameter, leaving their specific purpose and format ambiguous. For example, it is unclear if file_path is optional or relative to project_dir. The description fails to compensate for the complete lack of schema guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it parses AGENTS.md / CLAUDE.md / GEMINI.md / .agents/skills/*.md and returns declarative constraints. This specific verb+resource distinguishes it from the sibling tool get_constraints, which likely covers broader constraint sources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes that the data is 'Mixed into PreToolUse enforcement automatically' and that this tool 'exposes the same data for inspection,' implying it is intended for inspection/debugging rather than direct enforcement. It provides clear context but does not explicitly name alternatives or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_audit_log_headA

v0.13 tamper-evident audit log. Return the current head state (last log entry seq, last closed epoch seq, unclosed-entry count) plus the full closed-epoch chain with hybrid signature envelopes. Compliance auditors call this periodically to verify no operator misbehavior has occurred since the last check. Requires WORLD_MODEL_AUDIT_LOG=on at server startup.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the tamper-evident nature, the return content, and a runtime prerequisite (WORLD_MODEL_AUDIT_LOG=on). It adds meaningful operational context beyond a simple 'get' but stops short of stating read-only semantics explicitly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two sentences but packs in the purpose, return value details, use case, and a critical prerequisite. Every clause earns its place with no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter, read-only operation with no annotations or output schema, the description covers all essential aspects: what it does, what it returns, when to use it, and what server configuration is required. An agent has enough to deploy it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the schema is empty. The description explains what the tool returns, which is the only meaningful semantic content in this case. Baseline 4 applies since there are no parameters to describe.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Return' and names the resource 'audit log head state' plus the 'closed-epoch chain', clearly defining what the tool does. This differentiates it from other audit-related siblings like get_compaction_audit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states a use case: 'Compliance auditors call this periodically to verify no operator misbehavior has occurred since the last check.' This gives clear context on when to use it, but it does not mention alternatives or exclusions, so it misses the top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_co_edit_suggestionsB

Get files commonly edited alongside the given file based on historical patterns

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
file_pathYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden for behavioral disclosure. It mentions 'based on historical patterns' which adds some context, but it does not describe whether the operation is read-only, how suggestions are ranked, what format the results take, or any potential side effects. This is insufficient for full transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that is front-loaded with the action and resource. Every word earns its place, and there is no superfluous content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with two parameters, but the description leaves major gaps: no usage guidance, no parameter details, and no mention of return values (no output schema). For an agent to use this correctly, it needs more context about how suggestions are generated and what to expect from the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero description coverage (0%) for its parameters, and the description does not compensate. It references 'the given file' (mapping to file_path) but does not clarify the expected format, the meaning or usage of 'limit', or any constraints. The description adds minimal value beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'files commonly edited alongside the given file', with the basis 'historical patterns'. It is specific and distinguishes itself from the sibling tools, none of which share a similar focus on co-edit suggestions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any exclusions, prerequisites, or conditions, leaving the agent without context for selecting it among the many sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_compaction_auditA

List recent compaction audit entries, most-recent first. Filter by session_id or limit count.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
session_idNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the transparency burden. It discloses the ordering (most-recent first) and the available filters (session_id, limit), which is useful. However, it does not mention whether the operation is read-only, whether limit has a default, or describe the return format or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that delivers all key information without redundancy. Every clause adds functional value, and there is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list/filter tool, the description covers the essential purpose and options. However, with no output schema and no annotations, the agent is left without knowledge of the returned fields, default limit behavior, or error cases. This is adequate for simple use but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must add meaning. It does clarify that session_id filters and limit controls the count, which is helpful. However, the semantics are shallow: it does not specify whether limit is mandatory, its maximum/default value, or whether session_id requires exact or partial matching.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists compaction audit entries in reverse chronological order. It names the specific resource (compaction audit entries) and the action (list), which differentiates it from write-oriented siblings like record_compaction_audit. However, it does not explicitly distinguish from the similar get_audit_log_head tool, preventing a perfect score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for inspecting compaction audit history, and the mention of filtering suggests relevant scenarios. However, it provides no explicit guidance on when to prefer this over alternatives like get_audit_log_head, and there are no exclusions or prerequisites stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_constraintsA

Get constraints (linting rules, patterns, conventions) for a file

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes
constraint_typesNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It implies a read-only operation ('Get') but does not disclose potential behaviors such as error handling for missing files, whether it searches the entire repository, or if any filtering is applied beyond the optional constraint_types parameter. The description adds some clarity by defining constraints, but does not reveal side effects or edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that efficiently communicates the tool's purpose without extraneous words. It earns a high score for being concise and structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (read operation, 2 parameters, no output schema), the description is mostly complete. It states what the tool does and the schema covers parameter details. However, it does not specify the return format or behavior when no constraints are found, which could leave some ambiguity for the agent. Still, for a straightforward getter, it is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, so the description must compensate. It does so by explaining that constraints include 'linting rules, patterns, conventions', which helps interpret the constraint_types enum. However, it does not explicitly map parameters or explain the file_path semantics beyond the name. The schema itself provides the enum values, offering adequate baseline coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves constraints for a file, using the verb 'Get' and specifying the resource and scope. It also clarifies what constraints are (linting rules, patterns, conventions). However, it does not explicitly distinguish from the sibling tool 'get_agents_md_constraints', which may overlap in purpose for specific files.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when constraints for a file are needed, but provides no explicit guidance on when to use this tool versus alternatives like 'get_agents_md_constraints' or when not to use it. It lacks exclusions and alternative recommendations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_context_for_actionC

Pre-action context bundle: constraints, decisions, bugs, co-edits, related facts, and risk score for a file before editing

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes
action_typeYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It lists the content components but does not state whether the tool is read-only, how errors are handled, what the risk score means, or how the bundle is returned. This is a significant gap for a retrieval tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a concise noun phrase that front-loads the core purpose and lists contents, but it lacks a verb and reads more like a label than a full sentence. It is not overly verbose, yet it could be more structured with a clearer main clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two parameters, no annotations, and no output schema, the description is too sparse. It does not explain the output format, the meaning of the risk score, or how this bundle relates to the individual context tools. The description leaves significant gaps in understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage. The description mentions 'file' which maps to file_path, but action_type is only implied by 'editing' and not explicitly explained. The enum values (edit, create, delete, refactor) are not described, so the description adds minimal meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a pre-action context bundle for a file, listing specific content types such as constraints, decisions, bugs, co-edits, related facts, and risk score. This distinguishes it from sibling tools that target single context types. However, the phrase 'before editing' is slightly inconsistent with the action_type enum which includes create, delete, and refactor.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide explicit guidance on when to use this bundle versus the many sibling tools like get_constraints or get_related_bugs. It implies usage before an action, but there are no exclusions or alternative recommendations, leaving the agent to infer the appropriate context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_decision_logC

Get decision traces showing agent proposals and human corrections

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
file_pathNo
session_idNo
decision_typeNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It implies a read operation via 'Get', but does not disclose behavior such as default limits, ordering, filtering effects, or whether it is purely read-only. The description focuses on content, not operational behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no wasted words. It efficiently states the purpose without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has four parameters, no output schema, and no annotations, yet the description provides no details on return value, parameter usage, or edge cases. It is overly minimal for the tool's apparent complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description does not explain any parameters. It only hints at decision_type through 'corrections', but limit, file_path, and session_id are completely unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves 'decision traces' with specific content (agent proposals and human corrections), using a specific verb 'Get' that distinguishes it from sibling write tools like record_decision and record_correction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as get_audit_log_head or get_compaction_audit. There is no mention of scenarios, exclusions, or preferred contexts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_health_reportA

Memory health diagnostics: orphans, stale facts, contradictions, decay candidates, DB sizes

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden. The term 'diagnostics' implies a read-only operation, and the listed categories provide concrete insight into what the tool examines. However, it does not explicitly state whether any state is modified or note any performance implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise phrase that effectively communicates purpose and scope without wasted words. It is front-loaded and every element adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter diagnostic tool with no output schema, the description provides sufficient context about the report's contents. It could be improved by noting the output format, but the current level is adequate for understanding the tool's purpose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description need not document inputs. The baseline of 4 is appropriate because no parameter information is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function as memory health diagnostics and enumerates specific areas it covers (orphans, stale facts, contradictions, decay candidates, DB sizes). This distinguishes it from narrower sibling tools like find_contradictions or get_compaction_audit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use for overall health assessment, but it does not explicitly state when to prefer this tool over siblings or when not to use it. No alternative tools are mentioned, so guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_injection_contextB

Return a compact constraint+fact bundle for PostCompact / UserPromptSubmit hooks to re-inject after context loss.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_factsNo
event_typeYes
project_hintNo
max_constraintsNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden. It only says 'Return...' without disclosing whether this is a read-only operation, any side effects, how 'compact' is achieved (e.g., truncation, filtering), or what happens if parameters like max_facts are omitted. This lack of behavioral detail is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It is concise and easy to read, though it sacrifices substance for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 4 parameters, zero schema coverage, no annotations, and no output schema, the description is insufficiently complete. It gives a high-level purpose but lacks details on parameter behavior, return format, and how this tool fits with alternatives. More context is needed for an agent to use it reliably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it does not. It mentions two of the three event_type values but does not explain max_facts, max_constraints, project_hint, or the SessionStart event. The term 'compact' weakly implies size limits, but no explicit parameter meaning is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a 'compact constraint+fact bundle' for specific hooks (PostCompact/UserPromptSubmit), which distinguishes it from sibling tools like query_fact or get_constraints. However, it could be more explicit about what 'context loss' entails or how it differs from get_context_for_action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names specific trigger events (PostCompact/UserPromptSubmit) and the goal of re-injecting after context loss, giving clear context for when to use the tool. It does not explicitly mention alternatives or when not to use it, but the event-specific framing provides adequate guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ingest_pr_reviewsA

Pull GitHub PR review comments and convert them into learned constraints in the knowledge graph

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoGitHub repo (owner/repo). Auto-detected from git remote if omitted.
countNoNumber of recent PRs to scan (default 10, max 50)

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses that the tool converts PR comments into learned constraints, indicating a write operation to the knowledge graph. However, it does not mention idempotency, overwrite behavior, permissions, or any side effects beyond the conversion, which is minimal for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that directly states the action and outcome with no filler. Every word contributes to the meaning, making it highly efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has only two optional parameters and no output schema, so minimal description might suffice. However, because it is an ingest operation affecting the knowledge graph, a bit more context about result expectations or side effects would improve completeness. It is adequate but not fully comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds no extra parameter-specific meaning beyond what the schema already provides (repo auto-detection, count default/max). It does not compensate for any missing details, but none are missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('Pull' and 'convert') and identifies the resource ('GitHub PR review comments') and target ('knowledge graph'). It clearly distinguishes itself from sibling tools like get_constraints or record_event by describing a unique ingest-and-transform workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when you want to bring PR review comments into the knowledge graph, but it does not explicitly state when to prefer this over alternatives or provide exclusions. Sibling tools like record_correction or validate_change serve different purposes, yet no direct comparison is offered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pin_annotationA

Attach a signed human annotation (note, override rationale, or intervention record) to a span of agent events. Persists into the annotations table and chains into the same Merkle audit log as agent writes (v0.15.0, ADR-0001). Rationale limited to 8 KB.

ParametersJSON Schema
NameRequiredDescriptionDefault
authorYesAuthor identity. Self-asserted in OSS; KMS-verified in Etch hosted.
rationaleYesHuman rationale text (UTF-8). Max 8192 bytes.
session_idYesSession containing the annotated events.
annotation_typeYes
event_range_endYesLast event_id in the annotated span. Equals event_range_start for a single-event annotation.
event_range_startYesFirst event_id in the annotated span.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and exceeds expectations by disclosing persistence ('Persists into the annotations table'), audit integration ('chains into the same Merkle audit log as agent writes'), version/ADR references, and the rationale size limit. This gives the agent a clear behavioral model.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core action, and includes only essential extra context (persistence, audit log, limit). Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 6 required parameters, no output schema, and no annotations, the description provides solid context for a write operation: it explains persistence and audit chaining. It lacks explicit error handling or return behavior, but that is not critical for a basic mutation tool. The description is largely complete for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 83% coverage (5 of 6 parameters have descriptions), so the baseline is 3. The description adds little beyond the schema; it mentions the 8 KB rationale limit (already in schema) and does not elaborate on parameter meanings or relationships. The schema itself is well-documented, so a baseline score is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action ('Attach') and object ('a signed human annotation to a span of agent events'). It distinguishes from siblings like record_event by emphasizing human annotation vs agent events and mentions specific annotation types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case of attaching human notes or overrides to event spans, but provides no explicit guidance on when to choose this tool over alternatives such as record_correction or record_decision. The context is clear but lacks exclusions or alternative recommendations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

predict_regressionA

Score regression risk for a proposed change to a file based on past bugs, test failures, and constraint violations

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes
change_descriptionNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses what the tool uses (past bugs, test failures, constraint violations) but does not state whether it has side effects, permissions requirements, or what the output looks like. The methodology hint adds some behavioral context, but safety and operational traits remain undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It efficiently conveys the tool's purpose and key inputs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two simple parameters and no output schema, the description is adequate for tool selection but does not fully prepare the agent for invocation. It lacks details on input semantics and return value format, though the simplicity and sibling context make it acceptable but not thorough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions 'proposed change to a file' which loosely maps to file_path and change_description, but it does not clarify parameter formats, required fields, or examples. The description adds minimal meaning beyond what the property names already imply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Score') and resource ('regression risk') with the basis ('past bugs, test failures, and constraint violations'). It clearly distinguishes from siblings like predict_test_failures (which targets test failures specifically) and simulate_change (which simulates changes).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for assessing regression risk of a proposed change to a file, but it does not explicitly state when to prefer this tool over alternatives or provide exclusions. Sibling tools like predict_test_failures or validate_change could overlap, and no guidance is given on choosing among them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

predict_test_failuresA

Surface tests likely to fail given a set of edited files

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathsYes

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the full burden for behavioral transparency. It only states the core action without disclosing whether the operation is read-only, what data it relies on, what output format to expect, or any limitations. The description is minimal and leaves the agent without important contextual cues.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the action and avoids redundancy. Every word contributes to the core purpose, making it highly concise and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is relatively simple (one parameter, no output schema, no annotations), but the description is thin. It communicates the essential purpose, but does not explain what the tool returns (e.g., a list of test names) or provide operational boundaries. It is minimally complete for selection but not fully adequate for invocation without further inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no descriptions for the parameter file_paths (0% coverage). The description adds semantic value by indicating these are 'edited files', clarifying the parameter's intent. However, it does not specify path format, file existence requirements, or other constraints that would be useful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Surface' with a clear resource ('tests likely to fail') and a scope condition ('given a set of edited files'). This distinguishes it from sibling tools like predict_regression and get_related_bugs, which address different aspects of change impact.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'given a set of edited files' implies the primary use case, but the description does not explicitly state when to use this tool over alternatives or provide any exclusions. It lacks guidance on how to compare with similar tools like predict_regression.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

promote_constraintB

Promote a constraint from this project to all other registered projects

ParametersJSON Schema
NameRequiredDescriptionDefault
constraint_idYes
target_projectsNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals the action 'promote' but does not disclose side effects, permissions required, reversibility, or the fact that it likely mutates multiple projects. This leaves significant behavioral uncertainty.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the verb and object. It contains no filler or unnecessary detail, making it easy to parse and understand.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema and no annotations, and the description is minimal. It lacks essential context about the promotion's effects, error conditions, prerequisites, and return values. For a cross-project mutation tool, this is insufficient for an agent to use safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for parameter meaning. It implies constraint_id identifies the constraint, but it does not explain target_projects, its optionality, or how it interacts with the default 'all other registered projects'. The description adds minimal meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'promote' with a clear resource, 'constraint', and a clear scope, 'from this project to all other registered projects'. This distinguishes it from sibling tools like get_constraints (retrieval) and validate_change (validation), making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidance is provided. The description does not mention when to use this tool versus alternatives, nor any prerequisites or conditions for promotion. It only states the action itself, leaving the agent without context for decision-making.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prove_entry_inclusionA

v0.13 tamper-evident audit log. Return a cryptographic inclusion-proof bundle for a persisted row_id (fact, constraint, event, or decision ID). Bundle includes the entry, the containing signed epoch (Ed25519 + SLH-DSA hybrid signature envelope), an RFC 6962 Merkle inclusion proof, and the full epoch chain from genesis. Requires WORLD_MODEL_AUDIT_LOG=on at server startup; returns an error object when opt-in is off, when the row_id is not found, or when the entry is in the unclosed backlog.

ParametersJSON Schema
NameRequiredDescriptionDefault
row_idYesID of the fact / constraint / event / decision to prove inclusion for.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses the exact bundle components (entry, signed epoch, Merkle proof, epoch chain), required server flag, and all error cases, giving a thorough behavioral profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single information-dense sentence that front-loads the primary action ('Return a cryptographic inclusion-proof bundle') and then enumerates bundle contents and failure modes without unnecessary filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by listing the bundle components explicitly. It also covers prerequisites and error scenarios, making it complete enough for a single-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (row_id with description), but the description adds semantic value by specifying that row_id refers to fact, constraint, event, or decision ID, clarifying the expected input beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a cryptographic inclusion-proof bundle for a persisted row_id. It names the specific resource (audit log entries) and distinguishes it from siblings like get_audit_log_head or query_fact by focusing on inclusion proofs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use it (to prove inclusion of an entry), prerequisites (WORLD_MODEL_AUDIT_LOG=on), and error conditions (opt-in off, row_id not found, unclosed backlog). It does not explicitly mention alternatives, but the scope is clear enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_factB

Query the knowledge graph for facts about entities (APIs, functions, classes, etc.)

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesThe search query (e.g., 'User.findByEmail', 'JWT authentication')
contextNoAdditional context for the query
entity_typeNoOptional filter by entity type
content_typeNoOptional filter by content_type. Use 'procedure' to explicitly summon procedures (which are excluded from auto-injection by design).

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full responsibility for behavioral disclosure. It doesn't state whether the tool is read-only, what the response format is, or that it can retrieve 'rules' and 'procedures' despite being named 'facts'. The term 'facts' may mislead users into thinking only fact-type content is returned, leaving important behavior undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that is front-loaded with the action and resource. Every word contributes to the core purpose without fluff or repetition. It is appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a general-purpose query tool with no output schema and no annotations, the description is minimal but adequate. It does not explain return values, pagination, or how to decide between this and related sibling tools. The ambiguous scope of 'facts' (vs rules/procedures) is a notable gap, but the schema partially covers this via content_type descriptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema coverage is 100% with descriptive text for all parameters, so the baseline is 3. The tool description adds no parameter-level meaning beyond what the schema already provides; it only reiterates the general entity focus. The description does not compensate or enhance the schema's parameter docs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a clear verb ('Query') and resource ('knowledge graph') and specifies the object ('facts about entities'). It lists example entity types (APIs, functions, classes) which conveys scope. However, it doesn't distinguish itself from sibling tools like search_global that might also search the knowledge graph, so it lacks explicit differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance in the description about when to use this tool versus alternatives. The only usage hint ('Use "procedure" to explicitly summon procedures...') appears in the schema's content_type parameter, not the tool description, and it is parameter-level rather than tool-level. No when-not-to-use or alternative tool references are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recall_transcript_rangeC

Hydrate a Claude Code session transcript by line range. Lets agents trace a fact back to the exact conversation that produced it.

ParametersJSON Schema
NameRequiredDescriptionDefault
line_endNo
line_startNo
session_idYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only says 'Hydrate' without explaining whether the operation is read-only, what happens with invalid ranges, whether there are performance or memory implications, or what the response contains. This is a significant gap for a data-retrieval tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences and front-loaded with the core action. Each sentence contributes meaning: the first states the operational scope, the second provides the motivating use case. It is well-structured and free of fluff, though slightly vague.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and a 0% parameter coverage, the description should provide more contextual information about return values, error handling, and when to choose this tool. The current description is insufficient for an agent to confidently invoke the tool correctly in all cases, even though the tool itself is relatively simple.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions 'by line range,' which somewhat clarifies line_start and line_end, but it does not explain session_id, the inclusive/exclusive nature of the range, defaults, or behavior when only one line parameter is provided. The description adds minimal semantic value beyond the schema fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool hydrates a Claude Code session transcript by line range, with a specific use case of tracing facts to their source conversation. This distinguishes it from sibling tools like query_fact, which focus on facts rather than raw transcript retrieval. However, the verb 'hydrate' is somewhat jargon-heavy, and it lacks an explicit contrast with related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool ('Lets agents trace a fact back to the exact conversation that produced it'), which provides some context. However, it does not explicitly state when to prefer this over alternatives like query_fact or get_audit_log_head, nor does it give exclusions or prerequisites. The guidance is present but inferred rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_compaction_auditA

Record a context-compaction event with token counts and what was re-injected. Lets developers audit what was remembered across compaction boundaries.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNo
raw_summaryNo
facts_injectedNo
injection_eventNo
pre_compact_tokensNo
post_compact_tokensNo
constraints_injectedNo

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description must disclose behavioral traits. It states it records an event, implying a write operation, but does not mention side effects, whether data is appended or overwritten, permissions, or failure modes. This is a significant gap for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exactly two sentences, front-loaded with the core action and followed by the purpose. Every word adds value with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 optional parameters, no output schema, and no annotations, the description does not fully equip an agent to use the tool. It lacks information about expected return values, whether any parameters are required in practice, and how this differs from other record_* tools beyond the compaction focus. This is insufficient for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It mentions 'token counts' and 'what was re-injected,' which roughly maps to pre_compact_tokens, post_compact_tokens, facts_injected, and constraints_injected, but it does not explain parameters like injection_event, session_id, or raw_summary. This leaves meaningful ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Record' with the resource 'context-compaction event' and specifies token counts and re-injected content. It clearly distinguishes from siblings like get_compaction_audit (retrieval) and record_event (generic event) by focusing on compaction-specific auditing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: this is for recording compaction events for audit purposes. It does not explicitly mention alternatives or when not to use it, but the specialized language makes the use case evident. There are no exclusions or alternative references, but the context is strong enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_correctionC

Record a user correction to Claude's output (high-priority learning signal)

ParametersJSON Schema
NameRequiredDescriptionDefault
reasoningNoInferred reason for the correction
session_idYes
claude_actionYesWhat Claude did (tool, file, content)
user_correctionYesHow the user corrected it

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states the action is a 'high-priority learning signal', which hints at importance but does not describe side effects, persistence, reversibility, permissions, or any consequences of invoking the tool. This is a significant gap for a mutation-like recording tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, grammatically complete sentence of about ten words. It is front-loaded with the core verb and resource, and the parenthetical adds context without redundancy. Every word earns its place; there is no fluff or tail-heavy content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool involves nested object parameters, no annotations, and no output schema, yet the description stays at a high level. It does not explain how to structure claude_action or user_correction, what counts as a valid correction, or what the tool returns. This leaves an agent under-informed for correct invocation, especially given the complexity of the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, and the description adds no additional meaning beyond what the schema already provides. The phrase 'user correction to Claude's output' loosely maps to claude_action and user_correction but does not clarify their structure or relationships. The 'reasoning' and 'session_id' parameters are not addressed at all in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Record') and identifies the resource ('a user correction to Claude's output'), which clearly conveys the tool's function. The parenthetical '(high-priority learning signal)' adds useful context. However, it does not explicitly differentiate from sibling tools like record_event or record_decision, though 'correction' provides some distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool should be used when a user corrects Claude's output, but it gives no explicit 'when to use' vs. alternatives, no exclusions, and no prerequisites. Sibling tools with overlapping purposes (e.g., record_event, record_decision) are not referenced, leaving the agent without clear selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_decisionC

Record a decision trace: what the agent proposed and how the human responded

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathNo
reasoningNo
tool_nameNo
session_idYes
decision_typeYes
agent_proposalNo
human_correctionNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the full burden of behavioral disclosure. It explains what is recorded but not whether the operation is append-only, idempotent, permission-sensitive, or what happens on conflict. No mutation or safety details are given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that is front-loaded with the core action and object. No filler or redundant wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters including nested objects and required fields, the description is far too minimal. It leaves the agent without guidance on required inputs, the decision_type enum, or the structure of nested objects, and there is no output schema to aid expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description only hints at 'proposed' and 'human responded', which loosely map to agent_proposal and human_correction. It does not explain required parameters like session_id or decision_type, nor the enum values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool's purpose: recording a decision trace with the agent's proposal and human's response. The verb 'record' and resource 'decision trace' are specific, though it doesn't explicitly distinguish from siblings like record_correction or record_event.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided for when to use this tool versus the many sibling tools (e.g., record_correction, record_event). The description does not mention contexts, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_eventC

Record a development event (file edit, test run, etc.)

ParametersJSON Schema
NameRequiredDescriptionDefault
successNo
entitiesNoEntity names/paths involved
evidenceNoTool inputs/outputs, file contents, etc.
reasoningNo
event_typeYes
session_idYes
descriptionYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It only says 'Record' implying a write operation, but doesn't explain persistence, idempotency, success/failure effects, or whether it appends to a log. This is insufficient for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short sentence, which is concise, but it under-specifies the tool's behavior. While brevity is good, the sentence doesn't earn its place by providing necessary context, making it closer to under-specification than effective conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 7 parameters, no output schema, and no annotations, the description is far from complete. It doesn't explain required fields, return values, or how the event data is used. The description is only a high-level purpose statement, inadequate for safe and correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 29%, so the description must compensate for missing parameter details. It adds examples for event_type ('file edit, test run'), which is redundant with the enum, but it doesn't explain session_id, description, success, or reasoning. The description adds minimal semantic value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the general action ('Record a development event') with examples, so it's more than a tautology. However, it's vague about what constitutes a 'development event' and doesn't distinguish from sibling record tools like record_test_outcome or record_correction, which likely overlap in purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternative record tools. The description doesn't mention any conditions, prerequisites, or exclusions, leaving the agent without direction for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_test_outcomeC

Record test results and link failures to recent code changes

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes
test_resultsYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It implies a write operation but does not disclose side effects of linking failures, whether it is idempotent, or if it requires an existing session. This lack of behavioral detail leaves the agent uncertain about the tool's impact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler or unnecessary details. It efficiently conveys the main purpose and a secondary linking behavior, making it appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no annotations, no output schema, and a moderately complex input schema with a nested array. The description omits critical information about expected input formats, return behavior, and how the linking works. An agent would likely need additional schema inspection or external knowledge to use this tool confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not reference session_id or test_results at all. It fails to explain what session_id should be or how to structure the test_results array. The schema provides names/types, but the description adds no semantic value, making it insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the core action ('Record test results') and adds a distinguishing feature ('link failures to recent code changes') that separates it from sibling tools like record_event or record_decision. It is not a mere tautology because it specifies the linking behavior, though it could be more explicit about the scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool vs alternatives. It does not mention prerequisites, exclusions, or contexts where other record tools would be more appropriate. The agent must infer usage from the name and minimal description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_contradictionC

Pick a winner between two contradicting facts using a confidence-weighted strategy (auto, keep_higher_confidence, keep_most_recent, keep_most_sources, supersede_a, supersede_b, manual).

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNo
strategyNo
fact_a_idYes
fact_b_idYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of disclosing behavioral traits. It only mentions the strategy selection and does not disclose side effects, persistence, reversibility, or what happens to the losing fact. This is a material gap for a tool that presumably modifies state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that immediately states the primary action and enumerates strategies in a parenthetical list. Every element serves a purpose, with no fluff or redundant repetition of schema information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations, output schema, and any behavioral details, the description is far from complete. It fails to explain return values, the outcome of the resolution (e.g., which fact is updated), or any prerequisites. This is a mutation-like operation with substantial missing context, making the tool risky to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides no descriptions (0% coverage), so the description must compensate. It lists possible strategy values, which adds some meaning, but it does not explain the semantics of each strategy (e.g., what 'auto' does) or describe the 'fact_a_id', 'fact_b_id', and 'notes' parameters beyond their schema names. The compensation is partial and insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Pick a winner' and identifies the resource as 'two contradicting facts,' clearly indicating the tool's purpose. It differentiates from sibling tools like find_contradictions by focusing on resolution rather than detection, though it doesn't explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a contradiction exists between two facts and provides a list of strategies, which serves as guidance on how to resolve. However, it lacks explicit exclusions or directives about when not to use this tool versus alternatives, leaving some ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_globalC

Search entities across all registered world-model projects

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavioral traits. It does not mention whether the operation is read-only, how results are returned, or any side effects, leaving significant ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no redundant words. It is efficiently phrased, though it lacks additional structured information that a more complete description might include.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with no output schema and no annotations, the description is minimally sufficient but incomplete. It omits return format, pagination behavior, and the meaning of 'entities', leaving the agent without critical execution context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema includes 'query' and 'limit' with no descriptions, and the description adds no parameter-specific information. It does not explain what query syntax is expected or how limit affects results, failing to compensate for the 0% schema description coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (search), the resource (entities), and the scope (across all registered world-model projects). It distinguishes itself from potential siblings by emphasizing the global scope, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like query_fact. The description only states what it does, leaving it to the agent to infer appropriate usage without any exclusions or comparative context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

seed_projectA

Scan the project codebase and populate the knowledge graph with entities and relationships from existing code

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoRe-seed already processed files
project_dirNoProject directory path (defaults to current)

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states that the tool populates the knowledge graph but does not disclose whether the operation is idempotent, whether it modifies existing data, or any side effects. The 'force' parameter hints at re-seeding but this behavior is not explained in the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, direct, and front-loaded with the main action. No fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema or annotations, the description is the sole source of behavioral context. It covers the high-level operation but lacks guidance on prerequisites, idempotency, return values, or potential side effects of running a mutating operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters (force and project_dir), so the description does not need to compensate. The description itself adds no parameter semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Scan the project codebase and populate the knowledge graph with entities and relationships from existing code'. It uses a specific verb (scan/populate) and resource (codebase, knowledge graph), distinguishing it from siblings like query_fact (query) or record_event (record).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the tool name 'seed_project' (initial population), but the description provides no explicit when-to-use guidance or alternatives. No mention of when to run this versus other ingestion tools like ingest_pr_reviews or record_event.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

simulate_changeC

Project blast radius and historical outcomes for a proposed change

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes
change_descriptionYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states what the tool projects but does not indicate whether it is read-only, whether it requires specific permissions, or what outputs to expect. The lack of any side-effect or limitation details is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that front-loads the core action and object. Every word contributes meaning, and there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description should explain what the tool returns or how the output should be interpreted. It does not, and it also omits any caveats or prerequisites. For a simulation tool, this leaves the agent without enough context to use it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the tool description adds no parameter-level detail. The parameter names 'file_path' and 'change_description' are somewhat self-explanatory, but the description does not clarify expected formats, relationships, or how they map to the 'proposed change' concept, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Project') and resource ('blast radius and historical outcomes for a proposed change'), making the purpose clear. It does not explicitly distinguish from sibling tools like 'predict_regression' or 'validate_change', but the focus on blast radius and historical outcomes is distinctive enough for a 4.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. There is no mention of conditions like 'use when you need to assess impact before applying a change' or references to sibling tools, leaving the agent to infer usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_changeB

Validate a proposed code change against known constraints

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes
change_typeYes
proposed_contentYesThe new content to validate

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It only states the validation intent but doesn't explain whether validation is read-only, what happens on failure, whether it modifies anything, or what 'known constraints' refers to. This lack of transparency for a tool that could have side effects is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short sentence that gets to the point. However, at 8 words, it is extremely terse and could have expanded to include usage context without becoming verbose. It is efficient but slightly under-specified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema and no annotations, the description should explain what the validation result looks like and when to use the tool. It only provides the core purpose, missing critical context about return values, behavior on constraint violation, and relationship to other constraint-related tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers 33% of parameters with descriptions (only proposed_content). The tool description does not mention file_path, change_type, or explain the enum values. It adds no semantic value beyond the schema, leaving file_path and change_type under-specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'validate' with a clear object 'proposed code change' and scope 'against known constraints,' which distinguishes it from sibling tools like simulate_change (which implies running a simulation). It clearly states what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used for checking changes against constraints, but it does not explicitly state when to use it over simulate_change or how it relates to get_constraints. No alternatives are named, and no when-not-to-use guidance is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_retrievalA

Adversarially verify an answer is grounded in a specific set of facts. An independent Coach LLM call checks each material claim in the answer against the supplied source facts and returns confidence (HIGH / MEDIUM / LOW), verified + unverified claim lists, and per-claim source_pointers. Never raises; failures return LOW + error populated. v0.12.12.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesThe user query the answer responds to
answerYesThe candidate answer under verification
fact_idsYesIDs of facts the caller believes ground the answer. Missing IDs are silently dropped.
verification_modelNoOptional Coach model override. Defaults to config.verification_model (Haiku 4.5).

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses the independent Coach LLM call, the confidence levels (HIGH/MEDIUM/LOW), the verified and unverified claim lists, per-claim source_pointers, and that it never raises (failures return LOW with error populated). This is excellent behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the core purpose, followed by behavior and return details. The version tag 'v0.12.12' adds minor noise but does not detract significantly. Overall, every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Since there is no output schema, the description fully explains return values (confidence, claim lists, source_pointers) and error behavior. It is complete for a verification tool, covering what the agent needs to know to invoke it correctly and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description does not add parameter-specific semantics beyond what the schema already provides; it references 'supplied source facts' but leaves parameter details to the schema. This is adequate but not additive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Adversarially verify an answer is grounded in a specific set of facts.' It clearly distinguishes the tool from sibling tools like query_fact or validate_change by emphasizing adversarial verification against supplied source facts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use this tool: whenever an answer needs to be checked against a set of facts. It implies the use case without explicit exclusion or alternative reference, but the context is strong enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.15.5
    • Addedpin_annotation
  2. 2 tool updatesv0.13.0
    • Addedget_audit_log_head
    • Addedprove_entry_inclusion
  3. 1 tool updatev0.12.13
    • Addedverify_retrieval
  4. 1 tool updatev0.12.0
    • Changedquery_fact1 field changed
      • addedInput schema / properties / content_type
        Added value: +{
        +  "description": "Optional filter by content_type. Use 'procedure' to explicitly summon procedures (which are excluded from auto-injection by design).",
        +  "enum": [
        +    "rule",
        +    "fact",
        +    "procedure"
        +  ],
        +  "type": "string"
        +}
  5. 1 tool updatev0.7.4
    • Addedget_agents_md_constraints
  6. 26 tool updatesv0.7.3
    • First observedexport_claude_md
    • First observedfind_contradictions
    • First observedget_co_edit_suggestions
    • First observedget_compaction_audit
    • First observedget_constraints
    • First observedget_context_for_action
    • First observedget_decision_log
    • First observedget_health_report
    • First observedget_injection_context
    • First observedget_related_bugs
    • First observedingest_pr_reviews
    • First observedpredict_regression
    • First observedpredict_test_failures
    • First observedpromote_constraint
    • First observedquery_fact
    • First observedrecall_transcript_range
    • First observedrecord_compaction_audit
    • First observedrecord_correction
    • First observedrecord_decision
    • First observedrecord_event
    • First observedrecord_test_outcome
    • First observedresolve_contradiction
    • First observedsearch_global
    • First observedseed_project
    • First observedsimulate_change
    • First observedvalidate_change

TDQS

C2.9/5.0

Scored across 31 tools

Disambiguation2/5

Several tools have poorly separated boundaries: validate_change, simulate_change, predict_regression, and predict_test_failures all appear to assess the impact of a proposed change, differing mainly in subtle emphasis. Similarly, query_fact, search_global, and get_context_for_action overlap heavily in fact retrieval, and record_correction, record_decision, and pin_annotation all capture human feedback. An agent would frequently struggle to select the right tool without reading every description.

Naming Consistency5/5

Every tool follows a consistent snake_case verb_noun pattern, e.g., get_constraints, record_event, predict_regression, prove_entry_inclusion. Even the more unusual names like pin_annotation and seed_project fit the same imperative structure. This is a highly predictable and uniform naming convention.

Tool Count2/5

31 tools is well beyond the 25+ threshold for a single server and will overwhelm tool selection, especially given the many overlapping prediction and retrieval tools. The server would be more coherent with roughly half the current surface area, consolidating related reads and writes into broader commands.

Completeness4/5

The domain is broadly covered: it supports knowledge-graph population and queries, event and decision recording, constraint ingestion and validation, regression prediction, audit-log integrity, compaction auditing, and context export. Minor gaps exist, such as no explicit fact/constraint update or delete lifecycle and no project listing tool, but the core workflows are well supported.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    A temporal knowledge graph system that enables users to record and query architectural decisions, implementation patterns, and project failures. It integrates with Claude to provide hybrid search, timeline tracking, and automated knowledge gap detection using graph analysis.
    4
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides AI coding agents with pre-edit situational awareness by combining structural call graphs and co-change history to prevent incomplete edits. It surfaces files that historically change together, reducing missed coupled modules.
    3
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables coding agents to query a local, versioned knowledge graph of a software project, retrieving overviews, context packs, evidence, and explanations to make informed changes.
    MIT