world-model-mcp
World Model MCP
AI 코딩 에이전트를 위한 영구 메모리 + 선택적 서명 감사.
world-model-mcp는 AI 코딩 에이전트를 위한 영구 메모리와 포스트퀀텀 서명·오프라인 검증 가능한 감사 추적을 제공합니다 (FIPS 205 하이브리드 Ed25519 + SLH-DSA). MIT 라이선스이며 완전히 로컬에서 동작합니다.
world-model-mcp는 에이전트가 매 턴마다 조회하는 로컬 SQLite 지식 그래프를 제공합니다. 환각은 검증 가능해지고, 수정 사항은 세션을 넘어 유지되며, 회귀는 적용되기 전에 포착됩니다. 감사 체인을 켜면 모든 이벤트가 FIPS 205 하이브리드 Ed25519 + SLH-DSA로 서명되어 영구히 오프라인에서 검증할 수 있습니다. MIT 라이선스, 완전 로컬 실행, Claude Code, Cursor, Codex, Continue, Cline, Windsurf, GitHub Copilot Chat, pi, OpenClaw, Hermes Agent 등 10개 이상의 AI 코딩 에이전트와 호환됩니다.
최신 버전: v0.16.2.
world-model demo는 라이브 노터리 비트로 마무리됩니다. 세 개의 샘플 결정이 실제 하이브리드 서명 에포크(Ed25519 + SLH-DSA-SHA2-128f, FIPS 205)에 서명되고, VALID로 검증되며, 한 바이트를 변조하여 변조 탐지가 실시간으로 작동함을 증명한 뒤(INVALID) 복원됩니다. 가입 없음, 토큰 없음, 네트워크 없음. 현재 디렉터리에world-model-demo-receipt.json을 생성하고 복사-붙여넣기용etch.systems/verify#<manifest>URL을 출력합니다. v0.16.1에서는 로컬 liboqs 빌드가 SLH-DSA 없이 컴파일된 경우(데모가 충돌하는 대신) 정상적으로 건너뛰는 기능이 추가되었습니다. 전체 버전 기록은 CHANGELOG.md에서 확인하세요.
10초 만에 사용해 보기
pip install -U world-model-mcp && world-model demo변조 방지 에포크에 서명된 세 개의 샘플 결정을 확인한 다음, 데모가 한 바이트를 변조하고 재검증하여 변조가 즉시 탐지됨을 증명하고(INVALID), 복원 후 다시 검증합니다(VALID). 오프라인, 계정 없음, 네트워크 없음. 현재 디렉터리에 world-model-demo-receipt.json이 생성되고, 누구나 브라우저에서 동일한 영수증을 확인할 수 있는 공유 가능한 etch.systems/verify#… URL이 출력됩니다.
mcp-name: io.github.SaravananJaichandar/world-model-mcp
Related MCP server: Scrooge
관리형 호스팅 버전을 선호하시나요?
OSS 서버를 직접 실행하지 않고 동일한 서명 감사 체인을 원한다면, etch.systems는 이 패키지를 오프라인 참조 검증기로 사용하는 관리형 변형을 제공합니다.
실행할 인프라 없음(가입 없이 30초 만에 첫 이벤트 POST)
SR 11-7, EU AI Act 제12조, ISO 42001, NIST AI RMF, SOC 2 CC7 규정 준수 프레임워크 매핑
엔드투엔드로 제공되는 외부 앵커링(Sigstore Rekor + Bitcoin OpenTimestamps)
MCP 클라이언트 도구(Claude Code, Cursor, Continue, Cline, Codex) 전반에 걸친 키 순환 및 위임을 지원하는 휴대용 에이전트별 ID
오프라인 검증기는 동일합니다:
pip install world-model-mcp && etch-verify manifest.json은 호스팅 서비스가 온라인 상태일 필요 없이 고정된 공개 키로 호스팅 체인을 검증합니다.
curl -X POST https://etch.systems/v1/your-project베어러 토큰과 모든 MCP 호환 클라이언트가 연결할 수 있는 MCP 엔드포인트를 반환합니다. 전체 문서는 etch.systems에서 확인하세요.
자주 묻는 질문 (FAQ)
world-model-mcp란 무엇인가요?
world-model-mcp는 AI 코딩 에이전트를 위한 영구 메모리와 포스트퀀텀 서명·오프라인 검증 가능한 감사 추적을 제공하며, 에이전트가 매 턴마다 조회하는 MCP 서버로 노출됩니다. 완전히 로컬에서 실행되며, MIT 라이선스 Python + 10개 이상의 AI 코딩 에이전트용 선택적 어댑터로 제공됩니다.
Mem0, Letta 및 기타 에이전트 메모리 도구와 어떻게 다른가요?
world-model-mcp는 메모리와 함께 하이브리드 포스트퀀텀 서명 감사 체인(FIPS 205 Ed25519 + SLH-DSA-SHA2-128f)을 제공하고, 오프라인 참조 검증기(etch-verify)와 호스팅 etch.systems 컴패니언을 통해 제공되는 이중 외부 앵커링(Sigstore Rekor + Bitcoin OpenTimestamps)을 갖춘 유일한 에이전트 메모리 도구입니다. 유사 도구(Mem0, Letta, agentmemory)는 서명 감사 체인 없이 메모리만 제공합니다. 메커니즘별 세부 비교는 아래 비교 표를 참조하세요.
감사 추적은 포스트퀀텀 보안을 제공하나요?
네. 닫힌 모든 에포크는 하이브리드 서명을 포함합니다: Ed25519(클래식)와 SLH-DSA-SHA2-128f(FIPS 205 무상태 해시 기반 포스트퀀텀 서명)입니다. 미래의 양자 공격자가 Ed25519만 깨뜨린다 해도 동일한 페이로드에 대한 SLH-DSA 서명에 여전히 직면하게 되며, 체인을 위조하려면 두 서명을 모두 위조해야 합니다.
계정 없이 오프라인에서 레코드를 어떻게 검증하나요?
pip install -U world-model-mcp && world-model demo를 실행하세요. 데모는 세 개의 샘플 결정을 실제 하이브리드 서명 에포크에 서명하고, etch-verify CLI가 읽는 것과 동일한 매니페스트 형식을 내보낸 다음 VALID로 검증하고, 한 바이트를 변조하여 변조 탐지가 실제로 작동함을 증명한 뒤(INVALID) 복원하고 재검증합니다. 현재 디렉터리에 world-model-demo-receipt.json이 생성되고, 누구나 브라우저에서 영수증을 확인할 수 있는 etch.systems/verify#… URL이 출력됩니다. 가입 없음, 네트워크 없음.
어떤 AI 코딩 에이전트와 호환되나요?
Claude Code, Cursor, Codex, Continue, Cline, Windsurf, GitHub Copilot Chat, pi, OpenClaw, Hermes Agent와 호환됩니다. 각각 github.com/SaravananJaichandar/world-model-mcp-<agent>-starter 아래에 전용 스타터 저장소가 있습니다. MCP 와이어 형식이 표준이므로 MCP를 지원하는 모든 클라이언트가 동일한 도구를 사용할 수 있습니다.
이 OSS 패키지를 사용해야 할까요, 아니면 호스팅 etch.systems 서비스를 사용해야 할까요?
자체 인프라에서 감사 체인을 실행하거나 내부 도구에 서명 메모리를 추가하려면 이 OSS 패키지(MIT 라이선스, 완전 로컬)를 사용하세요. 규정 준수 프레임워크 매핑(SR 11-7, EU AI Act 제12조, ISO 42001, NIST AI RMF, SOC 2 CC7)과 MCP 클라이언트 도구 전반의 휴대용 에이전트별 ID를 갖춘 관리형 서비스로 동일한 감사 체인을 이용하려면 etch.systems을 사용하세요. 오프라인 참조 검증기는 두 경우 모두 동일합니다: pip install world-model-mcp && etch-verify manifest.json은 OSS 생성 또는 호스팅 생성 매니페스트를 고정된 공개 키로 검증합니다.
다른 에이전트 메모리 + 에이전트 감사 프로젝트와의 비교
에이전트 메모리 분야가 수렴하고 있는 감사/서명/앵커링 차원에서 8개 명시된 유사 프로젝트와의 1:1 비교입니다. 모든 world-model-mcp 셀에는 출처 플래그가 표시됩니다: own = 출시 제품에서 측정/관찰된 값, cited = 경쟁사의 공개 랜딩 페이지, 저장소 또는 보도 자료에서 가져온 값.
이 표의 기준 소스는 호스팅 서비스 저장소에 있습니다: world-model-mcp-hosted/src/etch/competitor_matrix.py. 둘 중 하나를 변경하면 두 곳 모두 업데이트하세요.
world-model-mcp (this repo) + Etch (hosted) | Mem0 | Letta | agentmemory | Unicity AOS | Repowise | Trinitite | Caura | FailproofAI | |
서명 체계 | Ed25519 + SLH-DSA-SHA2-128f(하이브리드) | 없음(저장 데이터 암호화만) | 없음 | 없음 | BLAKE3 해시 체인 | 결정적 신호(서명 없음) | 서명 + 해시 체인(알고리즘 미공개) | 없음 | 공개된 것 없음 |
포스트퀀텀 대비 | 예(SLH-DSA, NIST FIPS 205) | 아니요 | 아니요 | 아니요 | 아니요(BLAKE3 해시만) | 아니요 | 미공개 | 아니요 | 아니요 |
오프라인 참조 검증기 | etch-verify CLI, 스트리밍, PyPI 패키지에 포함 | 아니요 | 아니요 | 아니요 | 미공개 | 아니요(SaaS 전용) | 브라우저 기반 검증기 | 아니요 | 아니요 |
증명 주기 강제 | 예약된 systemd 타이머 + 온디맨드 검증 | 아니요 | 아니요 | 아니요 | 미공개 | 아니요 | 예약된 증명 실행 | 아니요 | 아니요 |
프레임워크 매핑(조문/통제 ID) | EU AI Act 제12-15조, SOC 2 CC6.6/7.2/7.3, ISO 27001 A.12.4/A.14.2 | SOC 2 + HIPAA(배지, 통제별 매핑 없음) | 미공개 | 아니요 | 미공개 | EU AI Act(네거티브 스페이스 주장) | 조문별 인용(EU AI Act, SOC 2, SR 11-7) | SOC 2 진행 중 배지 | SOC 2 엔터프라이즈 등급 전용 |
드리프트 감지 | 에포크별 체인 무결성 + 증명 추세 | 아니요 | 성공률 + 오류 추적 | 아니요 | 미공개 | 코드 건강 점수 델타 | 결정적 리플레이 + 위험 점수 델타 | 아니요 | 평가자 점수 델타 |
주석 지원(체인 내 서명된 사람 메모) | pin_annotation MCP 도구, 동일 체인에 서명 | 아니요 | 아니요 | 아니요 | 미공개 | 아니요 | 미공개 | 아니요 | 아니요 |
브라우저 검증기 | 예, /auditor/<slug>의 체인 무결성 위젯 | 아니요 | 아니요 | 아니요 | 미공개 | 아니요 | 예 | 아니요 | 아니요 |
공유 링크(감사자 접근, 로그인 불필요) | 예, 만료되는 공유 토큰(bs_ 접두사, 최대 30일) | 아니요 | 아니요 | 아니요 | 미공개 | 아니요 | 예, 만료되는 감사자 접근 | 아니요 | 아니요 |
외부 앵커(독립 증인 로그) | 이중: Sigstore Rekor + Bitcoin OpenTimestamps(공개, 프로젝트별 옵트아웃) | 아니요 | 아니요 | 아니요 | 내부 해시 체인만(외부 로그 없음) | 아니요 | 미공개 | 아니요 | 아니요 |
OSS 라이선스 | MIT(world-model-mcp) | Apache 2.0 | Apache 2.0 | Apache 2.0 | 미공개 | AGPL v3(코어) | OSS 없음 | Apache 2.0 | OSS 없음 |
GitHub 스타 수 | etch.systems/api/oss-stats를 통한 스냅샷 | 61.6k | 23.9k | 24.9k | 7.1k | 4.2k | 공개 저장소 없음 | 373 | 공개 저장소 없음 |
공개 자금 조달 | 부트스트랩 | $24.5M | $10M | 미공개 | 2026년 2월 $3M 시드 | 미공개 | 미공개 | 미공개 | 미공개 |
규정 준수 상태(랜딩 페이지 주장 기준) | SOC 2 Type I 진행 중(2026년 8월 목표) | SOC 2 + HIPAA(트러스트 서브도메인의 배지) | 미공개 | 규정 준수 상태 없음 | 규정 준수 상태 없음 | SOC 2 없음; EU AI Act 네거티브 스페이스 주장 | 조문별 규정 준수 프레임 | SOC 2 진행 중 배지 | SOC 2 엔터프라이즈 등급 |
이 표를 읽는 방법:
굵은 셀은 이 저장소(OSS) 또는 호스팅 서비스(etch.systems)에 포함된 메커니즘을 설명합니다.
경쟁사 셀은 각 경쟁사의 공개 랜딩 페이지/저장소/보도자료에서 인용한 것입니다. 당사 측정치가 아닙니다.
world-model-mcp셀의[own]은 해당 메커니즘을 당사가 직접 측정/관찰했음을 의미합니다.[cited]는 값이 제3자 출처에서 왔음을 의미합니다.이 표에는 감사/서명/앵커링 차원만 포함됩니다. 일반 메모리 기능(검색 정확도, 어댑터 범위, LLM 지원)은 아래 기능 섹션에서 다룹니다.
수치
Benchmark | Score | Details |
단일 시행 상한으로 +10.2포인트(49개 짝지어진 인스턴스에서 67.3% → 77.6%); 다중 시드 평균 효과는 인스턴스당 +0.24, 95% CI [0, 0.47] | 사전 등록, Claude Code 2.1.177 헤드리스, Zenodo DOI 10.5281/zenodo.21076824. 단일 시행 분할에서 도메인 내 +15.0포인트, 교차 도메인 +6.9포인트, 회귀 0건. | |
| 105쌍 × 19개 카테고리, 결정적(LLM 없음). v0.11.0부터 포함. | |
100.0% 정확히 일치 | 수동 라벨링된 12쌍(근거 기반 4, 부분 4, 환각 4). 독립 Coach LLM을 통한 레이어 3 적대적 검증. v0.12.12부터 포함. |
SWE-bench 수치가 핵심 경험적 주장입니다. 나머지 두 개는 출시된 구성 요소에 대한 내부 정확성 벤치마크입니다. 재현 스크립트는 각 벤치마크 디렉터리 또는 연결된 저장소에 있습니다.
테스트
출시된 코드베이스 전반에 걸친 단위 + 통합 + 퍼즈 테스트 1,494개. 커버리지 하한은 71% / 66%로 게이트됩니다(liboqs에 SLH-DSA가 있는 CI 러너와 없는 CI 러너를 위한 2계층).
# Run tests
pytest -q
# With coverage
pytest --cov=world_model_server --cov-report=term-missing
# Fuzz targets (Atheris, requires Clang / libFuzzer)
python fuzz/fuzz_verify_manifest.py fuzz/corpus/ -max_total_time=60
# Non-Atheris smoke fuzz (runs in every CI pass, no system deps)
pytest tests/test_fuzz_smoke_verify.py -q주목할 만한 주요 테스트 스위트:
FIPS 205 SLH-DSA 알려진 답변 테스트(
tests/test_fips_205_slh_dsa_kat.py): SLH-DSA-SHA2-128f의 매개변수 크기를 고정(공개 키 = 32바이트, 비밀 키 = 64바이트, 서명 = 17,088바이트), 잘못된 키 + 잘못된 메시지 + 변형된 서명 거부를 포함한 서명/검증 왕복, 그리고 영원히 true로 검증되어야 하는 고정 KAT 벡터 픽스처(tests/fixtures/slh_dsa_kat_vectors.json).스트리밍 검증기 바이트 동등성(
tests/test_etch_verify_streaming.py): 인메모리 익스포터와 스트리밍 익스포터 간의 바이트 단위 동일 출력을 고정하여, 감사자가 기록 아티팩트로 매니페스트를 해시할 때 운영자가 어떤 익스포터를 사용했는지와 관계없이 동일한manifest_sha256을 얻도록 합니다.모순 벤치마크(
benchmarks/contradictions-200/): 105쌍 × 19개 카테고리, 결정적. 모든 커밋에서auto전략을 100%로 고정합니다.
CI는 모든 푸시에서 커버리지 하한 + 전체 테스트 스위트를 게이트합니다. .github/workflows/pytest.yml을 참조하세요.
인증된 감사 체인(v0.13+, 옵트인)
감사 추적이 암호학적으로 검증 가능해야 하는 배포 환경(SOC 2, HIPAA, EU AI Act 또는 자체 내부 통제 목록)의 경우:
export WORLD_MODEL_AUDIT_LOG=on
# then start the world-model-mcp server as normal최초 옵트인 시작 시 서버는 기존 audit.db 파일에 두 개의 새 SQLite 테이블(tamper_evident_log 및 tamper_evident_epochs)을 생성하고, 첫 번째 에포크가 종료될 때 새 하이브리드 키페어를 생성합니다. 이후의 모든 이벤트는 SHA-256 Merkle 체인에 추가됩니다. 에포크가 종료되면(기본값: 1024개 이벤트) 체인 루트는 하이브리드 Ed25519 + SLH-DSA-SHA2-128f 엔벨로프로 서명됩니다. 클래식 + 포스트퀀텀 방식이므로, 타원곡선 암호화가 가상으로 깨지더라도 감사 추적은 그대로 유지됩니다.
오프라인에서 영구적으로 검증 가능합니다.
etch-verifyCLI는 PyPI 패키지와 함께 제공되며, 감사자는 노트북에서 실행하기만 하면 초기 다운로드 이후에는 네트워크 접근이 필요 없습니다.서명은 순서뿐만 아니라 작성자도 증명합니다. 해시 체인 대안은 재정렬이 없었음을 증명할 수 있지만, 누가 서명했는지는 증명할 수 없습니다. 이 체인은 둘 다 해결합니다.
대시보드도, 가입도, 외부 서비스도 없습니다. 로컬 SQLite를 대상으로 프로세스 안에서 완전히 실행됩니다.
전체 문서는 docs/AUDIT_LOG.md에서 확인하세요.
KMS 기반 키, 공개 투명성 로그, Sigstore Rekor + Bitcoin OpenTimestamps에 대한 외부 앵커링, 규정 준수용 운영자 대시보드를 추가한 호스팅 버전은 아래의 Etch 컴패니언을 참조하세요.
빠른 시작
가장 일반적인 세 가지 설치 방법입니다. 그 외 지원되는 모든 클라이언트(Cursor, Cline, Codex, Continue, Copilot, Windsurf, Goose, pi, OpenClaw, Hermes 등)는 etch.systems/docs/install을 참조하세요.
옵션 1: Claude Desktop(원클릭)
Releases에서 최신 .mcpb를 다운로드하여 Claude Desktop으로 끌어다 놓습니다. 후크, MCP 서버 구성, 종속성이 자동으로 설치됩니다.
옵션 2: Claude Code / IDE 플러그인(pip install)
# 1. Install the package
pip install world-model-mcp
# 2. Set up in your project (auto-seeds the knowledge graph from existing code)
cd /path/to/your/project
python -m world_model_server.cli setup
# 3. Restart Claude Code
# Done. The world model is pre-populated and active.다음 단계(선택 사항): 호스팅 Etch 노터리로 로컬 서명 감사 로그를 감사자가 검증할 수 있는 체인으로 전환하세요 — KMS 기반 키, 공개 투명성 로그, 외부 앵커링, 감사자용 공유 링크를 제공하며 실행할 인프라가 필요 없습니다. 호스팅 컴패니언: Etch를 참조하세요.
옵션 3: 원격 / MCP 터널 배포용 HTTP 전송
pip install 'world-model-mcp[http]'
python -m world_model_server.server --transport http --port 8000Streamable HTTP를 통해 MCP를 노출하여 원격 에이전트가 네트워크로 연결할 수 있게 합니다. 인증, CORS, 리버스 프록시 설정은 docs/http_transport.md를 참조하세요.
기타 클라이언트
OSS CLI를 통해 설치합니다. 각 명령은 해당 클라이언트에 맞는 올바른 구성을 작성합니다(인터프리터 경로는 기본적으로 sys.executable을 사용하며, 클라이언트별 형식을 인식하고, --force 및 --dry-run 플래그로 덮어쓰기로부터 안전합니다):
python -m world_model_server.cli install-cursor # Cursor
python -m world_model_server.cli install-cline # Cline
python -m world_model_server.cli install-codex # Codex
python -m world_model_server.cli install-continue # Continue (also --global)
python -m world_model_server.cli install-copilot # GitHub Copilot Chat (VS Code Insider)
python -m world_model_server.cli install-windsurf # Windsurf
python -m world_model_server.cli install-pi # pi
python -m world_model_server.cli install-openclaw # OpenClaw
python -m world_model_server.cli install-hermes # Hermes (MCP mode)
python -m world_model_server.cli install-hermes-provider # Hermes (Elixir-native provider)클라이언트별 전체 가이드(검증 명령 + 문제 해결 포함)는 etch.systems/docs/install에서 확인하세요.
하는 일
world-model-mcp는 AI 코딩 에이전트와 그 작업 사이에 위치하는 시간적 지식 그래프입니다. 코드베이스에서 사실, 엔티티, 제약 조건을 기록하고, 편집 경계에서 학습된 제약 조건에 대해 모든 코드 변경을 검증하며, 컨텍스트 창 압축 후 관련 컨텍스트를 다시 주입하고, 신뢰도 가중치 기반 해결로 모순을 추적하며, 독립적인 Coach LLM을 통해 검색 결과를 적대적으로 검증합니다.
기능
1. 환각 방지. 기록된 모든 사실에는 출처(asserted_by, confirmer, confirmation_state, evidence_type)가 포함됩니다. 에이전트가 사실을 조회하면 답변과 함께 신뢰도 점수를 받습니다. 두 사실이 모순될 때는 더 새롭거나 신뢰도가 높은 쪽이 승리하며, 패배한 쪽은 superseded_by와 함께 유지됩니다. 에이전트는 저장된 사실이 여전히 유효한지 추측할 필요가 없습니다.
2. 수정에서 학습. 에이전트를 수정하면(record_correction) 해당 수정 사항이 관련 엔티티와 함께 일급 이벤트로 저장됩니다. 다음에 에이전트가 해당 엔티티와 관련된 무엇이든 조회하면 수정 사항이 먼저 표시됩니다. 수정 사항은 세션을 넘어, 컨텍스트 압축을 넘어, 에이전트 재시작을 넘어 유지됩니다.
3. 회귀 방지. 모든 코드 변경 제안은 validate_change를 거치며, 이는 제약 조건 그래프를 탐색하여 편집이 적용되기 전에 위반 사항을 반환합니다. 제약 조건은 코드베이스에서 자동으로 학습되고(seed_project), PR 리뷰 댓글로 강화되며(ingest_pr_reviews), 수동으로 작성됩니다(record_event). 에이전트는 위반 사항, 제안, 제약 조건의 출처를 확인할 수 있습니다.
4. Coach-Player 적대적 검증. Player 에이전트가 결정을 초안으로 작성합니다. 독립적인 Coach 에이전트는 선례를 찾기 위해 그래프를 독립적으로 다시 조회하고, Merkle 증명을 따라가며, 하이브리드 서명을 확인한 후에만 승인합니다. Coach는 Player의 요약을 절대 신뢰하지 않으며, 매번 서명된 원장에 대해 재검증합니다. 12개의 수동 레이블링된 쌍에서 100% 정확히 일치합니다.
전체 기술 아키텍처는 docs/ARCHITECTURE.md에서 확인하세요.
MCP 도구
8개의 도구가 제공됩니다. 한 줄 요약이며, 각 도구를 클릭하면 docs/mcp/에서 시그니처와 예제를 확인할 수 있습니다.
도구 | 기능 |
| 출처 + 신뢰도 + 소스와 함께 저장된 사실을 가져옵니다 |
| 감사 체인에 이벤트를 추가합니다(고정 enum의 |
| 학습된 제약 조건에 대해 코드 변경 제안을 검사하고 위반 사항 + 제안을 반환합니다 |
| 엔티티 또는 파일 패턴과 일치하는 모든 제약 조건을 나열합니다 |
| 사용자 수정 사항을 저장하여 다음 관련 조회 시 다시 표시되게 합니다 |
| 지정된 파일과 관련된 버그 + 파일별 위험 점수를 가져옵니다 |
| 기존 코드베이스를 지식 그래프에 대량으로 수집합니다 |
| 최근 PR 리뷰 댓글을 제약 조건으로 변환합니다 |
전체 도구 문서는 docs/mcp/에서 확인하세요.
호스팅 컴패니언: Etch
world-model-mcp를 로컬에서 실행 중이신가요? Etch (etch.systems)는 동일한 OSS 코어 위에 구축된 호스팅 거버넌스 계층으로, 규정 준수 팀이 프로덕션 사용을 승인하는 데 필요한 기능을 추가합니다:
KMS 암호화 서명 키(저장 시 평문이 절대 아님)
서명된 헤드가 있는 공개 투명성 로그(분할 뷰 저항)
Sigstore Rekor + Bitcoin OpenTimestamps에 대한 외부 앵커링(이중 독립 증인)
세션 추적, 체인 무결성 뷰, PII 스캔, 클라이언트 응답 PDF 내보내기가 포함된 운영자 대시보드
/auditor/<slug>의 독립형 감사자 CLI + 브라우저 검증기무료 티어, 그 이상은 사용량에 따라 과금
동일한 암호화 프리미티브, 동일한 감사 로그 스키마, 제공되는 OSS에 대한 소스 변경이 전혀 없습니다. 독립 실행형으로 운영 중이라면 건너뛰어도 됩니다.
어떤 것이 나에게 적합할까요?
상황 | 방법 |
로컬 AI 코딩 에이전트를 위한 영구 메모리를 원한다 |
|
6개월 후 규제 기관에 에이전트 결정을 증명하고 싶다 | |
법원이 자체 인증할 수 있는 서명된 증거 번들을 원한다 | |
서명된 감사 기록의 팀 간 또는 벤더 간 연합을 원한다 | |
둘 다 시도해 보고 싶다 |
|
작동 방식
world-model-mcp는 SQLite로 뒷받침되는 시간적 지식 그래프를 노출하는 MCP 서버(stdio 또는 Streamable HTTP)입니다. 모든 에이전트 턴은 그래프를 조회하고, 모든 코드 변경 또는 수정은 다시 기록합니다. 시작 시 서버는 기존 코드베이스에서 자동으로 시드하며, 종료 시 깨끗하게 플러시합니다.
.claude/world-model/ 아래의 6개 데이터베이스:
entities.db: 파일, 함수, 클래스, 심볼facts.db: 출처 + 신뢰도가 있는 의미적 사실relationships.db: 종속성, 호출, 임포트constraints.db: 에이전트가 준수해야 하는 학습된 규칙sessions.db: 세션별 컨텍스트 추적events.db: 불변 이벤트 로그(옵트인 시 감사 체인)
다이어그램이 포함된 전체 아키텍처는 docs/ARCHITECTURE.md에서 확인하세요.
구성
환경 변수(모두 선택 사항):
변수 | 용도 | 기본값 |
| 서명된 감사 체인 활성화 |
|
| SQLite 데이터베이스 디렉터리 재정의 |
|
| 익명 OSS 텔레메트리 옵트인(개인정보 보호 참조) |
|
| 선택적 LLM 기반 기능을 활성화한 경우에만 사용 | 없음 |
전체 구성 참조는 docs/CONFIGURATION.md에서 확인하세요.
개인정보 보호 및 보안
텔레메트리는 기본적으로 꺼져 있습니다.
WORLD_MODEL_TELEMETRY=on으로 옵트인하면 집계된 설치 수준 메트릭이etch.systems/api/telemetry/ingest로 전송됩니다(엔드포인트 URL, 소스 코드 없음, 프롬프트 없음, PII 없음). 삭제 권리는DELETE /api/telemetry/install/{install_id}를 통해 지원됩니다.핵심 운영에는
ANTHROPIC_API_KEY가 필요하지 않습니다. 일부 선택적 기능(Coach-Player 레이어 3 검증, LLM 기반 재순위화)은 키가 제공된 경우 API를 사용합니다. 키가 없어도 다른 모든 기능은 작동합니다.보안 문제 신고: security@etch.systems로 이메일을 보내주세요. PGP 키는 요청 시 제공됩니다.
전체 문서는 docs/PRIVACY.md에서 확인하세요.
기여
기여는 언제나 환영합니다. 개발 환경 설정, 코딩 표준, 언어 지원 추가, 테스트 작성, PR 제출은 CONTRIBUTING.md를 참조하세요.
특히 도움이 필요한 분야:
언어 파서(Go, Rust, Java, C++)
추가 MCP 클라이언트 어댑터
프레임워크 통합(LangGraph, CrewAI, AutoGen, LlamaIndex; 호스팅 저장소에 스타터 shim이 있습니다)
benchmarks/의 벤치마크 기여
첫 PR 전에 CLA.md를 읽어주세요.
라이선스
MIT License. 상업적 및 개인적 사용이 무료입니다.
링크
전체 버전 기록: CHANGELOG.md
문서: docs/
호스팅 서비스: etch.systems
Zenodo(공식 인용): DOI 10.5281/zenodo.20834508
Available Tools
31 toolsexport_claude_mdB
Generate a CLAUDE.md document from the knowledge graph (top constraints, recent decisions, known bug regions, co-edit patterns).
| Name | Required | Description | Default |
|---|---|---|---|
| max_constraints | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavior. It states it generates a document but does not clarify whether it writes to a file, returns content, or has side effects. It also fails to mention any permissions, limits, or output format, leaving key behavioral aspects ambiguous.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with a clear verb and object, followed by a parenthetical list of content categories. It is concise and contains no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description should explain return values, side effects, and parameter behavior. It only covers purpose and content categories, leaving critical operational details absent for an export tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage for the only parameter, max_constraints. The description broadly references 'top constraints' but does not explain that max_constraints limits that section, nor its units or effect on other sections. The parameter semantics are only indirectly inferable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Generate') and identifies the resource ('CLAUDE.md document from the knowledge graph'), and it lists the included content categories (top constraints, recent decisions, known bug regions, co-edit patterns), distinguishing it from sibling tools that query individual aspects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for producing an aggregated CLAUDE.md from knowledge graph data, but it does not explicitly state when to use this tool versus querying individual components via siblings like get_constraints or get_decision_log. It lacks explicit exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_contradictionsC
Find pairs of facts that contradict each other based on similarity and status differences
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosing behavior. It doesn't state whether the tool is read-only, what the output looks like, or any side effects. The hint about 'similarity and status differences' is the only behavioral clue, but it's insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, which is concise and front-loaded, but it sacrifices necessary detail. It's not bloated, but it's under-specified. The brevity doesn't earn its place because it lacks critical information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has two parameters, no output schema, and no annotations. The description provides no information about return values, parameter usage, or behavioral context. It is far from complete even for a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has two parameters (limit, query) with zero descriptions, and the description doesn't mention them at all. Since schema coverage is 0%, the description fails to compensate, leaving the agent with no understanding of what these parameters do.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool finds pairs of contradicting facts, using a specific verb ('Find') and resource. It adds a hint of the method ('similarity and status differences'), but doesn't explicitly differentiate from sibling tools like resolve_contradiction, though the distinction is evident from the name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives. It doesn't mention any exclusions or prerequisites, leaving the agent to infer usage solely from the vague description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_agents_md_constraintsA
Parse AGENTS.md / CLAUDE.md / GEMINI.md / .agents/skills/*.md in the project and return declarative constraints. Mixed into PreToolUse enforcement automatically; this tool exposes the same data for inspection.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | No | ||
| project_dir | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden for behavioral disclosure. It explains the parsing action and its role in enforcement, but does not mention whether it performs a read-only operation, how missing files are handled, or any potential side effects. The context about automatic enforcement adds value but lacks depth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the main action and scoped to specific files. Every phrase contributes meaning, and the inspection/enforcement contrast adds valuable context without bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's purpose and its relationship to PreToolUse enforcement, giving useful context. However, it omits parameter semantics and the exact return format (beyond 'declarative constraints'), which is especially important since there is no output schema. The lack of detail on expected inputs and outputs leaves gaps for an agent trying to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has two parameters (file_path and project_dir) with no descriptions and 0% coverage. The description does not mention either parameter, leaving their specific purpose and format ambiguous. For example, it is unclear if file_path is optional or relative to project_dir. The description fails to compensate for the complete lack of schema guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it parses AGENTS.md / CLAUDE.md / GEMINI.md / .agents/skills/*.md and returns declarative constraints. This specific verb+resource distinguishes it from the sibling tool get_constraints, which likely covers broader constraint sources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description notes that the data is 'Mixed into PreToolUse enforcement automatically' and that this tool 'exposes the same data for inspection,' implying it is intended for inspection/debugging rather than direct enforcement. It provides clear context but does not explicitly name alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_audit_log_headA
v0.13 tamper-evident audit log. Return the current head state (last log entry seq, last closed epoch seq, unclosed-entry count) plus the full closed-epoch chain with hybrid signature envelopes. Compliance auditors call this periodically to verify no operator misbehavior has occurred since the last check. Requires WORLD_MODEL_AUDIT_LOG=on at server startup.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the tamper-evident nature, the return content, and a runtime prerequisite (WORLD_MODEL_AUDIT_LOG=on). It adds meaningful operational context beyond a simple 'get' but stops short of stating read-only semantics explicitly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only two sentences but packs in the purpose, return value details, use case, and a critical prerequisite. Every clause earns its place with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter, read-only operation with no annotations or output schema, the description covers all essential aspects: what it does, what it returns, when to use it, and what server configuration is required. An agent has enough to deploy it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the schema is empty. The description explains what the tool returns, which is the only meaningful semantic content in this case. Baseline 4 applies since there are no parameters to describe.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Return' and names the resource 'audit log head state' plus the 'closed-epoch chain', clearly defining what the tool does. This differentiates it from other audit-related siblings like get_compaction_audit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states a use case: 'Compliance auditors call this periodically to verify no operator misbehavior has occurred since the last check.' This gives clear context on when to use it, but it does not mention alternatives or exclusions, so it misses the top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_co_edit_suggestionsB
Get files commonly edited alongside the given file based on historical patterns
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| file_path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden for behavioral disclosure. It mentions 'based on historical patterns' which adds some context, but it does not describe whether the operation is read-only, how suggestions are ranked, what format the results take, or any potential side effects. This is insufficient for full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that is front-loaded with the action and resource. Every word earns its place, and there is no superfluous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with two parameters, but the description leaves major gaps: no usage guidance, no parameter details, and no mention of return values (no output schema). For an agent to use this correctly, it needs more context about how suggestions are generated and what to expect from the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero description coverage (0%) for its parameters, and the description does not compensate. It references 'the given file' (mapping to file_path) but does not clarify the expected format, the meaning or usage of 'limit', or any constraints. The description adds minimal value beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Get' and the resource 'files commonly edited alongside the given file', with the basis 'historical patterns'. It is specific and distinguishes itself from the sibling tools, none of which share a similar focus on co-edit suggestions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention any exclusions, prerequisites, or conditions, leaving the agent without context for selecting it among the many sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_compaction_auditA
List recent compaction audit entries, most-recent first. Filter by session_id or limit count.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| session_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the transparency burden. It discloses the ordering (most-recent first) and the available filters (session_id, limit), which is useful. However, it does not mention whether the operation is read-only, whether limit has a default, or describe the return format or pagination behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that delivers all key information without redundancy. Every clause adds functional value, and there is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list/filter tool, the description covers the essential purpose and options. However, with no output schema and no annotations, the agent is left without knowledge of the returned fields, default limit behavior, or error cases. This is adequate for simple use but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must add meaning. It does clarify that session_id filters and limit controls the count, which is helpful. However, the semantics are shallow: it does not specify whether limit is mandatory, its maximum/default value, or whether session_id requires exact or partial matching.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists compaction audit entries in reverse chronological order. It names the specific resource (compaction audit entries) and the action (list), which differentiates it from write-oriented siblings like record_compaction_audit. However, it does not explicitly distinguish from the similar get_audit_log_head tool, preventing a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for inspecting compaction audit history, and the mention of filtering suggests relevant scenarios. However, it provides no explicit guidance on when to prefer this over alternatives like get_audit_log_head, and there are no exclusions or prerequisites stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_constraintsA
Get constraints (linting rules, patterns, conventions) for a file
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| constraint_types | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It implies a read-only operation ('Get') but does not disclose potential behaviors such as error handling for missing files, whether it searches the entire repository, or if any filtering is applied beyond the optional constraint_types parameter. The description adds some clarity by defining constraints, but does not reveal side effects or edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that efficiently communicates the tool's purpose without extraneous words. It earns a high score for being concise and structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (read operation, 2 parameters, no output schema), the description is mostly complete. It states what the tool does and the schema covers parameter details. However, it does not specify the return format or behavior when no constraints are found, which could leave some ambiguity for the agent. Still, for a straightforward getter, it is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, so the description must compensate. It does so by explaining that constraints include 'linting rules, patterns, conventions', which helps interpret the constraint_types enum. However, it does not explicitly map parameters or explain the file_path semantics beyond the name. The schema itself provides the enum values, offering adequate baseline coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves constraints for a file, using the verb 'Get' and specifying the resource and scope. It also clarifies what constraints are (linting rules, patterns, conventions). However, it does not explicitly distinguish from the sibling tool 'get_agents_md_constraints', which may overlap in purpose for specific files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when constraints for a file are needed, but provides no explicit guidance on when to use this tool versus alternatives like 'get_agents_md_constraints' or when not to use it. It lacks exclusions and alternative recommendations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_context_for_actionC
Pre-action context bundle: constraints, decisions, bugs, co-edits, related facts, and risk score for a file before editing
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| action_type | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It lists the content components but does not state whether the tool is read-only, how errors are handled, what the risk score means, or how the bundle is returned. This is a significant gap for a retrieval tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a concise noun phrase that front-loads the core purpose and lists contents, but it lacks a verb and reads more like a label than a full sentence. It is not overly verbose, yet it could be more structured with a clearer main clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two parameters, no annotations, and no output schema, the description is too sparse. It does not explain the output format, the meaning of the risk score, or how this bundle relates to the individual context tools. The description leaves significant gaps in understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage. The description mentions 'file' which maps to file_path, but action_type is only implied by 'editing' and not explicitly explained. The enum values (edit, create, delete, refactor) are not described, so the description adds minimal meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a pre-action context bundle for a file, listing specific content types such as constraints, decisions, bugs, co-edits, related facts, and risk score. This distinguishes it from sibling tools that target single context types. However, the phrase 'before editing' is slightly inconsistent with the action_type enum which includes create, delete, and refactor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide explicit guidance on when to use this bundle versus the many sibling tools like get_constraints or get_related_bugs. It implies usage before an action, but there are no exclusions or alternative recommendations, leaving the agent to infer the appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_decision_logC
Get decision traces showing agent proposals and human corrections
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| file_path | No | ||
| session_id | No | ||
| decision_type | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It implies a read operation via 'Get', but does not disclose behavior such as default limits, ordering, filtering effects, or whether it is purely read-only. The description focuses on content, not operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no wasted words. It efficiently states the purpose without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has four parameters, no output schema, and no annotations, yet the description provides no details on return value, parameter usage, or edge cases. It is overly minimal for the tool's apparent complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description does not explain any parameters. It only hints at decision_type through 'corrections', but limit, file_path, and session_id are completely unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves 'decision traces' with specific content (agent proposals and human corrections), using a specific verb 'Get' that distinguishes it from sibling write tools like record_decision and record_correction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as get_audit_log_head or get_compaction_audit. There is no mention of scenarios, exclusions, or preferred contexts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_health_reportA
Memory health diagnostics: orphans, stale facts, contradictions, decay candidates, DB sizes
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. The term 'diagnostics' implies a read-only operation, and the listed categories provide concrete insight into what the tool examines. However, it does not explicitly state whether any state is modified or note any performance implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise phrase that effectively communicates purpose and scope without wasted words. It is front-loaded and every element adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter diagnostic tool with no output schema, the description provides sufficient context about the report's contents. It could be improved by noting the output format, but the current level is adequate for understanding the tool's purpose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description need not document inputs. The baseline of 4 is appropriate because no parameter information is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function as memory health diagnostics and enumerates specific areas it covers (orphans, stale facts, contradictions, decay candidates, DB sizes). This distinguishes it from narrower sibling tools like find_contradictions or get_compaction_audit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for overall health assessment, but it does not explicitly state when to prefer this tool over siblings or when not to use it. No alternative tools are mentioned, so guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_injection_contextB
Return a compact constraint+fact bundle for PostCompact / UserPromptSubmit hooks to re-inject after context loss.
| Name | Required | Description | Default |
|---|---|---|---|
| max_facts | No | ||
| event_type | Yes | ||
| project_hint | No | ||
| max_constraints | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden. It only says 'Return...' without disclosing whether this is a read-only operation, any side effects, how 'compact' is achieved (e.g., truncation, filtering), or what happens if parameters like max_facts are omitted. This lack of behavioral detail is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It is concise and easy to read, though it sacrifices substance for brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 4 parameters, zero schema coverage, no annotations, and no output schema, the description is insufficiently complete. It gives a high-level purpose but lacks details on parameter behavior, return format, and how this tool fits with alternatives. More context is needed for an agent to use it reliably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it does not. It mentions two of the three event_type values but does not explain max_facts, max_constraints, project_hint, or the SessionStart event. The term 'compact' weakly implies size limits, but no explicit parameter meaning is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns a 'compact constraint+fact bundle' for specific hooks (PostCompact/UserPromptSubmit), which distinguishes it from sibling tools like query_fact or get_constraints. However, it could be more explicit about what 'context loss' entails or how it differs from get_context_for_action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names specific trigger events (PostCompact/UserPromptSubmit) and the goal of re-injecting after context loss, giving clear context for when to use the tool. It does not explicitly mention alternatives or when not to use it, but the event-specific framing provides adequate guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ingest_pr_reviewsA
Pull GitHub PR review comments and convert them into learned constraints in the knowledge graph
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | GitHub repo (owner/repo). Auto-detected from git remote if omitted. | |
| count | No | Number of recent PRs to scan (default 10, max 50) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses that the tool converts PR comments into learned constraints, indicating a write operation to the knowledge graph. However, it does not mention idempotency, overwrite behavior, permissions, or any side effects beyond the conversion, which is minimal for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that directly states the action and outcome with no filler. Every word contributes to the meaning, making it highly efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has only two optional parameters and no output schema, so minimal description might suffice. However, because it is an ingest operation affecting the knowledge graph, a bit more context about result expectations or side effects would improve completeness. It is adequate but not fully comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds no extra parameter-specific meaning beyond what the schema already provides (repo auto-detection, count default/max). It does not compensate for any missing details, but none are missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs ('Pull' and 'convert') and identifies the resource ('GitHub PR review comments') and target ('knowledge graph'). It clearly distinguishes itself from sibling tools like get_constraints or record_event by describing a unique ingest-and-transform workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when you want to bring PR review comments into the knowledge graph, but it does not explicitly state when to prefer this over alternatives or provide exclusions. Sibling tools like record_correction or validate_change serve different purposes, yet no direct comparison is offered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pin_annotationA
Attach a signed human annotation (note, override rationale, or intervention record) to a span of agent events. Persists into the annotations table and chains into the same Merkle audit log as agent writes (v0.15.0, ADR-0001). Rationale limited to 8 KB.
| Name | Required | Description | Default |
|---|---|---|---|
| author | Yes | Author identity. Self-asserted in OSS; KMS-verified in Etch hosted. | |
| rationale | Yes | Human rationale text (UTF-8). Max 8192 bytes. | |
| session_id | Yes | Session containing the annotated events. | |
| annotation_type | Yes | ||
| event_range_end | Yes | Last event_id in the annotated span. Equals event_range_start for a single-event annotation. | |
| event_range_start | Yes | First event_id in the annotated span. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and exceeds expectations by disclosing persistence ('Persists into the annotations table'), audit integration ('chains into the same Merkle audit log as agent writes'), version/ADR references, and the rationale size limit. This gives the agent a clear behavioral model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action, and includes only essential extra context (persistence, audit log, limit). Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 6 required parameters, no output schema, and no annotations, the description provides solid context for a write operation: it explains persistence and audit chaining. It lacks explicit error handling or return behavior, but that is not critical for a basic mutation tool. The description is largely complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 83% coverage (5 of 6 parameters have descriptions), so the baseline is 3. The description adds little beyond the schema; it mentions the 8 KB rationale limit (already in schema) and does not elaborate on parameter meanings or relationships. The schema itself is well-documented, so a baseline score is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action ('Attach') and object ('a signed human annotation to a span of agent events'). It distinguishes from siblings like record_event by emphasizing human annotation vs agent events and mentions specific annotation types.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case of attaching human notes or overrides to event spans, but provides no explicit guidance on when to choose this tool over alternatives such as record_correction or record_decision. The context is clear but lacks exclusions or alternative recommendations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_regressionA
Score regression risk for a proposed change to a file based on past bugs, test failures, and constraint violations
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| change_description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses what the tool uses (past bugs, test failures, constraint violations) but does not state whether it has side effects, permissions requirements, or what the output looks like. The methodology hint adds some behavioral context, but safety and operational traits remain undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It efficiently conveys the tool's purpose and key inputs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two simple parameters and no output schema, the description is adequate for tool selection but does not fully prepare the agent for invocation. It lacks details on input semantics and return value format, though the simplicity and sibling context make it acceptable but not thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions 'proposed change to a file' which loosely maps to file_path and change_description, but it does not clarify parameter formats, required fields, or examples. The description adds minimal meaning beyond what the property names already imply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Score') and resource ('regression risk') with the basis ('past bugs, test failures, and constraint violations'). It clearly distinguishes from siblings like predict_test_failures (which targets test failures specifically) and simulate_change (which simulates changes).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for assessing regression risk of a proposed change to a file, but it does not explicitly state when to prefer this tool over alternatives or provide exclusions. Sibling tools like predict_test_failures or validate_change could overlap, and no guidance is given on choosing among them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_test_failuresA
Surface tests likely to fail given a set of edited files
| Name | Required | Description | Default |
|---|---|---|---|
| file_paths | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the full burden for behavioral transparency. It only states the core action without disclosing whether the operation is read-only, what data it relies on, what output format to expect, or any limitations. The description is minimal and leaves the agent without important contextual cues.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the action and avoids redundancy. Every word contributes to the core purpose, making it highly concise and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is relatively simple (one parameter, no output schema, no annotations), but the description is thin. It communicates the essential purpose, but does not explain what the tool returns (e.g., a list of test names) or provide operational boundaries. It is minimally complete for selection but not fully adequate for invocation without further inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has no descriptions for the parameter file_paths (0% coverage). The description adds semantic value by indicating these are 'edited files', clarifying the parameter's intent. However, it does not specify path format, file existence requirements, or other constraints that would be useful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Surface' with a clear resource ('tests likely to fail') and a scope condition ('given a set of edited files'). This distinguishes it from sibling tools like predict_regression and get_related_bugs, which address different aspects of change impact.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'given a set of edited files' implies the primary use case, but the description does not explicitly state when to use this tool over alternatives or provide any exclusions. It lacks guidance on how to compare with similar tools like predict_regression.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
promote_constraintB
Promote a constraint from this project to all other registered projects
| Name | Required | Description | Default |
|---|---|---|---|
| constraint_id | Yes | ||
| target_projects | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals the action 'promote' but does not disclose side effects, permissions required, reversibility, or the fact that it likely mutates multiple projects. This leaves significant behavioral uncertainty.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that front-loads the verb and object. It contains no filler or unnecessary detail, making it easy to parse and understand.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema and no annotations, and the description is minimal. It lacks essential context about the promotion's effects, error conditions, prerequisites, and return values. For a cross-project mutation tool, this is insufficient for an agent to use safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It implies constraint_id identifies the constraint, but it does not explain target_projects, its optionality, or how it interacts with the default 'all other registered projects'. The description adds minimal meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'promote' with a clear resource, 'constraint', and a clear scope, 'from this project to all other registered projects'. This distinguishes it from sibling tools like get_constraints (retrieval) and validate_change (validation), making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No usage guidance is provided. The description does not mention when to use this tool versus alternatives, nor any prerequisites or conditions for promotion. It only states the action itself, leaving the agent without context for decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prove_entry_inclusionA
v0.13 tamper-evident audit log. Return a cryptographic inclusion-proof bundle for a persisted row_id (fact, constraint, event, or decision ID). Bundle includes the entry, the containing signed epoch (Ed25519 + SLH-DSA hybrid signature envelope), an RFC 6962 Merkle inclusion proof, and the full epoch chain from genesis. Requires WORLD_MODEL_AUDIT_LOG=on at server startup; returns an error object when opt-in is off, when the row_id is not found, or when the entry is in the unclosed backlog.
| Name | Required | Description | Default |
|---|---|---|---|
| row_id | Yes | ID of the fact / constraint / event / decision to prove inclusion for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses the exact bundle components (entry, signed epoch, Merkle proof, epoch chain), required server flag, and all error cases, giving a thorough behavioral profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single information-dense sentence that front-loads the primary action ('Return a cryptographic inclusion-proof bundle') and then enumerates bundle contents and failure modes without unnecessary filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by listing the bundle components explicitly. It also covers prerequisites and error scenarios, making it complete enough for a single-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (row_id with description), but the description adds semantic value by specifying that row_id refers to fact, constraint, event, or decision ID, clarifying the expected input beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns a cryptographic inclusion-proof bundle for a persisted row_id. It names the specific resource (audit log entries) and distinguishes it from siblings like get_audit_log_head or query_fact by focusing on inclusion proofs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use it (to prove inclusion of an entry), prerequisites (WORLD_MODEL_AUDIT_LOG=on), and error conditions (opt-in off, row_id not found, unclosed backlog). It does not explicitly mention alternatives, but the scope is clear enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_factB
Query the knowledge graph for facts about entities (APIs, functions, classes, etc.)
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The search query (e.g., 'User.findByEmail', 'JWT authentication') | |
| context | No | Additional context for the query | |
| entity_type | No | Optional filter by entity type | |
| content_type | No | Optional filter by content_type. Use 'procedure' to explicitly summon procedures (which are excluded from auto-injection by design). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full responsibility for behavioral disclosure. It doesn't state whether the tool is read-only, what the response format is, or that it can retrieve 'rules' and 'procedures' despite being named 'facts'. The term 'facts' may mislead users into thinking only fact-type content is returned, leaving important behavior undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that is front-loaded with the action and resource. Every word contributes to the core purpose without fluff or repetition. It is appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a general-purpose query tool with no output schema and no annotations, the description is minimal but adequate. It does not explain return values, pagination, or how to decide between this and related sibling tools. The ambiguous scope of 'facts' (vs rules/procedures) is a notable gap, but the schema partially covers this via content_type descriptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema coverage is 100% with descriptive text for all parameters, so the baseline is 3. The tool description adds no parameter-level meaning beyond what the schema already provides; it only reiterates the general entity focus. The description does not compensate or enhance the schema's parameter docs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a clear verb ('Query') and resource ('knowledge graph') and specifies the object ('facts about entities'). It lists example entity types (APIs, functions, classes) which conveys scope. However, it doesn't distinguish itself from sibling tools like search_global that might also search the knowledge graph, so it lacks explicit differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance in the description about when to use this tool versus alternatives. The only usage hint ('Use "procedure" to explicitly summon procedures...') appears in the schema's content_type parameter, not the tool description, and it is parameter-level rather than tool-level. No when-not-to-use or alternative tool references are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recall_transcript_rangeC
Hydrate a Claude Code session transcript by line range. Lets agents trace a fact back to the exact conversation that produced it.
| Name | Required | Description | Default |
|---|---|---|---|
| line_end | No | ||
| line_start | No | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only says 'Hydrate' without explaining whether the operation is read-only, what happens with invalid ranges, whether there are performance or memory implications, or what the response contains. This is a significant gap for a data-retrieval tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences and front-loaded with the core action. Each sentence contributes meaning: the first states the operational scope, the second provides the motivating use case. It is well-structured and free of fluff, though slightly vague.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and a 0% parameter coverage, the description should provide more contextual information about return values, error handling, and when to choose this tool. The current description is insufficient for an agent to confidently invoke the tool correctly in all cases, even though the tool itself is relatively simple.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions 'by line range,' which somewhat clarifies line_start and line_end, but it does not explain session_id, the inclusive/exclusive nature of the range, defaults, or behavior when only one line parameter is provided. The description adds minimal semantic value beyond the schema fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool hydrates a Claude Code session transcript by line range, with a specific use case of tracing facts to their source conversation. This distinguishes it from sibling tools like query_fact, which focus on facts rather than raw transcript retrieval. However, the verb 'hydrate' is somewhat jargon-heavy, and it lacks an explicit contrast with related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool ('Lets agents trace a fact back to the exact conversation that produced it'), which provides some context. However, it does not explicitly state when to prefer this over alternatives like query_fact or get_audit_log_head, nor does it give exclusions or prerequisites. The guidance is present but inferred rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_compaction_auditA
Record a context-compaction event with token counts and what was re-injected. Lets developers audit what was remembered across compaction boundaries.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | No | ||
| raw_summary | No | ||
| facts_injected | No | ||
| injection_event | No | ||
| pre_compact_tokens | No | ||
| post_compact_tokens | No | ||
| constraints_injected | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description must disclose behavioral traits. It states it records an event, implying a write operation, but does not mention side effects, whether data is appended or overwritten, permissions, or failure modes. This is a significant gap for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is exactly two sentences, front-loaded with the core action and followed by the purpose. Every word adds value with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 7 optional parameters, no output schema, and no annotations, the description does not fully equip an agent to use the tool. It lacks information about expected return values, whether any parameters are required in practice, and how this differs from other record_* tools beyond the compaction focus. This is insufficient for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It mentions 'token counts' and 'what was re-injected,' which roughly maps to pre_compact_tokens, post_compact_tokens, facts_injected, and constraints_injected, but it does not explain parameters like injection_event, session_id, or raw_summary. This leaves meaningful ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Record' with the resource 'context-compaction event' and specifies token counts and re-injected content. It clearly distinguishes from siblings like get_compaction_audit (retrieval) and record_event (generic event) by focusing on compaction-specific auditing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: this is for recording compaction events for audit purposes. It does not explicitly mention alternatives or when not to use it, but the specialized language makes the use case evident. There are no exclusions or alternative references, but the context is strong enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_correctionC
Record a user correction to Claude's output (high-priority learning signal)
| Name | Required | Description | Default |
|---|---|---|---|
| reasoning | No | Inferred reason for the correction | |
| session_id | Yes | ||
| claude_action | Yes | What Claude did (tool, file, content) | |
| user_correction | Yes | How the user corrected it |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only states the action is a 'high-priority learning signal', which hints at importance but does not describe side effects, persistence, reversibility, permissions, or any consequences of invoking the tool. This is a significant gap for a mutation-like recording tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, grammatically complete sentence of about ten words. It is front-loaded with the core verb and resource, and the parenthetical adds context without redundancy. Every word earns its place; there is no fluff or tail-heavy content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool involves nested object parameters, no annotations, and no output schema, yet the description stays at a high level. It does not explain how to structure claude_action or user_correction, what counts as a valid correction, or what the tool returns. This leaves an agent under-informed for correct invocation, especially given the complexity of the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, and the description adds no additional meaning beyond what the schema already provides. The phrase 'user correction to Claude's output' loosely maps to claude_action and user_correction but does not clarify their structure or relationships. The 'reasoning' and 'session_id' parameters are not addressed at all in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Record') and identifies the resource ('a user correction to Claude's output'), which clearly conveys the tool's function. The parenthetical '(high-priority learning signal)' adds useful context. However, it does not explicitly differentiate from sibling tools like record_event or record_decision, though 'correction' provides some distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used when a user corrects Claude's output, but it gives no explicit 'when to use' vs. alternatives, no exclusions, and no prerequisites. Sibling tools with overlapping purposes (e.g., record_event, record_decision) are not referenced, leaving the agent without clear selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_decisionC
Record a decision trace: what the agent proposed and how the human responded
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | No | ||
| reasoning | No | ||
| tool_name | No | ||
| session_id | Yes | ||
| decision_type | Yes | ||
| agent_proposal | No | ||
| human_correction | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full burden of behavioral disclosure. It explains what is recorded but not whether the operation is append-only, idempotent, permission-sensitive, or what happens on conflict. No mutation or safety details are given.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that is front-loaded with the core action and object. No filler or redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 7 parameters including nested objects and required fields, the description is far too minimal. It leaves the agent without guidance on required inputs, the decision_type enum, or the structure of nested objects, and there is no output schema to aid expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only hints at 'proposed' and 'human responded', which loosely map to agent_proposal and human_correction. It does not explain required parameters like session_id or decision_type, nor the enum values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's purpose: recording a decision trace with the agent's proposal and human's response. The verb 'record' and resource 'decision trace' are specific, though it doesn't explicitly distinguish from siblings like record_correction or record_event.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided for when to use this tool versus the many sibling tools (e.g., record_correction, record_event). The description does not mention contexts, prerequisites, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_eventC
Record a development event (file edit, test run, etc.)
| Name | Required | Description | Default |
|---|---|---|---|
| success | No | ||
| entities | No | Entity names/paths involved | |
| evidence | No | Tool inputs/outputs, file contents, etc. | |
| reasoning | No | ||
| event_type | Yes | ||
| session_id | Yes | ||
| description | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It only says 'Record' implying a write operation, but doesn't explain persistence, idempotency, success/failure effects, or whether it appends to a log. This is insufficient for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one short sentence, which is concise, but it under-specifies the tool's behavior. While brevity is good, the sentence doesn't earn its place by providing necessary context, making it closer to under-specification than effective conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 7 parameters, no output schema, and no annotations, the description is far from complete. It doesn't explain required fields, return values, or how the event data is used. The description is only a high-level purpose statement, inadequate for safe and correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 29%, so the description must compensate for missing parameter details. It adds examples for event_type ('file edit, test run'), which is redundant with the enum, but it doesn't explain session_id, description, success, or reasoning. The description adds minimal semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the general action ('Record a development event') with examples, so it's more than a tautology. However, it's vague about what constitutes a 'development event' and doesn't distinguish from sibling record tools like record_test_outcome or record_correction, which likely overlap in purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternative record tools. The description doesn't mention any conditions, prerequisites, or exclusions, leaving the agent without direction for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_test_outcomeC
Record test results and link failures to recent code changes
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | ||
| test_results | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It implies a write operation but does not disclose side effects of linking failures, whether it is idempotent, or if it requires an existing session. This lack of behavioral detail leaves the agent uncertain about the tool's impact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler or unnecessary details. It efficiently conveys the main purpose and a secondary linking behavior, making it appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no annotations, no output schema, and a moderately complex input schema with a nested array. The description omits critical information about expected input formats, return behavior, and how the linking works. An agent would likely need additional schema inspection or external knowledge to use this tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not reference session_id or test_results at all. It fails to explain what session_id should be or how to structure the test_results array. The schema provides names/types, but the description adds no semantic value, making it insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the core action ('Record test results') and adds a distinguishing feature ('link failures to recent code changes') that separates it from sibling tools like record_event or record_decision. It is not a mere tautology because it specifies the linking behavior, though it could be more explicit about the scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool vs alternatives. It does not mention prerequisites, exclusions, or contexts where other record tools would be more appropriate. The agent must infer usage from the name and minimal description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_contradictionC
Pick a winner between two contradicting facts using a confidence-weighted strategy (auto, keep_higher_confidence, keep_most_recent, keep_most_sources, supersede_a, supersede_b, manual).
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| strategy | No | ||
| fact_a_id | Yes | ||
| fact_b_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of disclosing behavioral traits. It only mentions the strategy selection and does not disclose side effects, persistence, reversibility, or what happens to the losing fact. This is a material gap for a tool that presumably modifies state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that immediately states the primary action and enumerates strategies in a parenthetical list. Every element serves a purpose, with no fluff or redundant repetition of schema information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of annotations, output schema, and any behavioral details, the description is far from complete. It fails to explain return values, the outcome of the resolution (e.g., which fact is updated), or any prerequisites. This is a mutation-like operation with substantial missing context, making the tool risky to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no descriptions (0% coverage), so the description must compensate. It lists possible strategy values, which adds some meaning, but it does not explain the semantics of each strategy (e.g., what 'auto' does) or describe the 'fact_a_id', 'fact_b_id', and 'notes' parameters beyond their schema names. The compensation is partial and insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Pick a winner' and identifies the resource as 'two contradicting facts,' clearly indicating the tool's purpose. It differentiates from sibling tools like find_contradictions by focusing on resolution rather than detection, though it doesn't explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when a contradiction exists between two facts and provides a list of strategies, which serves as guidance on how to resolve. However, it lacks explicit exclusions or directives about when not to use this tool versus alternatives, leaving some ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_globalC
Search entities across all registered world-model projects
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavioral traits. It does not mention whether the operation is read-only, how results are returned, or any side effects, leaving significant ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no redundant words. It is efficiently phrased, though it lacks additional structured information that a more complete description might include.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with no output schema and no annotations, the description is minimally sufficient but incomplete. It omits return format, pagination behavior, and the meaning of 'entities', leaving the agent without critical execution context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema includes 'query' and 'limit' with no descriptions, and the description adds no parameter-specific information. It does not explain what query syntax is expected or how limit affects results, failing to compensate for the 0% schema description coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (search), the resource (entities), and the scope (across all registered world-model projects). It distinguishes itself from potential siblings by emphasizing the global scope, though it does not explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives like query_fact. The description only states what it does, leaving it to the agent to infer appropriate usage without any exclusions or comparative context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
seed_projectA
Scan the project codebase and populate the knowledge graph with entities and relationships from existing code
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | Re-seed already processed files | |
| project_dir | No | Project directory path (defaults to current) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states that the tool populates the knowledge graph but does not disclose whether the operation is idempotent, whether it modifies existing data, or any side effects. The 'force' parameter hints at re-seeding but this behavior is not explained in the description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, direct, and front-loaded with the main action. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema or annotations, the description is the sole source of behavioral context. It covers the high-level operation but lacks guidance on prerequisites, idempotency, return values, or potential side effects of running a mutating operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters (force and project_dir), so the description does not need to compensate. The description itself adds no parameter semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Scan the project codebase and populate the knowledge graph with entities and relationships from existing code'. It uses a specific verb (scan/populate) and resource (codebase, knowledge graph), distinguishing it from siblings like query_fact (query) or record_event (record).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the tool name 'seed_project' (initial population), but the description provides no explicit when-to-use guidance or alternatives. No mention of when to run this versus other ingestion tools like ingest_pr_reviews or record_event.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
simulate_changeC
Project blast radius and historical outcomes for a proposed change
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| change_description | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only states what the tool projects but does not indicate whether it is read-only, whether it requires specific permissions, or what outputs to expect. The lack of any side-effect or limitation details is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the core action and object. Every word contributes meaning, and there is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description should explain what the tool returns or how the output should be interpreted. It does not, and it also omits any caveats or prerequisites. For a simulation tool, this leaves the agent without enough context to use it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the tool description adds no parameter-level detail. The parameter names 'file_path' and 'change_description' are somewhat self-explanatory, but the description does not clarify expected formats, relationships, or how they map to the 'proposed change' concept, so it fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Project') and resource ('blast radius and historical outcomes for a proposed change'), making the purpose clear. It does not explicitly distinguish from sibling tools like 'predict_regression' or 'validate_change', but the focus on blast radius and historical outcomes is distinctive enough for a 4.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. There is no mention of conditions like 'use when you need to assess impact before applying a change' or references to sibling tools, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_changeB
Validate a proposed code change against known constraints
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| change_type | Yes | ||
| proposed_content | Yes | The new content to validate |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It only states the validation intent but doesn't explain whether validation is read-only, what happens on failure, whether it modifies anything, or what 'known constraints' refers to. This lack of transparency for a tool that could have side effects is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one short sentence that gets to the point. However, at 8 words, it is extremely terse and could have expanded to include usage context without becoming verbose. It is efficient but slightly under-specified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema and no annotations, the description should explain what the validation result looks like and when to use the tool. It only provides the core purpose, missing critical context about return values, behavior on constraint violation, and relationship to other constraint-related tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers 33% of parameters with descriptions (only proposed_content). The tool description does not mention file_path, change_type, or explain the enum values. It adds no semantic value beyond the schema, leaving file_path and change_type under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'validate' with a clear object 'proposed code change' and scope 'against known constraints,' which distinguishes it from sibling tools like simulate_change (which implies running a simulation). It clearly states what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used for checking changes against constraints, but it does not explicitly state when to use it over simulate_change or how it relates to get_constraints. No alternatives are named, and no when-not-to-use guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_retrievalA
Adversarially verify an answer is grounded in a specific set of facts. An independent Coach LLM call checks each material claim in the answer against the supplied source facts and returns confidence (HIGH / MEDIUM / LOW), verified + unverified claim lists, and per-claim source_pointers. Never raises; failures return LOW + error populated. v0.12.12.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The user query the answer responds to | |
| answer | Yes | The candidate answer under verification | |
| fact_ids | Yes | IDs of facts the caller believes ground the answer. Missing IDs are silently dropped. | |
| verification_model | No | Optional Coach model override. Defaults to config.verification_model (Haiku 4.5). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses the independent Coach LLM call, the confidence levels (HIGH/MEDIUM/LOW), the verified and unverified claim lists, per-claim source_pointers, and that it never raises (failures return LOW with error populated). This is excellent behavioral transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the core purpose, followed by behavior and return details. The version tag 'v0.12.12' adds minor noise but does not detract significantly. Overall, every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Since there is no output schema, the description fully explains return values (confidence, claim lists, source_pointers) and error behavior. It is complete for a verification tool, covering what the agent needs to know to invoke it correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description does not add parameter-specific semantics beyond what the schema already provides; it references 'supplied source facts' but leaves parameter details to the schema. This is adequate but not additive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Adversarially verify an answer is grounded in a specific set of facts.' It clearly distinguishes the tool from sibling tools like query_fact or validate_change by emphasizing adversarial verification against supplied source facts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool: whenever an answer needs to be checked against a set of facts. It implies the use case without explicit exclusion or alternative reference, but the context is strong enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.15.5- Added
pin_annotation
2 tool updates
v0.13.0- Added
get_audit_log_head - Added
prove_entry_inclusion
1 tool update
v0.12.13- Added
verify_retrieval
1 tool update
v0.12.0- Changed
query_fact1 field changed- added
Input schema / properties / content_typeAdded value: +{ + "description": "Optional filter by content_type. Use 'procedure' to explicitly summon procedures (which are excluded from auto-injection by design).", + "enum": [ + "rule", + "fact", + "procedure" + ], + "type": "string" +}
1 tool update
v0.7.4- Added
get_agents_md_constraints
26 tool updates
v0.7.3- First observed
export_claude_md - First observed
find_contradictions - First observed
get_co_edit_suggestions - First observed
get_compaction_audit - First observed
get_constraints - First observed
get_context_for_action - First observed
get_decision_log - First observed
get_health_report - First observed
get_injection_context - First observed
get_related_bugs - First observed
ingest_pr_reviews - First observed
predict_regression - First observed
predict_test_failures - First observed
promote_constraint - First observed
query_fact - First observed
recall_transcript_range - First observed
record_compaction_audit - First observed
record_correction - First observed
record_decision - First observed
record_event - First observed
record_test_outcome - First observed
resolve_contradiction - First observed
search_global - First observed
seed_project - First observed
simulate_change - First observed
validate_change
TDQS
Scored across 31 tools
Several tools have poorly separated boundaries: validate_change, simulate_change, predict_regression, and predict_test_failures all appear to assess the impact of a proposed change, differing mainly in subtle emphasis. Similarly, query_fact, search_global, and get_context_for_action overlap heavily in fact retrieval, and record_correction, record_decision, and pin_annotation all capture human feedback. An agent would frequently struggle to select the right tool without reading every description.
Every tool follows a consistent snake_case verb_noun pattern, e.g., get_constraints, record_event, predict_regression, prove_entry_inclusion. Even the more unusual names like pin_annotation and seed_project fit the same imperative structure. This is a highly predictable and uniform naming convention.
31 tools is well beyond the 25+ threshold for a single server and will overwhelm tool selection, especially given the many overlapping prediction and retrieval tools. The server would be more coherent with roughly half the current surface area, consolidating related reads and writes into broader commands.
The domain is broadly covered: it supports knowledge-graph population and queries, event and decision recording, constraint ingestion and validation, regression prediction, audit-log integrity, compaction auditing, and context export. Minor gaps exist, such as no explicit fact/constraint update or delete lifecycle and no project listing tool, but the core workflows are well supported.
Maintenance
Related MCP Connectors
Shared project memory for AI coding agents: decisions, lessons, risks and tasks in one graph.
Codebase graphs, caller impact analysis, and recorded project context for AI coding agents.
Governance layer for AI coding agents: knowledge-graph grounding, session audit, policy controls.
Change-aware CI validation and affected-test guidance for coding agents.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceA temporal knowledge graph system that enables users to record and query architectural decisions, implementation patterns, and project failures. It integrates with Claude to provide hybrid search, timeline tracking, and automated knowledge gap detection using graph analysis.4MIT
- AlicenseNot gradedqualityDmaintenanceProvides AI coding agents with pre-edit situational awareness by combining structural call graphs and co-change history to prevent incomplete edits. It surfaces files that historically change together, reducing missed coupled modules.3MIT
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to query a local, versioned knowledge graph of a software project, retrieving overviews, context packs, evidence, and explanations to make informed changes.MIT
- AlicenseAqualityBmaintenanceEnables AI agents to maintain a persistent knowledge graph of a project, providing dependency context, impact analysis, side-effect discovery, and session recording for more informed coding decisions.5MIT