Skip to main content
Glama

World Model MCP

Persistenter Speicher + optional signiertes Audit für KI-Codierungsagenten.

world-model-mcp ist persistenter Speicher plus ein post-quanten-signierter, offline-verifizierbarer Audit-Trail für KI-Codierungsagenten (FIPS 205 hybrid Ed25519 + SLH-DSA), MIT-lizenziert und vollständig lokal.

world-model-mcp liefert einen lokalen SQLite-Wissensgraphen, den Ihr Agent bei jedem Zug abfragt: Halluzinationen werden überprüfbar, Korrekturen bleiben über Sitzungen hinweg erhalten und Regressionen werden erkannt, bevor sie landen. Aktivieren Sie die Audit-Kette und jedes Ereignis wird mit FIPS 205 hybrid Ed25519 + SLH-DSA signiert, dauerhaft offline verifizierbar. MIT-lizenziert, läuft vollständig lokal, funktioniert mit 10+ KI-Codierungsagenten, darunter Claude Code, Cursor, Codex, Continue, Cline, Windsurf, GitHub Copilot Chat, pi, OpenClaw und Hermes Agent.

Neueste Version: v0.16.2. world-model demo endet mit einem Live-Notar-Beat: Drei Beispielentscheidungen werden in eine echte hybrid-signierte Epoche (Ed25519 + SLH-DSA-SHA2-128f, FIPS 205) signiert, als VALID verifiziert, ein Byte mutiert, um zu beweisen, dass die Manipulationserkennung live auslöst (INVALID), dann wiederhergestellt. Null Anmeldung, null Token, null Netzwerk; legt eine world-model-demo-receipt.json im aktuellen Verzeichnis ab und gibt eine Copy-Paste-etch.systems/verify#<manifest>-URL aus. v0.16.1 fügt einen eleganten Übersprung hinzu, wenn der lokale liboqs-Build ohne SLH-DSA kompiliert wurde (anstatt die Demo zum Absturz zu bringen). Vollständige Versionshistorie in CHANGELOG.md.

Testen Sie es in 10 Sekunden

pip install -U world-model-mcp && world-model demo

Sie sehen drei Beispielentscheidungen, die in eine manipulationssichere Epoche signiert werden, dann mutiert die Demo ein Byte und verifiziert erneut, um zu beweisen, dass die Manipulation sofort erkannt wird (INVALID), dann stellt sie wieder her und verifiziert erneut (wieder VALID). Offline, kein Konto, kein Netzwerk. Eine world-model-demo-receipt.json landet in Ihrem aktuellen Verzeichnis und eine teilbare etch.systems/verify#…-URL wird ausgegeben, damit jeder dieselbe Quittung im Browser überprüfen kann.

PyPI Downloads License: MIT Python 3.11+ world-model-mcp MCP server DOI

mcp-name: io.github.SaravananJaichandar/world-model-mcp


Related MCP server: Scrooge

FAQ

Was ist world-model-mcp? world-model-mcp ist persistenter Speicher plus ein post-quanten-signierter, offline-verifizierbarer Audit-Trail für KI-Codierungsagenten, bereitgestellt als MCP-Server, den der Agent bei jedem Zug abfragt. Es läuft vollständig lokal, wird als MIT-lizenziertes Python + optionale Adapter für 10+ KI-Codierungsagenten ausgeliefert.

Wie unterscheidet es sich von Mem0, Letta und anderen Agent-Speicher-Tools? world-model-mcp ist das einzige Agent-Speicher-Tool, das eine hybride post-quanten-signierte Audit-Kette (FIPS 205 Ed25519 + SLH-DSA-SHA2-128f) zusammen mit dem Speicher ausliefert, mit einem Offline-Referenzverifizierer (etch-verify) und doppelter externer Verankerung (Sigstore Rekor + Bitcoin OpenTimestamps), verfügbar über den gehosteten etch.systems-Begleiter. Vergleichbare Tools (Mem0, Letta, agentmemory) liefern Speicher ohne signierte Audit-Kette. Siehe die Tabelle „Compared to“ unten für eine Aufschlüsselung nach Mechanismus.

Ist der Audit-Trail post-quanten-sicher? Ja. Jede geschlossene Epoche trägt eine hybride Signatur: Ed25519 (klassisch) plus SLH-DSA-SHA2-128f (FIPS 205 zustandslose hashbasierte Post-Quanten-Signatur). Ein zukünftiger Quantengegner, der allein Ed25519 bricht, steht immer noch vor der SLH-DSA-Signatur über dieselbe Nutzlast; beide Signaturen müssen gefälscht werden, damit die Kette gefälscht werden kann.

Wie verifiziere ich einen Datensatz offline, ohne Konto? Führen Sie pip install -U world-model-mcp && world-model demo aus. Die Demo signiert drei Beispielentscheidungen in eine echte hybrid-signierte Epoche, exportiert dasselbe Manifestformat, das die etch-verify-CLI liest, verifiziert VALID, mutiert ein Byte, um zu beweisen, dass die Manipulationserkennung live ist (INVALID), stellt dann wieder her und verifiziert erneut. Eine world-model-demo-receipt.json landet in Ihrem aktuellen Verzeichnis und eine etch.systems/verify#…-URL wird ausgegeben, damit jeder die Quittung im Browser überprüfen kann. Null Anmeldung, null Netzwerk.

Mit welchen KI-Codierungsagenten funktioniert es? Claude Code, Cursor, Codex, Continue, Cline, Windsurf, GitHub Copilot Chat, pi, OpenClaw und Hermes Agent. Jeder hat ein eigenes Starter-Repository unter github.com/SaravananJaichandar/world-model-mcp-<agent>-starter. Das MCP-Wire-Format ist Standard, sodass jeder MCP-fähige Client dieselben Tools verwenden kann.


Vergleich mit anderen Agent-Speicher- und Agent-Audit-Projekten

Direkter Vergleich mit acht benannten Mitbewerbern in den Dimensionen Audit / Signierung / Verankerung, auf die sich der Agent-Speicher-Bereich zubewegt. Jede world-model-mcp-Zelle trägt eine Herkunftsflagge: own = gemessen / beobachtet im ausgelieferten Produkt; cited = aus der eigenen öffentlichen Landingpage, dem Repository oder der Presse des Mitbewerbers übernommen.

Die Quelle der Wahrheit für diese Tabelle befindet sich im Repository des gehosteten Dienstes: world-model-mcp-hosted/src/etch/competitor_matrix.py. Aktualisieren Sie beide Stellen, wenn Sie eine davon ändern.

world-model-mcp (this repo) + Etch (hosted)

Mem0

Letta

agentmemory

Unicity AOS

Repowise

Trinitite

Caura

FailproofAI

Signaturschema

Ed25519 + SLH-DSA-SHA2-128f (hybrid) [own]

Keine (nur Verschlüsselung im Ruhezustand)

Keine

Keine

BLAKE3-Hash-Kette

Deterministische Signale (keine Signierung)

Signiert + hash-verkettet (Algorithmus nicht offengelegt)

Keine

Keine offengelegt

Post-Quanten-bereit

Ja (SLH-DSA, NIST FIPS 205) [own]

Nein

Nein

Nein

Nein (nur BLAKE3-Hash)

Nein

Nicht offengelegt

Nein

Nein

Offline-Referenzverifizierer

etch-verify-CLI, Streaming, wird mit PyPI-Paket ausgeliefert [own]

Nein

Nein

Nein

Nicht offengelegt

Nein (nur SaaS)

Browserbasierter Verifizierer

Nein

Nein

Durchsetzung der Attestierungsfrequenz

Geplanter systemd-Timer + On-Demand-Verifizierung [own]

Nein

Nein

Nein

Nicht offengelegt

Nein

Geplante Attestierungsläufe

Nein

Nein

Framework-Zuordnung (Artikel / Kontroll-IDs)

EU AI Act Art. 12-15, SOC 2 CC6.6/7.2/7.3, ISO 27001 A.12.4/A.14.2 [own]

SOC 2 + HIPAA (Abzeichen, keine Zuordnung pro Kontrolle)

Nicht offengelegt

Nein

Nicht offengelegt

EU AI Act (Negativraum-Behauptung)

Zitate pro Artikel (EU AI Act, SOC 2, SR 11-7)

SOC-2-Abzeichen in Arbeit

Nur SOC-2-Enterprise-Tier

Drift-Erkennung

Kettenintegrität pro Epoche + Attestierungstrend [own]

Nein

Erfolgsrate + Fehlerverfolgung

Nein

Nicht offengelegt

Deltas des Code-Gesundheits-Scores

Deterministische Wiedergabe + Deltas des Risiko-Scores

Nein

Deltas des Bewertungs-Scores

Anmerkungsunterstützung (signierte menschliche Notizen in der Kette)

pin_annotation MCP-Tool, in dieselbe Kette signiert [own]

Nein

Nein

Nein

Nicht offengelegt

Nein

Nicht offengelegt

Nein

Nein

Browser-Verifizierer

Ja, Widget für Kettenintegrität auf /auditor/<slug> [own]

Nein

Nein

Nein

Nicht offengelegt

Nein

Ja

Nein

Nein

Freigabelink (Auditor-Zugriff, ohne Anmeldung)

Ja, ablaufende Freigabe-Tokens (bs_-Präfix, max. 30 Tage) [own]

Nein

Nein

Nein

Nicht offengelegt

Nein

Ja, ablaufender Auditor-Zugriff

Nein

Nein

Externer Anker (unabhängiges Zeugenprotokoll)

Doppelt: Sigstore Rekor + Bitcoin OpenTimestamps (öffentlich, Opt-out pro Projekt) [own]

Nein

Nein

Nein

Nur interne Hash-Kette (kein externes Protokoll)

Nein

Nicht offengelegt

Nein

Nein

OSS-Lizenz

MIT (world-model-mcp) [own]

Apache 2.0

Apache 2.0

Apache 2.0

Nicht offengelegt

AGPL v3 (Kern)

Kein OSS

Apache 2.0

Kein OSS

GitHub-Sterne

Momentaufnahme über etch.systems/api/oss-stats [cited]

61,6k [cited]

23,9k [cited]

24,9k [cited]

7,1k [cited]

4,2k [cited]

Kein öffentliches Repository

373 [cited]

Kein öffentliches Repository

Öffentliche Finanzierung

Bootstrapped [own]

$24,5M [cited]

$10M [cited]

Nicht offengelegt

$3M Seed Feb 2026 [cited]

Nicht offengelegt

Nicht offengelegt

Nicht offengelegt

Nicht offengelegt

Compliance-Haltung (wie auf der Landingpage behauptet)

SOC 2 Type I in Arbeit (Ziel: Aug 2026) [own]

SOC 2 + HIPAA (Abzeichen auf Trust-Subdomain)

Nicht offengelegt

Keine Compliance-Haltung

Keine Compliance-Haltung

Kein SOC 2; EU-AI-Act-Negativraum-Behauptung

Compliance-Rahmen pro Artikel

SOC-2-Abzeichen in Arbeit

SOC-2-Enterprise-Tier

So lesen Sie diese Tabelle:

  • Fettgedruckte Zellen beschreiben Mechanismen, die in diesem Repository (OSS) oder im gehosteten Dienst (etch.systems) ausgeliefert werden.

  • Zellen von Mitbewerbern stammen aus der öffentlichen Landingpage / dem Repository / der Presse des jeweiligen Wettbewerbers. Nicht unsere eigene Messung.

  • [own] in einer world-model-mcp-Zelle bedeutet, dass wir den Mechanismus selbst gemessen / beobachtet haben. [cited] bedeutet, dass der Wert aus einer Drittquelle stammt.

  • Nur die Audit-/Signatur-/Verankerungs-Dimensionen sind in dieser Tabelle enthalten. Allgemeine Gedächtnisfunktionen (Abrufgenauigkeit, Adapterbreite, LLM-Unterstützung) werden im Abschnitt Features unten behandelt.


Zahlen

Benchmark

Ergebnis

Details

SWE-bench Verified repeat-mistake

+10,2 Punkte als Obergrenze für einen einzelnen Versuch (67,3 % → 77,6 % bei 49 gepaarten Instanzen); der Mittelwert über mehrere Seeds beträgt +0,24 pro Instanz, 95 %-KI [0, 0,47]

Vorab registriert, Claude Code 2.1.177 headless, Zenodo-DOI 10.5281/zenodo.21076824. Innerhalb der Domäne +15,0 Punkte, domänenübergreifend +6,9 Punkte mit null Regressionen beim Einzelversuchs-Split.

Widerspruchsauflösung

100,0 % bei auto-Strategie

105 Paare × 19 Kategorien, deterministisch (kein LLM). Ausgeliefert seit v0.11.0.

Coach-Player-Verifizierung

100,0 % exakte Übereinstimmung

12 manuell gekennzeichnete Paare (4 fundiert, 4 teilweise, 4 halluziniert). Layer-3-adversariale Verifizierung über unabhängiges Coach-LLM. Ausgeliefert seit v0.12.12.

Die SWE-bench-Zahl ist die tragende empirische Behauptung. Die anderen beiden sind interne Korrektheits-Benchmarks für ausgelieferte Komponenten. Reproduzierbarkeitsskripte in jedem Benchmark-Verzeichnis oder im verlinkten Repository.


Testen

1.494 Unit-, Integrations- und Fuzz-Tests in der gesamten ausgelieferten Codebasis. Abdeckungsuntergrenze bei 71 % / 66 % (zweistufig für CI-Runner mit und ohne SLH-DSA in liboqs).

# Run tests
pytest -q

# With coverage
pytest --cov=world_model_server --cov-report=term-missing

# Fuzz targets (Atheris, requires Clang / libFuzzer)
python fuzz/fuzz_verify_manifest.py fuzz/corpus/ -max_total_time=60

# Non-Atheris smoke fuzz (runs in every CI pass, no system deps)
pytest tests/test_fuzz_smoke_verify.py -q

Ausgewählte Testsuiten, die erwähnenswert sind:

  • FIPS-205-SLH-DSA-Tests mit bekannten Antworten (tests/test_fips_205_slh_dsa_kat.py): legt Parametergrößen für SLH-DSA-SHA2-128f fest (öffentlicher Schlüssel = 32 Bytes, geheimer Schlüssel = 64 Bytes, Signatur = 17.088 Bytes), Signier-/Verifizierungs-Roundtrip mit Ablehnung von falschem Schlüssel + falscher Nachricht + mutierter Signatur, plus eine feste KAT-Vektor-Fixtur (tests/fixtures/slh_dsa_kat_vectors.json), die für immer wahr ergeben muss.

  • Byte-Parität des Streaming-Verifizierers (tests/test_etch_verify_streaming.py): legt byteidentische Ausgabe zwischen In-Memory- und Streaming-Exporteuren fest, sodass ein Auditor, der das Manifest als Aufzeichnungsartefakt hasht, dieselbe manifest_sha256 erhält, unabhängig davon, welchen Exporteur der Betreiber verwendet hat.

  • Widerspruchs-Benchmark (benchmarks/contradictions-200/): 105 Paare × 19 Kategorien, deterministisch. Legt die auto-Strategie bei jedem Commit auf 100 % fest.

CI erzwingt Abdeckungsuntergrenze + vollständige Testsuiten bei jedem Push. Siehe .github/workflows/pytest.yml.


Authentifizierte Audit-Kette (v0.13+, Opt-in)

Für Bereitstellungen, bei denen der Audit-Trail kryptografisch verifizierbar sein muss (SOC 2, HIPAA, EU AI Act oder Ihre eigene interne Kontrollliste):

export WORLD_MODEL_AUDIT_LOG=on
# then start the world-model-mcp server as normal

Beim ersten Opt-in-Start erstellt der Server zwei neue SQLite-Tabellen in der vorhandenen audit.db-Datei (tamper_evident_log und tamper_evident_epochs) und generiert beim ersten Epochenabschluss ein neues Hybrid-Schlüsselpaar. Jedes nachfolgende Ereignis wird an eine SHA-256-Merkle-Kette angehängt. Wenn eine Epoche abgeschlossen wird (Standard: 1024 Ereignisse), wird die Kettenwurzel mit einem Hybrid-Ed25519 + SLH-DSA-SHA2-128f-Umschlag signiert: klassisch + post-quanten, sodass ein hypothetischer Bruch der elliptischen Kurven-Kryptografie den Prüfpfad intakt lässt.

  • Offline verifizierbar, für immer. Die etch-verify-CLI wird mit dem PyPI-Paket ausgeliefert; Prüfer führen sie auf ihrem Laptop aus, nach dem ersten Download ist keine Netzwerkverbindung erforderlich.

  • Signaturen beweisen die Urheberschaft, nicht nur die Reihenfolge. Hash-Ketten-Alternativen können beweisen, dass nichts umsortiert wurde, aber nicht, WER signiert hat. Diese Kette beantwortet beides.

  • Kein Dashboard, keine Anmeldung, kein externer Dienst. Läuft vollständig in Ihrem Prozess gegen lokales SQLite.

Ausführliche Beschreibung in docs/AUDIT_LOG.md.

Für eine gehostete Version, die KMS-gestützte Schlüssel, ein öffentliches Transparenzprotokoll, externe Verankerung bei Sigstore Rekor + Bitcoin OpenTimestamps sowie ein compliance-orientiertes Operator-Dashboard hinzufügt, siehe den Etch-Begleiter unten.


Schnellstart

Drei häufigste Installationspfade. Für jeden anderen unterstützten Client (Cursor, Cline, Codex, Continue, Copilot, Windsurf, Goose, pi, OpenClaw, Hermes und mehr) siehe etch.systems/docs/install.

Option 1: Claude Desktop (Ein-Klick)

Laden Sie die neueste .mcpb von Releases herunter und ziehen Sie sie in Claude Desktop. Installiert automatisch Hooks, MCP-Server-Konfiguration und Abhängigkeiten.

Option 2: Claude Code / IDE-Plugins (pip install)

# 1. Install the package
pip install world-model-mcp

# 2. Set up in your project (auto-seeds the knowledge graph from existing code)
cd /path/to/your/project
python -m world_model_server.cli setup

# 3. Restart Claude Code
# Done. The world model is pre-populated and active.

Nächster Schritt (optional): Verwandeln Sie das lokale signierte Audit-Log mit dem gehosteten Etch-Notar in eine prüferverifizierbare Kette — KMS-gestützte Schlüssel, öffentliches Transparenzprotokoll, externe Verankerung und einen Freigabelink für Prüfer, ohne eigene Infrastruktur. Siehe Gehosteter Begleiter: Etch.

Option 3: HTTP-Transport für Remote-/MCP-Tunnel-Bereitstellung

pip install 'world-model-mcp[http]'
python -m world_model_server.server --transport http --port 8000

Stellt MCP über Streamable HTTP bereit, sodass Remote-Agenten über das Netzwerk eine Verbindung herstellen können. Siehe docs/http_transport.md für Auth, CORS und Reverse-Proxy-Einrichtung.

Andere Clients

Installation über die OSS-CLI. Jeder Befehl schreibt die korrekte Konfiguration für diesen Client (standardmäßig sys.executable als Interpreterpfad, formatbewusst pro Client, sicher gegen Überschreiben über die Flags --force und --dry-run):

python -m world_model_server.cli install-cursor     # Cursor
python -m world_model_server.cli install-cline      # Cline
python -m world_model_server.cli install-codex      # Codex
python -m world_model_server.cli install-continue   # Continue (also --global)
python -m world_model_server.cli install-copilot    # GitHub Copilot Chat (VS Code Insider)
python -m world_model_server.cli install-windsurf   # Windsurf
python -m world_model_server.cli install-pi         # pi
python -m world_model_server.cli install-openclaw   # OpenClaw
python -m world_model_server.cli install-hermes     # Hermes (MCP mode)
python -m world_model_server.cli install-hermes-provider  # Hermes (Elixir-native provider)

Vollständige Schritt-für-Schritt-Anleitungen pro Client (mit Verifizierungsbefehlen + Fehlerbehebung) unter etch.systems/docs/install.


Was es tut

world-model-mcp ist ein temporaler Wissensgraph, der zwischen Ihrem KI-Codierungsagenten und seiner Arbeit sitzt. Er erfasst Fakten, Entitäten und Einschränkungen aus Ihrer Codebasis; validiert jede Codeänderung an der Bearbeitungsgrenze gegen gelernte Einschränkungen; injiziert nach der Kontextfenster-Kompaktierung relevanten Kontext erneut; verfolgt Widersprüche mit konfidenzgewichteter Auflösung; und verifiziert Abrufe adversarisch über ein unabhängiges Coach-LLM.

Funktionen

1. Halluzinationsprävention. Jeder erfasste Fakt trägt eine Herkunft (asserted_by, confirmer, confirmation_state, evidence_type). Wenn der Agent einen Fakt abfragt, erhält er den Konfidenzwert zusammen mit der Antwort. Wenn zwei Fakten widersprüchlich sind, gewinnt der neuere/konfidentere; der Verlierer wird mit superseded_by beibehalten. Der Agent muss nie raten, ob ein gespeicherter Fakt noch aktuell ist.

2. Lernen aus Korrekturen. Wenn Sie einen Agenten korrigieren (record_correction), wird die Korrektur als Ereignis erster Klasse mit den beteiligten Entitäten gespeichert. Wenn der Agent das nächste Mal etwas abfragt, das diese Entitäten betrifft, taucht die Korrektur zuerst auf. Korrekturen bleiben über Sitzungen, Kontextkompaktierungen und Agentenneustarts hinweg bestehen.

3. Regressionsprävention. Jeder Codeänderungsvorschlag läuft durch validate_change, das den Einschränkungsgraphen durchläuft und Verstöße zurückgibt, bevor die Änderung angewendet wird. Einschränkungen werden automatisch aus Ihrer Codebasis gelernt (seed_project), durch PR-Review-Kommentare verstärkt (ingest_pr_reviews) und manuell verfasst (record_event). Der Agent sieht den Verstoß, den Vorschlag und die Quelle der Einschränkung.

4. Coach-Player-adversarische Verifizierung. Ein Player-Agent entwirft eine Entscheidung. Ein unabhängiger Coach-Agent fragt den Graphen unabhängig nach Präzedenzfällen ab, durchläuft den Merkle-Beweis, prüft die Hybrid-Signatur und genehmigt erst dann. Der Coach vertraut nie der Zusammenfassung des Players; er verifiziert jedes Mal gegen ein signiertes Ledger. 100 % exakte Übereinstimmung bei 12 handbeschrifteten Paaren.

Vollständige technische Architektur in docs/ARCHITECTURE.md.


MCP-Tools

Acht Tools werden mitgeliefert. Einzeilige Zusammenfassung; klicken Sie für Signatur + Beispiele in docs/mcp/ auf jedes.

Tool

Was es tut

query_fact

Ruft einen gespeicherten Fakt mit Herkunft + Konfidenz + Quelle ab

record_event

Hängt ein Ereignis an die Audit-Kette an (event_type aus einem festen Enum)

validate_change

Prüft einen Codeänderungsvorschlag gegen gelernte Einschränkungen, gibt Verstöße + Vorschläge zurück

get_constraints

Listet alle Einschränkungen auf, die zu einer Entität oder einem Dateimuster passen

record_correction

Speichert eine Benutzerkorrektur, sodass sie bei der nächsten relevanten Abfrage wieder auftaucht

get_related_bugs

Ruft Bugs ab, die die angegebenen Dateien betreffen + ein Risikobewertung pro Datei

seed_project

Massenimport einer bestehenden Codebasis in den Wissensgraph

ingest_pr_reviews

Wandelt aktuelle PR-Review-Kommentare in Einschränkungen um

Vollständige Tool-Dokumentation unter docs/mcp/.


Gehosteter Begleiter: Etch

Sie betreiben world-model-mcp lokal? Etch (etch.systems) ist die gehostete Governance-Ebene, die auf demselben OSS-Kern aufbaut und hinzufügt, was ein Compliance-Team benötigt, um den Produktionseinsatz freizugeben:

  • KMS-verschlüsselte Signaturschlüssel (niemals Klartext im Ruhezustand)

  • Öffentliches Transparenzprotokoll mit signiertem Kopf (Split-View-Resistenz)

  • Externe Verankerung bei Sigstore Rekor + Bitcoin OpenTimestamps (doppelt unabhängiger Zeuge)

  • Operator-Dashboard mit Sitzungsabläufen, Kettenintegritätsansicht, PII-Scanning und PDF-Export der Client-Antworten

  • Eigenständige Auditor-CLI + Browser-Verifizierer unter /auditor/<slug>

  • Kostenlose Stufe, darüber hinaus pay-as-you-grow

Dieselben Krypto-Primitive, dasselbe Audit-Log-Schema, null Quellcodeänderungen am ausgelieferten OSS. Überspringen Sie dies, wenn Sie eigenständig arbeiten.

Welches ist das Richtige für mich?

Situation

Route

Ich möchte persistenten Speicher für meinen lokalen KI-Codierungsagenten

world-model-mcp (dieses Repo, MIT)

Ich möchte Agentenentscheidungen sechs Monate später einer Regulierungsbehörde nachweisen

etch.systems

Ich möchte signierte Beweisbündel, die ein Gericht selbst authentifizieren kann

etch.systems

Ich möchte team- oder anbieterübergreifende Föderation signierter Audit-Historie

etch.systems

Ich möchte beides ausprobieren

Lokal mit world-model-mcp starten; Etch hinzufügen, wenn Sie signierte Beweise für andere zur Verifizierung benötigen


So funktioniert es

world-model-mcp ist ein MCP-Server (stdio oder Streamable HTTP), der einen temporalen Wissensgraph bereitstellt, der von SQLite unterstützt wird. Jede Agentenrunde fragt den Graphen ab; jede Codeänderung oder Korrektur schreibt zurück. Beim Start seedet der Server automatisch aus Ihrer bestehenden Codebasis; beim Herunterfahren leert er sauber.

Sechs Datenbanken unter .claude/world-model/:

  • entities.db: Dateien, Funktionen, Klassen, Symbole

  • facts.db: semantische Fakten mit Herkunft + Konfidenz

  • relationships.db: Abhängigkeiten, Aufrufe, Importe

  • constraints.db: gelernte Regeln, die der Agent beachten muss

  • sessions.db: sitzungsbezogene Kontextverfolgung

  • events.db: unveränderliches Ereignisprotokoll (Audit-Kette bei Opt-in)

Vollständige Architektur mit Diagrammen in docs/ARCHITECTURE.md.


Konfiguration

Umgebungsvariablen (alle optional):

Variable

Zweck

Standard

WORLD_MODEL_AUDIT_LOG

Aktiviert die signierte Audit-Kette

off

WORLD_MODEL_DB_PATH

Überschreibt das SQLite-Datenbankverzeichnis

.claude/world-model

WORLD_MODEL_TELEMETRY

Opt-in für anonyme OSS-Telemetrie (siehe Datenschutz)

off

ANTHROPIC_API_KEY

Nur verwendet, wenn Sie optionale LLM-gestützte Funktionen aktivieren

keine

Vollständige Konfigurationsreferenz in docs/CONFIGURATION.md.


Datenschutz und Sicherheit

  • Telemetrie ist standardmäßig deaktiviert. Wenn Sie mit WORLD_MODEL_TELEMETRY=on opt-in, werden aggregierte Installationsmetriken an etch.systems/api/telemetry/ingest gesendet (Endpunkt-URL, kein Quellcode, keine Prompts, keine PII). Recht auf Löschung wird über DELETE /api/telemetry/install/{install_id} unterstützt.

  • Für den Kernbetrieb ist kein ANTHROPIC_API_KEY erforderlich. Einige optionale Funktionen (Coach-Player-Layer-3-Verifizierung, LLM-gestütztes Re-Ranking) verwenden die API, wenn ein Schlüssel bereitgestellt wird. Ohne ihn funktioniert alles andere.

  • Sicherheitsproblem melden: E-Mail an security@etch.systems. PGP-Schlüssel auf Anfrage.

Ausführliche Beschreibung in docs/PRIVACY.md.


Mitwirken

Beiträge sind willkommen. Siehe CONTRIBUTING.md für Entwicklungseinrichtung, Codierungsstandards, Hinzufügen von Sprachunterstützung, Schreiben von Tests und Einreichen von PRs.

Bereiche, in denen Hilfe besonders erwünscht ist:

  • Sprachparser (Go, Rust, Java, C++)

  • Zusätzliche MCP-Client-Adapter

  • Framework-Integrationen (LangGraph, CrewAI, AutoGen, LlamaIndex; Starter-Shims existieren im gehosteten Repo)

  • Benchmark-Beiträge in benchmarks/

Bitte lesen Sie CLA.md vor Ihrem ersten PR.


Lizenz

MIT-Lizenz. Kostenlos für kommerzielle und persönliche Nutzung.


Available Tools

31 tools
export_claude_mdB

Generate a CLAUDE.md document from the knowledge graph (top constraints, recent decisions, known bug regions, co-edit patterns).

ParametersJSON Schema
NameRequiredDescriptionDefault
max_constraintsNo

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavior. It states it generates a document but does not clarify whether it writes to a file, returns content, or has side effects. It also fails to mention any permissions, limits, or output format, leaving key behavioral aspects ambiguous.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with a clear verb and object, followed by a parenthetical list of content categories. It is concise and contains no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description should explain return values, side effects, and parameter behavior. It only covers purpose and content categories, leaving critical operational details absent for an export tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for the only parameter, max_constraints. The description broadly references 'top constraints' but does not explain that max_constraints limits that section, nor its units or effect on other sections. The parameter semantics are only indirectly inferable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Generate') and identifies the resource ('CLAUDE.md document from the knowledge graph'), and it lists the included content categories (top constraints, recent decisions, known bug regions, co-edit patterns), distinguishing it from sibling tools that query individual aspects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for producing an aggregated CLAUDE.md from knowledge graph data, but it does not explicitly state when to use this tool versus querying individual components via siblings like get_constraints or get_decision_log. It lacks explicit exclusions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_contradictionsC

Find pairs of facts that contradict each other based on similarity and status differences

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of disclosing behavior. It doesn't state whether the tool is read-only, what the output looks like, or any side effects. The hint about 'similarity and status differences' is the only behavioral clue, but it's insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise and front-loaded, but it sacrifices necessary detail. It's not bloated, but it's under-specified. The brevity doesn't earn its place because it lacks critical information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has two parameters, no output schema, and no annotations. The description provides no information about return values, parameter usage, or behavioral context. It is far from complete even for a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has two parameters (limit, query) with zero descriptions, and the description doesn't mention them at all. Since schema coverage is 0%, the description fails to compensate, leaving the agent with no understanding of what these parameters do.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool finds pairs of contradicting facts, using a specific verb ('Find') and resource. It adds a hint of the method ('similarity and status differences'), but doesn't explicitly differentiate from sibling tools like resolve_contradiction, though the distinction is evident from the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus alternatives. It doesn't mention any exclusions or prerequisites, leaving the agent to infer usage solely from the vague description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agents_md_constraintsA

Parse AGENTS.md / CLAUDE.md / GEMINI.md / .agents/skills/*.md in the project and return declarative constraints. Mixed into PreToolUse enforcement automatically; this tool exposes the same data for inspection.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathNo
project_dirNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden for behavioral disclosure. It explains the parsing action and its role in enforcement, but does not mention whether it performs a read-only operation, how missing files are handled, or any potential side effects. The context about automatic enforcement adds value but lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the main action and scoped to specific files. Every phrase contributes meaning, and the inspection/enforcement contrast adds valuable context without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's purpose and its relationship to PreToolUse enforcement, giving useful context. However, it omits parameter semantics and the exact return format (beyond 'declarative constraints'), which is especially important since there is no output schema. The lack of detail on expected inputs and outputs leaves gaps for an agent trying to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has two parameters (file_path and project_dir) with no descriptions and 0% coverage. The description does not mention either parameter, leaving their specific purpose and format ambiguous. For example, it is unclear if file_path is optional or relative to project_dir. The description fails to compensate for the complete lack of schema guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it parses AGENTS.md / CLAUDE.md / GEMINI.md / .agents/skills/*.md and returns declarative constraints. This specific verb+resource distinguishes it from the sibling tool get_constraints, which likely covers broader constraint sources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes that the data is 'Mixed into PreToolUse enforcement automatically' and that this tool 'exposes the same data for inspection,' implying it is intended for inspection/debugging rather than direct enforcement. It provides clear context but does not explicitly name alternatives or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_audit_log_headA

v0.13 tamper-evident audit log. Return the current head state (last log entry seq, last closed epoch seq, unclosed-entry count) plus the full closed-epoch chain with hybrid signature envelopes. Compliance auditors call this periodically to verify no operator misbehavior has occurred since the last check. Requires WORLD_MODEL_AUDIT_LOG=on at server startup.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the tamper-evident nature, the return content, and a runtime prerequisite (WORLD_MODEL_AUDIT_LOG=on). It adds meaningful operational context beyond a simple 'get' but stops short of stating read-only semantics explicitly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two sentences but packs in the purpose, return value details, use case, and a critical prerequisite. Every clause earns its place with no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter, read-only operation with no annotations or output schema, the description covers all essential aspects: what it does, what it returns, when to use it, and what server configuration is required. An agent has enough to deploy it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the schema is empty. The description explains what the tool returns, which is the only meaningful semantic content in this case. Baseline 4 applies since there are no parameters to describe.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Return' and names the resource 'audit log head state' plus the 'closed-epoch chain', clearly defining what the tool does. This differentiates it from other audit-related siblings like get_compaction_audit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states a use case: 'Compliance auditors call this periodically to verify no operator misbehavior has occurred since the last check.' This gives clear context on when to use it, but it does not mention alternatives or exclusions, so it misses the top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_co_edit_suggestionsB

Get files commonly edited alongside the given file based on historical patterns

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
file_pathYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden for behavioral disclosure. It mentions 'based on historical patterns' which adds some context, but it does not describe whether the operation is read-only, how suggestions are ranked, what format the results take, or any potential side effects. This is insufficient for full transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that is front-loaded with the action and resource. Every word earns its place, and there is no superfluous content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with two parameters, but the description leaves major gaps: no usage guidance, no parameter details, and no mention of return values (no output schema). For an agent to use this correctly, it needs more context about how suggestions are generated and what to expect from the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero description coverage (0%) for its parameters, and the description does not compensate. It references 'the given file' (mapping to file_path) but does not clarify the expected format, the meaning or usage of 'limit', or any constraints. The description adds minimal value beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'files commonly edited alongside the given file', with the basis 'historical patterns'. It is specific and distinguishes itself from the sibling tools, none of which share a similar focus on co-edit suggestions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any exclusions, prerequisites, or conditions, leaving the agent without context for selecting it among the many sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_compaction_auditA

List recent compaction audit entries, most-recent first. Filter by session_id or limit count.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
session_idNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the transparency burden. It discloses the ordering (most-recent first) and the available filters (session_id, limit), which is useful. However, it does not mention whether the operation is read-only, whether limit has a default, or describe the return format or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that delivers all key information without redundancy. Every clause adds functional value, and there is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list/filter tool, the description covers the essential purpose and options. However, with no output schema and no annotations, the agent is left without knowledge of the returned fields, default limit behavior, or error cases. This is adequate for simple use but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must add meaning. It does clarify that session_id filters and limit controls the count, which is helpful. However, the semantics are shallow: it does not specify whether limit is mandatory, its maximum/default value, or whether session_id requires exact or partial matching.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists compaction audit entries in reverse chronological order. It names the specific resource (compaction audit entries) and the action (list), which differentiates it from write-oriented siblings like record_compaction_audit. However, it does not explicitly distinguish from the similar get_audit_log_head tool, preventing a perfect score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for inspecting compaction audit history, and the mention of filtering suggests relevant scenarios. However, it provides no explicit guidance on when to prefer this over alternatives like get_audit_log_head, and there are no exclusions or prerequisites stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_constraintsA

Get constraints (linting rules, patterns, conventions) for a file

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes
constraint_typesNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It implies a read-only operation ('Get') but does not disclose potential behaviors such as error handling for missing files, whether it searches the entire repository, or if any filtering is applied beyond the optional constraint_types parameter. The description adds some clarity by defining constraints, but does not reveal side effects or edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that efficiently communicates the tool's purpose without extraneous words. It earns a high score for being concise and structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (read operation, 2 parameters, no output schema), the description is mostly complete. It states what the tool does and the schema covers parameter details. However, it does not specify the return format or behavior when no constraints are found, which could leave some ambiguity for the agent. Still, for a straightforward getter, it is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, so the description must compensate. It does so by explaining that constraints include 'linting rules, patterns, conventions', which helps interpret the constraint_types enum. However, it does not explicitly map parameters or explain the file_path semantics beyond the name. The schema itself provides the enum values, offering adequate baseline coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves constraints for a file, using the verb 'Get' and specifying the resource and scope. It also clarifies what constraints are (linting rules, patterns, conventions). However, it does not explicitly distinguish from the sibling tool 'get_agents_md_constraints', which may overlap in purpose for specific files.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when constraints for a file are needed, but provides no explicit guidance on when to use this tool versus alternatives like 'get_agents_md_constraints' or when not to use it. It lacks exclusions and alternative recommendations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_context_for_actionC

Pre-action context bundle: constraints, decisions, bugs, co-edits, related facts, and risk score for a file before editing

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes
action_typeYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It lists the content components but does not state whether the tool is read-only, how errors are handled, what the risk score means, or how the bundle is returned. This is a significant gap for a retrieval tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a concise noun phrase that front-loads the core purpose and lists contents, but it lacks a verb and reads more like a label than a full sentence. It is not overly verbose, yet it could be more structured with a clearer main clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two parameters, no annotations, and no output schema, the description is too sparse. It does not explain the output format, the meaning of the risk score, or how this bundle relates to the individual context tools. The description leaves significant gaps in understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage. The description mentions 'file' which maps to file_path, but action_type is only implied by 'editing' and not explicitly explained. The enum values (edit, create, delete, refactor) are not described, so the description adds minimal meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a pre-action context bundle for a file, listing specific content types such as constraints, decisions, bugs, co-edits, related facts, and risk score. This distinguishes it from sibling tools that target single context types. However, the phrase 'before editing' is slightly inconsistent with the action_type enum which includes create, delete, and refactor.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide explicit guidance on when to use this bundle versus the many sibling tools like get_constraints or get_related_bugs. It implies usage before an action, but there are no exclusions or alternative recommendations, leaving the agent to infer the appropriate context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_decision_logC

Get decision traces showing agent proposals and human corrections

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
file_pathNo
session_idNo
decision_typeNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It implies a read operation via 'Get', but does not disclose behavior such as default limits, ordering, filtering effects, or whether it is purely read-only. The description focuses on content, not operational behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no wasted words. It efficiently states the purpose without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has four parameters, no output schema, and no annotations, yet the description provides no details on return value, parameter usage, or edge cases. It is overly minimal for the tool's apparent complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description does not explain any parameters. It only hints at decision_type through 'corrections', but limit, file_path, and session_id are completely unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves 'decision traces' with specific content (agent proposals and human corrections), using a specific verb 'Get' that distinguishes it from sibling write tools like record_decision and record_correction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as get_audit_log_head or get_compaction_audit. There is no mention of scenarios, exclusions, or preferred contexts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_health_reportA

Memory health diagnostics: orphans, stale facts, contradictions, decay candidates, DB sizes

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden. The term 'diagnostics' implies a read-only operation, and the listed categories provide concrete insight into what the tool examines. However, it does not explicitly state whether any state is modified or note any performance implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise phrase that effectively communicates purpose and scope without wasted words. It is front-loaded and every element adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter diagnostic tool with no output schema, the description provides sufficient context about the report's contents. It could be improved by noting the output format, but the current level is adequate for understanding the tool's purpose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description need not document inputs. The baseline of 4 is appropriate because no parameter information is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function as memory health diagnostics and enumerates specific areas it covers (orphans, stale facts, contradictions, decay candidates, DB sizes). This distinguishes it from narrower sibling tools like find_contradictions or get_compaction_audit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use for overall health assessment, but it does not explicitly state when to prefer this tool over siblings or when not to use it. No alternative tools are mentioned, so guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_injection_contextB

Return a compact constraint+fact bundle for PostCompact / UserPromptSubmit hooks to re-inject after context loss.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_factsNo
event_typeYes
project_hintNo
max_constraintsNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden. It only says 'Return...' without disclosing whether this is a read-only operation, any side effects, how 'compact' is achieved (e.g., truncation, filtering), or what happens if parameters like max_facts are omitted. This lack of behavioral detail is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It is concise and easy to read, though it sacrifices substance for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 4 parameters, zero schema coverage, no annotations, and no output schema, the description is insufficiently complete. It gives a high-level purpose but lacks details on parameter behavior, return format, and how this tool fits with alternatives. More context is needed for an agent to use it reliably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it does not. It mentions two of the three event_type values but does not explain max_facts, max_constraints, project_hint, or the SessionStart event. The term 'compact' weakly implies size limits, but no explicit parameter meaning is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a 'compact constraint+fact bundle' for specific hooks (PostCompact/UserPromptSubmit), which distinguishes it from sibling tools like query_fact or get_constraints. However, it could be more explicit about what 'context loss' entails or how it differs from get_context_for_action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names specific trigger events (PostCompact/UserPromptSubmit) and the goal of re-injecting after context loss, giving clear context for when to use the tool. It does not explicitly mention alternatives or when not to use it, but the event-specific framing provides adequate guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ingest_pr_reviewsA

Pull GitHub PR review comments and convert them into learned constraints in the knowledge graph

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoGitHub repo (owner/repo). Auto-detected from git remote if omitted.
countNoNumber of recent PRs to scan (default 10, max 50)

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses that the tool converts PR comments into learned constraints, indicating a write operation to the knowledge graph. However, it does not mention idempotency, overwrite behavior, permissions, or any side effects beyond the conversion, which is minimal for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that directly states the action and outcome with no filler. Every word contributes to the meaning, making it highly efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has only two optional parameters and no output schema, so minimal description might suffice. However, because it is an ingest operation affecting the knowledge graph, a bit more context about result expectations or side effects would improve completeness. It is adequate but not fully comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds no extra parameter-specific meaning beyond what the schema already provides (repo auto-detection, count default/max). It does not compensate for any missing details, but none are missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('Pull' and 'convert') and identifies the resource ('GitHub PR review comments') and target ('knowledge graph'). It clearly distinguishes itself from sibling tools like get_constraints or record_event by describing a unique ingest-and-transform workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when you want to bring PR review comments into the knowledge graph, but it does not explicitly state when to prefer this over alternatives or provide exclusions. Sibling tools like record_correction or validate_change serve different purposes, yet no direct comparison is offered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pin_annotationA

Attach a signed human annotation (note, override rationale, or intervention record) to a span of agent events. Persists into the annotations table and chains into the same Merkle audit log as agent writes (v0.15.0, ADR-0001). Rationale limited to 8 KB.

ParametersJSON Schema
NameRequiredDescriptionDefault
authorYesAuthor identity. Self-asserted in OSS; KMS-verified in Etch hosted.
rationaleYesHuman rationale text (UTF-8). Max 8192 bytes.
session_idYesSession containing the annotated events.
annotation_typeYes
event_range_endYesLast event_id in the annotated span. Equals event_range_start for a single-event annotation.
event_range_startYesFirst event_id in the annotated span.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and exceeds expectations by disclosing persistence ('Persists into the annotations table'), audit integration ('chains into the same Merkle audit log as agent writes'), version/ADR references, and the rationale size limit. This gives the agent a clear behavioral model.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core action, and includes only essential extra context (persistence, audit log, limit). Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 6 required parameters, no output schema, and no annotations, the description provides solid context for a write operation: it explains persistence and audit chaining. It lacks explicit error handling or return behavior, but that is not critical for a basic mutation tool. The description is largely complete for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 83% coverage (5 of 6 parameters have descriptions), so the baseline is 3. The description adds little beyond the schema; it mentions the 8 KB rationale limit (already in schema) and does not elaborate on parameter meanings or relationships. The schema itself is well-documented, so a baseline score is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action ('Attach') and object ('a signed human annotation to a span of agent events'). It distinguishes from siblings like record_event by emphasizing human annotation vs agent events and mentions specific annotation types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case of attaching human notes or overrides to event spans, but provides no explicit guidance on when to choose this tool over alternatives such as record_correction or record_decision. The context is clear but lacks exclusions or alternative recommendations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

predict_regressionA

Score regression risk for a proposed change to a file based on past bugs, test failures, and constraint violations

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes
change_descriptionNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses what the tool uses (past bugs, test failures, constraint violations) but does not state whether it has side effects, permissions requirements, or what the output looks like. The methodology hint adds some behavioral context, but safety and operational traits remain undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It efficiently conveys the tool's purpose and key inputs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two simple parameters and no output schema, the description is adequate for tool selection but does not fully prepare the agent for invocation. It lacks details on input semantics and return value format, though the simplicity and sibling context make it acceptable but not thorough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions 'proposed change to a file' which loosely maps to file_path and change_description, but it does not clarify parameter formats, required fields, or examples. The description adds minimal meaning beyond what the property names already imply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Score') and resource ('regression risk') with the basis ('past bugs, test failures, and constraint violations'). It clearly distinguishes from siblings like predict_test_failures (which targets test failures specifically) and simulate_change (which simulates changes).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for assessing regression risk of a proposed change to a file, but it does not explicitly state when to prefer this tool over alternatives or provide exclusions. Sibling tools like predict_test_failures or validate_change could overlap, and no guidance is given on choosing among them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

predict_test_failuresA

Surface tests likely to fail given a set of edited files

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathsYes

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the full burden for behavioral transparency. It only states the core action without disclosing whether the operation is read-only, what data it relies on, what output format to expect, or any limitations. The description is minimal and leaves the agent without important contextual cues.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the action and avoids redundancy. Every word contributes to the core purpose, making it highly concise and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is relatively simple (one parameter, no output schema, no annotations), but the description is thin. It communicates the essential purpose, but does not explain what the tool returns (e.g., a list of test names) or provide operational boundaries. It is minimally complete for selection but not fully adequate for invocation without further inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no descriptions for the parameter file_paths (0% coverage). The description adds semantic value by indicating these are 'edited files', clarifying the parameter's intent. However, it does not specify path format, file existence requirements, or other constraints that would be useful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Surface' with a clear resource ('tests likely to fail') and a scope condition ('given a set of edited files'). This distinguishes it from sibling tools like predict_regression and get_related_bugs, which address different aspects of change impact.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'given a set of edited files' implies the primary use case, but the description does not explicitly state when to use this tool over alternatives or provide any exclusions. It lacks guidance on how to compare with similar tools like predict_regression.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

promote_constraintB

Promote a constraint from this project to all other registered projects

ParametersJSON Schema
NameRequiredDescriptionDefault
constraint_idYes
target_projectsNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals the action 'promote' but does not disclose side effects, permissions required, reversibility, or the fact that it likely mutates multiple projects. This leaves significant behavioral uncertainty.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the verb and object. It contains no filler or unnecessary detail, making it easy to parse and understand.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema and no annotations, and the description is minimal. It lacks essential context about the promotion's effects, error conditions, prerequisites, and return values. For a cross-project mutation tool, this is insufficient for an agent to use safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for parameter meaning. It implies constraint_id identifies the constraint, but it does not explain target_projects, its optionality, or how it interacts with the default 'all other registered projects'. The description adds minimal meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'promote' with a clear resource, 'constraint', and a clear scope, 'from this project to all other registered projects'. This distinguishes it from sibling tools like get_constraints (retrieval) and validate_change (validation), making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidance is provided. The description does not mention when to use this tool versus alternatives, nor any prerequisites or conditions for promotion. It only states the action itself, leaving the agent without context for decision-making.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prove_entry_inclusionA

v0.13 tamper-evident audit log. Return a cryptographic inclusion-proof bundle for a persisted row_id (fact, constraint, event, or decision ID). Bundle includes the entry, the containing signed epoch (Ed25519 + SLH-DSA hybrid signature envelope), an RFC 6962 Merkle inclusion proof, and the full epoch chain from genesis. Requires WORLD_MODEL_AUDIT_LOG=on at server startup; returns an error object when opt-in is off, when the row_id is not found, or when the entry is in the unclosed backlog.

ParametersJSON Schema
NameRequiredDescriptionDefault
row_idYesID of the fact / constraint / event / decision to prove inclusion for.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses the exact bundle components (entry, signed epoch, Merkle proof, epoch chain), required server flag, and all error cases, giving a thorough behavioral profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single information-dense sentence that front-loads the primary action ('Return a cryptographic inclusion-proof bundle') and then enumerates bundle contents and failure modes without unnecessary filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by listing the bundle components explicitly. It also covers prerequisites and error scenarios, making it complete enough for a single-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (row_id with description), but the description adds semantic value by specifying that row_id refers to fact, constraint, event, or decision ID, clarifying the expected input beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a cryptographic inclusion-proof bundle for a persisted row_id. It names the specific resource (audit log entries) and distinguishes it from siblings like get_audit_log_head or query_fact by focusing on inclusion proofs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use it (to prove inclusion of an entry), prerequisites (WORLD_MODEL_AUDIT_LOG=on), and error conditions (opt-in off, row_id not found, unclosed backlog). It does not explicitly mention alternatives, but the scope is clear enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_factB

Query the knowledge graph for facts about entities (APIs, functions, classes, etc.)

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesThe search query (e.g., 'User.findByEmail', 'JWT authentication')
contextNoAdditional context for the query
entity_typeNoOptional filter by entity type
content_typeNoOptional filter by content_type. Use 'procedure' to explicitly summon procedures (which are excluded from auto-injection by design).

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full responsibility for behavioral disclosure. It doesn't state whether the tool is read-only, what the response format is, or that it can retrieve 'rules' and 'procedures' despite being named 'facts'. The term 'facts' may mislead users into thinking only fact-type content is returned, leaving important behavior undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that is front-loaded with the action and resource. Every word contributes to the core purpose without fluff or repetition. It is appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a general-purpose query tool with no output schema and no annotations, the description is minimal but adequate. It does not explain return values, pagination, or how to decide between this and related sibling tools. The ambiguous scope of 'facts' (vs rules/procedures) is a notable gap, but the schema partially covers this via content_type descriptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema coverage is 100% with descriptive text for all parameters, so the baseline is 3. The tool description adds no parameter-level meaning beyond what the schema already provides; it only reiterates the general entity focus. The description does not compensate or enhance the schema's parameter docs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a clear verb ('Query') and resource ('knowledge graph') and specifies the object ('facts about entities'). It lists example entity types (APIs, functions, classes) which conveys scope. However, it doesn't distinguish itself from sibling tools like search_global that might also search the knowledge graph, so it lacks explicit differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance in the description about when to use this tool versus alternatives. The only usage hint ('Use "procedure" to explicitly summon procedures...') appears in the schema's content_type parameter, not the tool description, and it is parameter-level rather than tool-level. No when-not-to-use or alternative tool references are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recall_transcript_rangeC

Hydrate a Claude Code session transcript by line range. Lets agents trace a fact back to the exact conversation that produced it.

ParametersJSON Schema
NameRequiredDescriptionDefault
line_endNo
line_startNo
session_idYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only says 'Hydrate' without explaining whether the operation is read-only, what happens with invalid ranges, whether there are performance or memory implications, or what the response contains. This is a significant gap for a data-retrieval tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences and front-loaded with the core action. Each sentence contributes meaning: the first states the operational scope, the second provides the motivating use case. It is well-structured and free of fluff, though slightly vague.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and a 0% parameter coverage, the description should provide more contextual information about return values, error handling, and when to choose this tool. The current description is insufficient for an agent to confidently invoke the tool correctly in all cases, even though the tool itself is relatively simple.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions 'by line range,' which somewhat clarifies line_start and line_end, but it does not explain session_id, the inclusive/exclusive nature of the range, defaults, or behavior when only one line parameter is provided. The description adds minimal semantic value beyond the schema fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool hydrates a Claude Code session transcript by line range, with a specific use case of tracing facts to their source conversation. This distinguishes it from sibling tools like query_fact, which focus on facts rather than raw transcript retrieval. However, the verb 'hydrate' is somewhat jargon-heavy, and it lacks an explicit contrast with related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool ('Lets agents trace a fact back to the exact conversation that produced it'), which provides some context. However, it does not explicitly state when to prefer this over alternatives like query_fact or get_audit_log_head, nor does it give exclusions or prerequisites. The guidance is present but inferred rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_compaction_auditA

Record a context-compaction event with token counts and what was re-injected. Lets developers audit what was remembered across compaction boundaries.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNo
raw_summaryNo
facts_injectedNo
injection_eventNo
pre_compact_tokensNo
post_compact_tokensNo
constraints_injectedNo

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description must disclose behavioral traits. It states it records an event, implying a write operation, but does not mention side effects, whether data is appended or overwritten, permissions, or failure modes. This is a significant gap for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exactly two sentences, front-loaded with the core action and followed by the purpose. Every word adds value with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 optional parameters, no output schema, and no annotations, the description does not fully equip an agent to use the tool. It lacks information about expected return values, whether any parameters are required in practice, and how this differs from other record_* tools beyond the compaction focus. This is insufficient for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It mentions 'token counts' and 'what was re-injected,' which roughly maps to pre_compact_tokens, post_compact_tokens, facts_injected, and constraints_injected, but it does not explain parameters like injection_event, session_id, or raw_summary. This leaves meaningful ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Record' with the resource 'context-compaction event' and specifies token counts and re-injected content. It clearly distinguishes from siblings like get_compaction_audit (retrieval) and record_event (generic event) by focusing on compaction-specific auditing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: this is for recording compaction events for audit purposes. It does not explicitly mention alternatives or when not to use it, but the specialized language makes the use case evident. There are no exclusions or alternative references, but the context is strong enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_correctionC

Record a user correction to Claude's output (high-priority learning signal)

ParametersJSON Schema
NameRequiredDescriptionDefault
reasoningNoInferred reason for the correction
session_idYes
claude_actionYesWhat Claude did (tool, file, content)
user_correctionYesHow the user corrected it

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states the action is a 'high-priority learning signal', which hints at importance but does not describe side effects, persistence, reversibility, permissions, or any consequences of invoking the tool. This is a significant gap for a mutation-like recording tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, grammatically complete sentence of about ten words. It is front-loaded with the core verb and resource, and the parenthetical adds context without redundancy. Every word earns its place; there is no fluff or tail-heavy content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool involves nested object parameters, no annotations, and no output schema, yet the description stays at a high level. It does not explain how to structure claude_action or user_correction, what counts as a valid correction, or what the tool returns. This leaves an agent under-informed for correct invocation, especially given the complexity of the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, and the description adds no additional meaning beyond what the schema already provides. The phrase 'user correction to Claude's output' loosely maps to claude_action and user_correction but does not clarify their structure or relationships. The 'reasoning' and 'session_id' parameters are not addressed at all in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Record') and identifies the resource ('a user correction to Claude's output'), which clearly conveys the tool's function. The parenthetical '(high-priority learning signal)' adds useful context. However, it does not explicitly differentiate from sibling tools like record_event or record_decision, though 'correction' provides some distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool should be used when a user corrects Claude's output, but it gives no explicit 'when to use' vs. alternatives, no exclusions, and no prerequisites. Sibling tools with overlapping purposes (e.g., record_event, record_decision) are not referenced, leaving the agent without clear selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_decisionC

Record a decision trace: what the agent proposed and how the human responded

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathNo
reasoningNo
tool_nameNo
session_idYes
decision_typeYes
agent_proposalNo
human_correctionNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the full burden of behavioral disclosure. It explains what is recorded but not whether the operation is append-only, idempotent, permission-sensitive, or what happens on conflict. No mutation or safety details are given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that is front-loaded with the core action and object. No filler or redundant wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters including nested objects and required fields, the description is far too minimal. It leaves the agent without guidance on required inputs, the decision_type enum, or the structure of nested objects, and there is no output schema to aid expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description only hints at 'proposed' and 'human responded', which loosely map to agent_proposal and human_correction. It does not explain required parameters like session_id or decision_type, nor the enum values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool's purpose: recording a decision trace with the agent's proposal and human's response. The verb 'record' and resource 'decision trace' are specific, though it doesn't explicitly distinguish from siblings like record_correction or record_event.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided for when to use this tool versus the many sibling tools (e.g., record_correction, record_event). The description does not mention contexts, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_eventC

Record a development event (file edit, test run, etc.)

ParametersJSON Schema
NameRequiredDescriptionDefault
successNo
entitiesNoEntity names/paths involved
evidenceNoTool inputs/outputs, file contents, etc.
reasoningNo
event_typeYes
session_idYes
descriptionYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It only says 'Record' implying a write operation, but doesn't explain persistence, idempotency, success/failure effects, or whether it appends to a log. This is insufficient for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short sentence, which is concise, but it under-specifies the tool's behavior. While brevity is good, the sentence doesn't earn its place by providing necessary context, making it closer to under-specification than effective conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 7 parameters, no output schema, and no annotations, the description is far from complete. It doesn't explain required fields, return values, or how the event data is used. The description is only a high-level purpose statement, inadequate for safe and correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 29%, so the description must compensate for missing parameter details. It adds examples for event_type ('file edit, test run'), which is redundant with the enum, but it doesn't explain session_id, description, success, or reasoning. The description adds minimal semantic value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the general action ('Record a development event') with examples, so it's more than a tautology. However, it's vague about what constitutes a 'development event' and doesn't distinguish from sibling record tools like record_test_outcome or record_correction, which likely overlap in purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternative record tools. The description doesn't mention any conditions, prerequisites, or exclusions, leaving the agent without direction for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_test_outcomeC

Record test results and link failures to recent code changes

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes
test_resultsYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It implies a write operation but does not disclose side effects of linking failures, whether it is idempotent, or if it requires an existing session. This lack of behavioral detail leaves the agent uncertain about the tool's impact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler or unnecessary details. It efficiently conveys the main purpose and a secondary linking behavior, making it appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no annotations, no output schema, and a moderately complex input schema with a nested array. The description omits critical information about expected input formats, return behavior, and how the linking works. An agent would likely need additional schema inspection or external knowledge to use this tool confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not reference session_id or test_results at all. It fails to explain what session_id should be or how to structure the test_results array. The schema provides names/types, but the description adds no semantic value, making it insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the core action ('Record test results') and adds a distinguishing feature ('link failures to recent code changes') that separates it from sibling tools like record_event or record_decision. It is not a mere tautology because it specifies the linking behavior, though it could be more explicit about the scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool vs alternatives. It does not mention prerequisites, exclusions, or contexts where other record tools would be more appropriate. The agent must infer usage from the name and minimal description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_contradictionC

Pick a winner between two contradicting facts using a confidence-weighted strategy (auto, keep_higher_confidence, keep_most_recent, keep_most_sources, supersede_a, supersede_b, manual).

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNo
strategyNo
fact_a_idYes
fact_b_idYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of disclosing behavioral traits. It only mentions the strategy selection and does not disclose side effects, persistence, reversibility, or what happens to the losing fact. This is a material gap for a tool that presumably modifies state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that immediately states the primary action and enumerates strategies in a parenthetical list. Every element serves a purpose, with no fluff or redundant repetition of schema information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations, output schema, and any behavioral details, the description is far from complete. It fails to explain return values, the outcome of the resolution (e.g., which fact is updated), or any prerequisites. This is a mutation-like operation with substantial missing context, making the tool risky to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides no descriptions (0% coverage), so the description must compensate. It lists possible strategy values, which adds some meaning, but it does not explain the semantics of each strategy (e.g., what 'auto' does) or describe the 'fact_a_id', 'fact_b_id', and 'notes' parameters beyond their schema names. The compensation is partial and insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Pick a winner' and identifies the resource as 'two contradicting facts,' clearly indicating the tool's purpose. It differentiates from sibling tools like find_contradictions by focusing on resolution rather than detection, though it doesn't explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a contradiction exists between two facts and provides a list of strategies, which serves as guidance on how to resolve. However, it lacks explicit exclusions or directives about when not to use this tool versus alternatives, leaving some ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_globalC

Search entities across all registered world-model projects

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavioral traits. It does not mention whether the operation is read-only, how results are returned, or any side effects, leaving significant ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no redundant words. It is efficiently phrased, though it lacks additional structured information that a more complete description might include.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with no output schema and no annotations, the description is minimally sufficient but incomplete. It omits return format, pagination behavior, and the meaning of 'entities', leaving the agent without critical execution context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema includes 'query' and 'limit' with no descriptions, and the description adds no parameter-specific information. It does not explain what query syntax is expected or how limit affects results, failing to compensate for the 0% schema description coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (search), the resource (entities), and the scope (across all registered world-model projects). It distinguishes itself from potential siblings by emphasizing the global scope, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like query_fact. The description only states what it does, leaving it to the agent to infer appropriate usage without any exclusions or comparative context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

seed_projectA

Scan the project codebase and populate the knowledge graph with entities and relationships from existing code

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoRe-seed already processed files
project_dirNoProject directory path (defaults to current)

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states that the tool populates the knowledge graph but does not disclose whether the operation is idempotent, whether it modifies existing data, or any side effects. The 'force' parameter hints at re-seeding but this behavior is not explained in the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, direct, and front-loaded with the main action. No fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema or annotations, the description is the sole source of behavioral context. It covers the high-level operation but lacks guidance on prerequisites, idempotency, return values, or potential side effects of running a mutating operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters (force and project_dir), so the description does not need to compensate. The description itself adds no parameter semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Scan the project codebase and populate the knowledge graph with entities and relationships from existing code'. It uses a specific verb (scan/populate) and resource (codebase, knowledge graph), distinguishing it from siblings like query_fact (query) or record_event (record).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the tool name 'seed_project' (initial population), but the description provides no explicit when-to-use guidance or alternatives. No mention of when to run this versus other ingestion tools like ingest_pr_reviews or record_event.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

simulate_changeC

Project blast radius and historical outcomes for a proposed change

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes
change_descriptionYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states what the tool projects but does not indicate whether it is read-only, whether it requires specific permissions, or what outputs to expect. The lack of any side-effect or limitation details is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that front-loads the core action and object. Every word contributes meaning, and there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description should explain what the tool returns or how the output should be interpreted. It does not, and it also omits any caveats or prerequisites. For a simulation tool, this leaves the agent without enough context to use it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the tool description adds no parameter-level detail. The parameter names 'file_path' and 'change_description' are somewhat self-explanatory, but the description does not clarify expected formats, relationships, or how they map to the 'proposed change' concept, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Project') and resource ('blast radius and historical outcomes for a proposed change'), making the purpose clear. It does not explicitly distinguish from sibling tools like 'predict_regression' or 'validate_change', but the focus on blast radius and historical outcomes is distinctive enough for a 4.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. There is no mention of conditions like 'use when you need to assess impact before applying a change' or references to sibling tools, leaving the agent to infer usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_changeB

Validate a proposed code change against known constraints

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes
change_typeYes
proposed_contentYesThe new content to validate

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It only states the validation intent but doesn't explain whether validation is read-only, what happens on failure, whether it modifies anything, or what 'known constraints' refers to. This lack of transparency for a tool that could have side effects is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short sentence that gets to the point. However, at 8 words, it is extremely terse and could have expanded to include usage context without becoming verbose. It is efficient but slightly under-specified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema and no annotations, the description should explain what the validation result looks like and when to use the tool. It only provides the core purpose, missing critical context about return values, behavior on constraint violation, and relationship to other constraint-related tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers 33% of parameters with descriptions (only proposed_content). The tool description does not mention file_path, change_type, or explain the enum values. It adds no semantic value beyond the schema, leaving file_path and change_type under-specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'validate' with a clear object 'proposed code change' and scope 'against known constraints,' which distinguishes it from sibling tools like simulate_change (which implies running a simulation). It clearly states what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used for checking changes against constraints, but it does not explicitly state when to use it over simulate_change or how it relates to get_constraints. No alternatives are named, and no when-not-to-use guidance is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_retrievalA

Adversarially verify an answer is grounded in a specific set of facts. An independent Coach LLM call checks each material claim in the answer against the supplied source facts and returns confidence (HIGH / MEDIUM / LOW), verified + unverified claim lists, and per-claim source_pointers. Never raises; failures return LOW + error populated. v0.12.12.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesThe user query the answer responds to
answerYesThe candidate answer under verification
fact_idsYesIDs of facts the caller believes ground the answer. Missing IDs are silently dropped.
verification_modelNoOptional Coach model override. Defaults to config.verification_model (Haiku 4.5).

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses the independent Coach LLM call, the confidence levels (HIGH/MEDIUM/LOW), the verified and unverified claim lists, per-claim source_pointers, and that it never raises (failures return LOW with error populated). This is excellent behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the core purpose, followed by behavior and return details. The version tag 'v0.12.12' adds minor noise but does not detract significantly. Overall, every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Since there is no output schema, the description fully explains return values (confidence, claim lists, source_pointers) and error behavior. It is complete for a verification tool, covering what the agent needs to know to invoke it correctly and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description does not add parameter-specific semantics beyond what the schema already provides; it references 'supplied source facts' but leaves parameter details to the schema. This is adequate but not additive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Adversarially verify an answer is grounded in a specific set of facts.' It clearly distinguishes the tool from sibling tools like query_fact or validate_change by emphasizing adversarial verification against supplied source facts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use this tool: whenever an answer needs to be checked against a set of facts. It implies the use case without explicit exclusion or alternative reference, but the context is strong enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.15.5
    • Addedpin_annotation
  2. 2 tool updatesv0.13.0
    • Addedget_audit_log_head
    • Addedprove_entry_inclusion
  3. 1 tool updatev0.12.13
    • Addedverify_retrieval
  4. 1 tool updatev0.12.0
    • Changedquery_fact1 field changed
      • addedInput schema / properties / content_type
        Added value: +{
        +  "description": "Optional filter by content_type. Use 'procedure' to explicitly summon procedures (which are excluded from auto-injection by design).",
        +  "enum": [
        +    "rule",
        +    "fact",
        +    "procedure"
        +  ],
        +  "type": "string"
        +}
  5. 1 tool updatev0.7.4
    • Addedget_agents_md_constraints
  6. 26 tool updatesv0.7.3
    • First observedexport_claude_md
    • First observedfind_contradictions
    • First observedget_co_edit_suggestions
    • First observedget_compaction_audit
    • First observedget_constraints
    • First observedget_context_for_action
    • First observedget_decision_log
    • First observedget_health_report
    • First observedget_injection_context
    • First observedget_related_bugs
    • First observedingest_pr_reviews
    • First observedpredict_regression
    • First observedpredict_test_failures
    • First observedpromote_constraint
    • First observedquery_fact
    • First observedrecall_transcript_range
    • First observedrecord_compaction_audit
    • First observedrecord_correction
    • First observedrecord_decision
    • First observedrecord_event
    • First observedrecord_test_outcome
    • First observedresolve_contradiction
    • First observedsearch_global
    • First observedseed_project
    • First observedsimulate_change
    • First observedvalidate_change

TDQS

C2.9/5.0

Scored across 31 tools

Disambiguation2/5

Several tools have poorly separated boundaries: validate_change, simulate_change, predict_regression, and predict_test_failures all appear to assess the impact of a proposed change, differing mainly in subtle emphasis. Similarly, query_fact, search_global, and get_context_for_action overlap heavily in fact retrieval, and record_correction, record_decision, and pin_annotation all capture human feedback. An agent would frequently struggle to select the right tool without reading every description.

Naming Consistency5/5

Every tool follows a consistent snake_case verb_noun pattern, e.g., get_constraints, record_event, predict_regression, prove_entry_inclusion. Even the more unusual names like pin_annotation and seed_project fit the same imperative structure. This is a highly predictable and uniform naming convention.

Tool Count2/5

31 tools is well beyond the 25+ threshold for a single server and will overwhelm tool selection, especially given the many overlapping prediction and retrieval tools. The server would be more coherent with roughly half the current surface area, consolidating related reads and writes into broader commands.

Completeness4/5

The domain is broadly covered: it supports knowledge-graph population and queries, event and decision recording, constraint ingestion and validation, regression prediction, audit-log integrity, compaction auditing, and context export. Minor gaps exist, such as no explicit fact/constraint update or delete lifecycle and no project listing tool, but the core workflows are well supported.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    A temporal knowledge graph system that enables users to record and query architectural decisions, implementation patterns, and project failures. It integrates with Claude to provide hybrid search, timeline tracking, and automated knowledge gap detection using graph analysis.
    4
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides AI coding agents with pre-edit situational awareness by combining structural call graphs and co-change history to prevent incomplete edits. It surfaces files that historically change together, reducing missed coupled modules.
    3
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables coding agents to query a local, versioned knowledge graph of a software project, retrieving overviews, context packs, evidence, and explanations to make informed changes.
    MIT