CogZ
CogZ
Local-first, code-aware engineering cognition for AI coding agents.
CogZ gives a coding agent persistent memory, contextual retrieval, and continuous cognition about a software repository — all running locally on your machine, no cloud services required.
Works with Claude Code, Cursor, Codex, Gemini CLI, GitHub Copilot, Devin, and any MCP-compatible agent.
What it looks like
Real output from CogZ running on its own codebase:
$ cogz context --mode task "token budget estimation and context pack compression"
Context pack (mode: task)
Query: token budget estimation and context pack compression
Search mode: hybrid
Sections: 91
Token estimate: 8188
Dropped: 6 sections over token budget
---
## 1. [rule] New expansion channels: emit early, filter before seen-mark, sort deterministically, never displace directs (relevance: 0.5148)
Conventions proven across the sibling and co-change channels:
1. Emit before the generic expansion loops — candidates emitted
later get claimed-and-floored by graph traversal …
2. Apply entity-type/test filters BEFORE `seen.insert` …
…
## 2. [rule] cfg-gated code must be typechecked per-target before release (relevance: 0.4690)
Code behind #[cfg(unix)]/cfg(target_os = ...) is invisible to host
builds, tests, and clippy — a compile error in a cfg'd branch ships
silently until a real target build sees it. The v0.5.0 Windows leg
failure is the canonical example.
## 3. [rule] Degradation must be loud, never silent (relevance: 0.3680)
Every degraded or failed code path must surface a signal …
## 4. [identity] CogZ (relevance: —)
Project: CogZ
## 5. [file] assemble.rs (relevance: 0.6993)
//! Context pack assembly — the tiered-push pipeline.
//! Tier 0 (baseline: identity + top rules) always ships for task and
//! escalation packs …
… 86 more sections …That's not a text chunk from a vector search. The pack leads with validated rules — one learned from a release failure on this very project — plus the identity baseline and the actual source file, all ranked, traceable, and budgeted.
This repository already contains real dogfooding knowledge — CogZ has been used on its own codebase throughout development. You can clone it, install CogZ, and try the commands above against it directly.
Related MCP server: Serena
What it does
CogZ maintains a project-specific knowledge layer that connects what an agent learns to the code it is working with.
Memory
CogZ stores three kinds of project knowledge:
Observations — things an agent has learned or noticed. Raw, unvalidated experience: bugs found, decisions made, patterns noticed.
Rules — validated knowledge that should influence future work. Coding standards, design decisions, confirmed patterns.
Knowledge — structured information about the codebase. Architecture explanations, module responsibilities, trade-off rationale.
These are stored as Markdown files with YAML frontmatter, linked to each other and to code entities in the repository. The files are the canonical source of truth — SQLite is a derived index, disposable and rebuildable. Your knowledge is portable, version-controlled, and editable by hand.
Context
Instead of giving an agent everything it knows, CogZ builds scoped context packs for the current situation. A context pack combines relevant rules, observations, knowledge, and code structures — ranked by relevance, traceable through the code graph, and limited by a token budget so the agent gets what matters for the task rather than the entire project history.
Cognition
CogZ periodically consolidates what has been learned: deduplicates entries, detects contradictions, promotes well-supported observations to rules, merges superseded entries, and flags knowledge as stale when the code it references changes.
Quick start
Linux / macOS / Windows (Git Bash):
# Install
curl -fsSL https://raw.githubusercontent.com/balaianu/CogZ/master/install.sh | bash
# Initialize in a repo (add --configure auto to wire MCP + hooks for detected agents)
cd ~/your-project
cogz init
# Index (downloads models on first run, or use --no-download for FTS-only)
cogz index
# Verify it's working — entity counts, model status, DB stats
cogz statusWindows (PowerShell):
# Install
irm https://raw.githubusercontent.com/balaianu/CogZ/master/install.ps1 | iex
# Initialize in a repo
cd your-project
cogz init
cogz indexSee Getting Started for the mental model and a complete walkthrough.
MCP integration
CogZ runs as a stateless MCP server over stdio. Every tool call specifies which repo it targets via a required repo parameter — no Roots, no session state, no fallbacks.
{
"mcpServers": {
"cogz": {
"command": "cogz",
"args": ["mcp-stdio"]
}
}
}The server exposes 15 tools: create_entity, update_knowledge, verify_knowledge, reject_entity, query_entities, search, get_context, get_status, list_entities, consolidate, capture_event, get_callers, get_impact, find_orphans, suggest_observations.
See MCP Tools for full parameter reference and example responses. See Agent Setup for per-agent config files, hook formats, and verified capability notes for all six supported agents — or just run cogz configure auto.
Hook integration
Hooks capture lifecycle events and inject context packs into agent sessions. CogZ's binary is the hook handler — no wrapper scripts needed.
{
"hooks": {
"SessionStart": [{
"matcher": "",
"hooks": [{
"type": "command",
"command": "cogz capture-event session_start --hook-json",
"timeout": 15
}]
}]
}
}See Hooks for all 7 event types and per-agent wiring guides.
CLI commands
Normal operation is automatic: hooks fire on lifecycle events, the agent drives CogZ through MCP. The CLI is not needed for day-to-day use — it's available for setup, manual exploration, and automation if you want or need it.
Command | Description |
| Initialize |
| Write agent MCP + hook config ( |
| Sync files to DB + index source code |
| Incremental reindex (changed files only) |
| Hybrid FTS + vector + graph search |
| Assemble context pack |
| DB stats, entity counts, model status |
| Run promotion and merge |
| List mined observation candidates |
| Re-stamp a drifted entity's provenance |
| Mark an entity rejected ( |
| Capture lifecycle event from hooks |
| Model management |
| Health check, policy violations, usage metrics |
| Self-update from GitHub releases |
| Drop DB (optionally purge observations) |
| Run MCP server over stdio |
See CLI Reference for all flags and options.
Requirements
Minimum (FTS-only mode)
Resource | Requirement |
RAM | 256 MB free |
Disk | 50 MB (binary + DB, no models) |
CPU | any x86_64 or ARM64 |
Works without ONNX Runtime or model downloads. All hooks, FTS search, context packs, consolidation, doctor, and prune are functional. Vector search, embedding-based dedup, and contradiction detection are not available.
Recommended (hybrid search mode)
Resource | Requirement |
RAM | 2 GB free |
Disk | 550 MB (binary + ONNX Runtime + 3 models + DB) |
CPU | any x86_64 or ARM64, 4+ cores speeds up batch embedding |
Full functionality including vector search, semantic dedup, and NLI contradiction detection. Models auto-download on first use and auto-unload after 5 min idle (RAM drops back to ~11 MB). See Evaluations for the full resource consumption profile.
Benchmarks
CogZ ships a reproducible suite (benchmark/) run on pinned public corpora — httpx, cobra, clap, each injected with memory seeds mined from its real git history — plus this repository's own .cogz corpus. Seeded ground truth:
Corpus | P@5 | MRR | Recall@20 |
cobra | 0.200 | 0.531 | 0.967 |
httpx | 0.173 | 0.358 | 0.917 |
clap | 0.185 | 0.278 | 0.839 |
Channel ablations on commit queries: removing graph expansion costs 10–16pt recall@20 on every corpus; FTS-only mode retains ~75–85% of hybrid recall with ~745 MB less RSS. Context packs keep 0.70–0.90 expected-entity recall at the default 8K budget. Reruns are byte-identical. Full methodology, per-phase numbers, and the raw artifacts: benchmark/README.md.
What using it buys (measured): in a 14-task agent replay, the seeded-knowledge arm finished ~2x faster than bare (871s vs 1748s average) and completed more runs (14/14 vs 10/14) at equal correctness. Consolidation machinery is precise: dedup precision/recall 1.0, NLI contradiction detection 4/4 with zero false alarms, drift marking exact.
Honest limits: top-5 precision is weak on mixed corpora (P@5 <= 0.20; code entities outrank knowledge at the top of the ranking), commit-intent queries reach 0.36–0.56 recall@20, adjacent-domain negative queries leak confident hits (silence-gate clean rate 0–0.4 across corpora), and at n=14 tasks there is no measurable task-correctness lift yet.
Architecture
Single Rust binary — no runtime dependencies except optional ONNX models for vector search.
Files are canonical — all entities are Markdown files. The SQLite DB is a derived index, disposable and rebuildable.
Code-aware — tree-sitter indexes source code as first-class graph entities. Supported languages: Rust, Python, Go, JavaScript, TypeScript, TSX, Bash.
Graceful degradation — works without ML models in FTS-only mode.
Local-first — no cloud, no telemetry, no accounts. The only network access is optional model downloads.
See Architecture for the full system design.
Compatibility
Platform | Support | Embeddings | FTS-only | Install |
Linux x86_64 | Full | Auto-download | Yes |
|
Linux aarch64 | Full | Auto-download | Yes |
|
macOS arm64 (Apple Silicon) | Full | Auto-download | Yes |
|
macOS x86_64 (Intel) | Not supported | — | — | — |
Windows x86_64 | Full | Auto-download | Yes |
|
macOS Intel is not supported because Microsoft dropped ONNX Runtime macOS Intel binaries after v1.22. Intel Mac users can run the arm64 binary under Rosetta 2 (with a compatible ORT build) or use cargo install cogz for FTS-only mode.
Windows 10+ is required (bsdtar is bundled since build 17063, needed for ONNX Runtime auto-extraction).
Cross-platform team collaboration is supported: code entity UUIDs use forward-slash path normalization so the same source file produces the same entity ID on all platforms.
Documentation
User guides:
Getting Started — mental model and walkthrough
Configuration — full
config.tomlreferenceCLI Reference — every command and flag
Integration:
MCP Tools — all 15 tool signatures and response shapes
Hooks — lifecycle events and output format
Agent Setup — all six agents + generic MCP, with per-agent effect coverage
Design:
Architecture — system overview and module map
Entity Model — entity types, frontmatter, state machine
Search — hybrid FTS + vector, RRF, graph expansion
Consolidation — dedup, contradiction, promotion, merge
Degradation — FTS-only mode and fallback behavior
Contributing:
Building — build, release, cross-compile
Testing — test categories and mock models
Conventions — code patterns and invariants
Dependencies — pinned versions and supply-chain policy
Schema — DB schema and migrations
Contributing
See CONTRIBUTING.md for build, test, and PR guidelines.
License
MIT — see LICENSE.
Support
If you find this tool useful, consider buying me a coffee:
Available Tools
15 toolscapture_eventA
Capture a lifecycle event — called by hook scripts, not intended for direct use. session_start/prompt_submit return context packs; file_save reindexes and may return rules governing the edited file; session_end runs consolidation.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | Absolute path to the project root containing `.cogz/`. | |
| prompt | No | Prompt text (for prompt_submit). | |
| file_path | No | Saved file path, relative to repo root (for file_save). | |
| tool_name | No | Tool name (for pre_tool_use, post_tool_use). | |
| event_type | Yes | Event type: session_start, prompt_submit, pre_tool_use, post_tool_use, file_save, session_end. | |
| tool_result | No | Tool result summary (for post_tool_use). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does disclose side effects and return behavior: 'session_start/prompt_submit return context packs', 'file_save reindexes and may return rules', 'session_end runs consolidation'. It omits any behavior for pre_tool_use/post_tool_use, leaving those two event types opaque.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with the caller/scope constraint front-loaded and zero filler; every clause conveys distinct information about invocation or per-event behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no annotations and no output schema, the description usefully explains the return behavior of most events, but two of the six event types (pre_tool_use, post_tool_use) get no behavioral treatment, leaving a gap an agent cannot close from structured fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by mapping event_type values to distinct behaviors (context packs, reindexing, consolidation) that the raw enum listing does not convey. The other five parameters are covered by their own schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Capture a lifecycle event') and immediately scopes who invokes it ('called by hook scripts, not intended for direct use'). An agent can distinguish this event-ingest tool from sibling query/entity tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit when-not ('not intended for direct use') and identifies the actual caller (hook scripts), which is exactly the routing information an agent needs to avoid invoking it manually. Context for each event type is also implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
consolidateA
Run deferred consolidation — promote supported observations to rules, merge confirmed duplicates. Housekeeping, not a write path; dedup and contradiction checks already run on every insert.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | Absolute path to the project root containing `.cogz/`. | |
| dry_run | No | If true, report what would be consolidated without making changes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does useful work: it names the two mutating effects and clarifies this is deferred housekeeping rather than the ingest write path. But it never states whether merged duplicates are destroyed irreversibly, whether the run is idempotent, or what permissions the repo path requires — significant omissions for a tool that promotes and merges stored knowledge.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, both earning their place, with the core action front-loaded before the scoping caveat. No filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description is the sole source of behavioral context. It covers purpose and the not-a-write-path distinction adequately, but leaves irreversibility, idempotency, and the trigger condition unspecified — meaningful gaps for a mutation tool with a dry_run escape hatch.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both `repo` (absolute path to project root containing `.cogz/`) and `dry_run` are already fully documented in the schema. The description adds no parameter-level syntax or constraints beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (deferred consolidation) and enumerates its two concrete effects: promoting supported observations to rules and merging confirmed duplicates. The clause 'Housekeeping, not a write path' implicitly separates it from mutation siblings like create_entity or update_knowledge, but no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The note that 'dedup and contradiction checks already run on every insert' tells the agent this tool is not for routine dedup, which implicitly scopes usage. However, it never states when consolidation should actually be triggered (e.g., after N inserts, on a schedule, or when stale rules surface), so the trigger condition is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_entityA
Create a knowledge-layer entity. entity_type picks the lifecycle, not the topic — ask what the entry IS: 'observation' = something that happened (bug found, surprising behavior, decision noticed) — raw, append-only, unvalidated; consolidation promotes the good ones to rules. 'rule' = a verified directive agents must always follow (conventions, constraints) — pushed into every context pack; change via supersede, not edits. 'knowledge' = a curated reference doc (architecture, gotchas, design decisions) — the only editable type (update_knowledge). Requires: content always; title+category for knowledge.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | Absolute path to the project root containing `.cogz/`. | |
| tags | No | `knowledge` only. | |
| title | No | Title. Required for `knowledge`; auto-generated for observation/rule when omitted. | |
| source | No | `observation` only: who or what produced it. Default: "agent". | |
| content | Yes | Entity body. Required for all types. | |
| category | No | Category — required for `knowledge` (e.g. architecture, decisions, gotchas). Ignored otherwise. | |
| confidence | No | `rule` only: confidence 0..1. | |
| references | No | UUIDs of entities this entry references (code or knowledge). | |
| entity_type | Yes | Which lifecycle class to create: `observation` (raw finding, append-only, unvalidated), `rule` (verified directive, always delivered in packs, supersede to change), or `knowledge` (curated reference doc, editable via update_knowledge). | |
| supporting_ids | No | `observation` only: UUIDs of observations this one supports — creates `supports` edges for promotion consolidation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it does substantial work: observation is append-only/unvalidated with consolidation promoting good ones, rules are pushed into every context pack and changed via supersede rather than edits, knowledge is the only editable type. It omits permissions/auth needs and any notion of the create response, but the lifecycle consequences are unusually well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then organized by entity_type with em-dash clauses that keep related facts together. It is information-dense and slightly long, but every clause (lifecycle, editability, promotion) earns its place, so no significant waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter create tool with no output schema and no annotations, the description covers the behavioral model thoroughly and summarizes the required fields. Return-value details are unnecessary given no output schema, though the absence of any auth/permission note leaves a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 10 parameters; the description only restates the requirements ("content always; title+category for knowledge"). That summary is useful but adds little beyond what the schema fields already say, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Create a knowledge-layer entity") and then disambiguates the lifecycle meaning of entity_type far beyond a tautology, explicitly noting that entity_type selects the lifecycle, not the topic. It distinguishes the three resulting lifecycles clearly enough that an agent can tell what kind of object it is producing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong per-type usage guidance ('observation' = raw append-only, 'rule' = verified directive, 'knowledge' = editable doc) and routes the agent to siblings — 'change via supersede, not edits' and 'the only editable type (update_knowledge)'. It lacks an explicit when-not-to-create statement, so it falls short of the top band, but the conditions for each branch are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_orphansA
Find code entities with no incoming calls/imports/extends edges — dead-code candidates. Use during cleanup audits. Entry points like main() surface by design; test code is excluded.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | Absolute path to the project root containing `.cogz/`. | |
| limit | No | ||
| entity_type | No | Code entity type to check: function, class, file, or module. Default: functions and classes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does real work: it warns that entry points like main() surface by design (a known false-positive class) and that test code is excluded. It does not state read-only nature, performance characteristics, or result ordering, but the two disclosed caveats are exactly the traps an agent would otherwise mis-handle.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core definition, followed by usage context and two caveats. No filler, no restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analysis tool with no output schema, the description covers purpose, when-to-use, and two important caveats (entry points, test exclusion). It stops short of describing what a result item contains or how to disambiguate entity_type/limit, but that is a modest gap rather than a blocking one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: `repo` and `entity_type` are documented in the schema, while `limit` has no description anywhere. The description adds the edge-relation vocabulary (calls/imports/extends) and confirms the default entity scope, but does nothing to clarify `limit` or how entity_type values map to results, so it only marginally exceeds the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('find code entities') and defines the exact condition that qualifies one ('no incoming calls/imports/extends edges'), which is genuinely distinguishing information an agent cannot get from the name alone. It is clearly separable from siblings like list_entities, query_entities, and get_callers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use context ('use during cleanup audits'), which is more than most siblings offer, but names no explicit alternative or exclusion condition (e.g., when to prefer get_callers or query_entities instead).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_callersA
Find entities that call this function — answers 'who calls X?' with resolved call edges, not text matches. Use instead of grepping for the name when you need real callers (not comments, strings, or same-named functions).
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | Absolute path to the project root containing `.cogz/`. | |
| limit | No | ||
| entity_id | Yes | Entity ID whose callers to find (function, method). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose a meaningful behavioral trait — results are resolved call edges, not text matches — which sets expectations about precision and false positives. However, it says nothing about permissions, result ordering, pagination, or what happens when no callers exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with zero waste, front-loading what the tool returns and following immediately with the when-to-use contrast against grep.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description must cover behavior and it does so at a conceptual level (call edges, not text matches). It stops short of describing the shape or ordering of returned caller records or how limit interacts with results, leaving a small but real gap for a tool with a bare schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: repo and entity_id are documented in-schema, while limit has no description at all. The description adds no parameter-level detail (e.g. that entity_id must be a function/method ID) beyond what the schema already states, so it neither compensates for the limit gap nor adds value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Find entities that call this function') and frames it as answering 'who calls X?' with resolved call edges rather than text matches. This distinguishes it from grep-style search siblings without needing to open any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use it 'instead of grepping for the name' and gives the disqualifying conditions (comments, strings, same-named functions). It names an alternative approach but not the sibling tools (e.g. search, query_entities) an agent might otherwise pick, so routing is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_contextA
Assemble a context pack — a scoped, ranked bundle of orientation, rules, relevant code, and knowledge for a task query. Use at the start of substantial work on a topic instead of reading files broadly — narrower than a session-start pack, broader than a single search.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| repo | Yes | Absolute path to the project root containing `.cogz/`. | |
| query | No | ||
| max_tokens | No | ||
| include_stale | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no output schema, the description carries the full burden. It does disclose useful behavior: the result is a scoped, ranked bundle with a known content mix, which tells the agent what it will get back. However, it says nothing about cost, token budgeting, staleness handling, or failure modes (e.g. missing `.cogz/`), all of which are live concerns given the max_tokens/include_stale parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, zero filler, with the core purpose front-loaded and the routing guidance second. Every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter read-style tool with no output schema, the description adequately covers purpose, output content, and when to reach for it. It is incomplete only on parameter behavior (mode, max_tokens, include_stale), which is left undocumented everywhere, a notable but bounded gap rather than a missing core.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% (repo alone is documented), so the description must compensate for mode, query, max_tokens, and include_stale — and it does not mention any of them. "For a task query" faintly gestures at the query parameter and "scoped" at max_tokens, but there is no explanation of modes or what include_stale changes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ("Assemble") plus a clearly defined resource ("context pack") whose contents are enumerated: orientation, rules, relevant code, and knowledge. It also positions itself against adjacent behaviors ("reading files broadly", "a single search"), so an agent can distinguish it from siblings like search or list_entities without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ("at the start of substantial work on a topic") plus named alternatives it replaces ("instead of reading files broadly") and a scope calibration against both a broader pack and a single search. The only minor gap is no explicit when-not-to-use for trivial queries, but "substantial work" implies that boundary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_impactA
Transitive dependents of an entity — what breaks or needs updating when it changes (incoming calls/imports/extends up to max_depth hops), plus knowledge that references it. Use before renaming, deleting, or changing a signature.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | Absolute path to the project root containing `.cogz/`. | |
| limit | No | ||
| entity_id | Yes | Entity ID to analyze. | |
| max_depth | No | Max dependency hops to traverse (default 2, capped at 4). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden; it does disclose the traversal semantics (transitive, up to max_depth hops, incoming edges) and that results include knowledge references. However, it omits the depth cap of 4 (only in the schema) and says nothing about truncation, cost of a deep traversal, or the effect of the undocumented limit parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core output description front-loaded and the usage trigger second; no filler, and both ideas earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations and no output schema, so the description must frame the return payload; it does so conceptually (dependents plus referencing knowledge). It is slightly incomplete on result shape, depth-cap truncation, and the mystery 'limit' parameter for a 4-parameter analysis tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% and the description adds real meaning to max_depth ('dependency hops to traverse'), which the schema only partially captures. But the 'limit' parameter is undocumented in both the schema and the description, so the description does not fully compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact resource (transitive dependents of an entity) and the concrete edge types traversed (incoming calls/imports/extends), plus a second payload (knowledge referencing it). The word 'transitive' implicitly separates it from the direct-caller sibling get_callers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use before renaming, deleting, or changing a signature' gives an explicit triggering context for the tool. It stops short of naming an alternative or a when-not condition, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_statusA
CogZ system status — entity counts, DB stats, model availability, staleness. Use to check the index is fresh and retrieval is at full capability before relying on it.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | Absolute path to the project root containing `.cogz/`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the diagnostic dimensions (staleness, model availability) but never states that this is a non-mutating read, whether it is cheap or expensive to call, or what happens when the repo path is invalid — gaps that matter more given zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the scope of returned data front-loaded before the usage cue. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating what the status report contains, which is what an agent needs to interpret the result. It is nearly complete for a low-risk diagnostic tool, missing only edge-case behavior (e.g., invalid or unindexed repo).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with a single parameter, and the schema already explains that `repo` is the absolute path to the project root containing `.cogz/`. The description adds nothing about the parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (CogZ system status) and enumerates the concrete contents returned: entity counts, DB stats, model availability, and staleness. This is specific enough to distinguish it from the entity/query-oriented siblings, though it does not explicitly contrast itself with any sibling by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use to check the index is fresh and retrieval is at full capability before relying on it" gives a clear triggering condition tied to a real workflow decision. It stops short of naming alternatives or stating when-not to use it, so it is not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_entitiesA
List IDs and titles for all entities of a type. Use to enumerate a type or resolve an entity_id for get_callers/get_impact — returns no content; for substance use search or query_* tools.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | Absolute path to the project root containing `.cogz/`. | |
| status | No | ||
| entity_type | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose the key trait that this returns only IDs and titles with no content, which tells the agent this is a cheap read-only enumeration. However, for a tool that lists ALL entities it says nothing about result size, pagination, or truncation, and no permissions/rate context is given.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the primary purpose followed immediately by routing guidance. Every clause earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully specifies the return shape ('IDs and titles'), covering the biggest gap. But with no annotations and 33% parameter coverage, the meaning of entity_type and status remains undefined in both the schema and the description, leaving the agent under-equipped for a required-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: repo is documented, but entity_type has no description and status has no description, no enum, and a null default, so an agent cannot tell what valid values or filtering semantics apply. The description never mentions any parameter, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List IDs and titles for all entities of a type') and narrows the payload explicitly ('returns no content'), which separates it from the content-returning siblings. An agent can distinguish it from query_entities, search, and get_impact without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use ('enumerate a type or resolve an entity_id for get_callers/get_impact') and an explicit when-not ('for substance use search or query_* tools'), naming the alternative tools by name. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_entitiesA
Browse knowledge-layer entities of one type — 'observation' (raw findings, recency order), 'rule' (verified directives, confidence order), or 'knowledge' (curated docs). Use to enumerate what exists before writing (avoid duplicates) or to review a type — for ranked retrieval on a question use search; for code entities use list_entities or search.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | Absolute path to the project root containing `.cogz/`. | |
| tags | No | `knowledge` only. | |
| limit | No | ||
| status | No | Status filter. Default: active only; `all` for every status. | |
| category | No | `knowledge` only. | |
| references | No | `observation`/`rule` only: filter to entities referencing this target UUID. | |
| entity_type | Yes | `observation` | `rule` | `knowledge`. For code entities use `list_entities` or `search` instead. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the ordering semantics per entity type (recency vs confidence) and that this is a browsing/enumeration operation, but it never explicitly states the operation is read-only/non-destructive, nor does it cover pagination or limit behavior. Adequate but with real gaps for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded: the purpose and type enumeration come first, then the routing guidance. It is a single heavily em-dashed sentence, which packs a lot but stays readable and wastes little.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 7 parameters, no annotations, and no output schema, the description needs to carry usage and behavioral context, and it does cover purpose, routing, and ordering. Remaining gaps (pagination, read-only confirmation) are minor since the schema documents most parameters and no return-value explanation is required.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 86%, so the schema already documents repo, tags, status, category, references, and entity_type. The description adds the ordering meaning of entity_type values but nothing about limit, tags, category, or references beyond what the schema states. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Browse knowledge-layer entities of one type') and then enumerates the three valid types with their ordering semantics (observation=recency, rule=confidence, knowledge=curated docs). This lets an agent distinguish it from search, list_entities, and get_context without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('enumerate what exists before writing (avoid duplicates)' or 'review a type') and explicit when-not with named alternatives ('for ranked retrieval on a question use search; for code entities use list_entities or search'). Routing is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reject_entityA
Reject a knowledge-layer entity — writes status: rejected (and an optional rejected_reason) into its canonical file, then syncs so the status lattice validates the transition. Only active entities can be rejected; verify a stale one first if it must be ruled wrong. This is a verdict, not an edit — do not use it for content changes. Rejected entities stay on record: retrieval filters them out, and dedup can warn when a matching claim resurfaces.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | UUID of the entity to reject. | |
| repo | Yes | Absolute path to the project root containing `.cogz/`. | |
| reason | No | Why the entity is rejected — stored as `rejected_reason` in the file's frontmatter so the verdict carries its evidence. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the state-machine constraint (only active entities), the sync/validation step ('the status lattice validates the transition'), and the lasting consequences (rejected entities stay on record, retrieval filters them out, dedup can warn on resurfacing). This is exactly the behavioral context an agent needs for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core action and state change, followed by preconditions, exclusions, and consequences. Every clause earns its place and nothing is repeated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema or annotations, it covers preconditions, side effects, and persistence semantics thoroughly. It stops short of describing the response payload or the failure mode when the entity is not active, which would round it out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three parameters are already documented, including that reason is stored as 'rejected_reason' in frontmatter. The description largely restates that same fact (reason becomes 'rejected_reason') without adding format, length, or constraints. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Reject a knowledge-layer entity') and immediately distinguishes itself from siblings by contrasting with edits (update_knowledge) and verification, and it names the exact state transition it performs ('writes status: rejected'). An agent can tell this apart from verify_knowledge or update_knowledge without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit preconditions ('Only active entities can be rejected'), a routing rule to an alternative ('verify a stale one first if it must be ruled wrong'), and a clear exclusion ('This is a verdict, not an edit — do not use it for content changes'). When-to-use, when-not, and the alternative path are all present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchA
Ranked search across all entities — code (functions, files, classes) plus rules, observations, and knowledge — using hybrid lexical + semantic retrieval with graph expansion. Use when you know the concept but not the exact name, or want related entities surfaced automatically. For exact identifier/text matches, grep is faster and equally precise.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | Absolute path to the project root containing `.cogz/`. | |
| limit | No | ||
| query | Yes | ||
| before | No | Only use git history committed before this unix timestamp for the co-change channel. | |
| expand | No | ||
| status | No | ||
| code_search | No | Use the code model (CodeRankEmbed) for query embedding. Applies the CodeRankEmbed query prefix for code-focused search. | |
| entity_type | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the disclosure burden. It explains the retrieval strategy (hybrid lexical+semantic, graph expansion), which is useful behavioral context, but says nothing about read-only safety, rate limits, result shape, or pagination/limit behavior for a tool that takes a `limit` parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the core purpose front-loaded and the alternative-usage caveat last. No filler, though the density comes at the cost of parameter coverage rather than excess verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Purpose and usage routing are complete, but with 8 parameters at low coverage, no annotations, and no output schema, the definition leaves an agent unable to determine how to scope results by type, status, or time. Adequate for intent, incomplete for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 38% across 8 parameters. The description alludes to graph expansion (expand) and the code model (code_search) but never explains `before` (git-timestamp co-change scoping), `status`, `entity_type`, or `limit`, leaving several params undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (ranked search) over a specific resource (all entities: code, rules, observations, knowledge) and names the retrieval mechanism (hybrid lexical + semantic with graph expansion). An agent can immediately tell what it returns and how it differs from a plain text matcher.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it ('know the concept but not the exact name', 'want related entities surfaced automatically') and names the competing approach (grep for exact identifier/text matches) with the tradeoff. This is the strongest part of the definition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_observationsA
Mine recent session usage for observation candidates — zero-hit packs followed by edits, hot files, error→fix sequences. Use at natural stopping points to capture what the session learned; confirm salient suggestions via create_entity.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | How far back to mine events and usage, in days (default 7). | |
| repo | Yes | Absolute path to the project root containing `.cogz/`. | |
| limit | No | Max candidates to return (default 10, capped at 50). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose meaningful traits: it mines existing session data rather than persisting anything, outputs suggestions that must be confirmed via create_entity, and lists the heuristics used. It does not mention whether anything is written, rate limits, or the candidate object's shape, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both front-loaded: the first defines the mining behavior with concrete signals, the second gives the usage trigger and the follow-up tool. No filler or restated name/title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description should signal the return shape; it indicates "observation candidates" and "salient suggestions" but does not describe the candidate object or count/pagination behavior explicitly. For a read-and-suggest tool it is largely complete, with a minor gap on return-value detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so days, repo, and limit are fully documented in the schema (including defaults and the 50 cap). The description adds no parameter-level detail beyond what is already structured, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (mine recent session usage for observation candidates) and enumerates the exact signals it looks for (zero-hit packs followed by edits, hot files, error→fix sequences). It also names the sibling tool (create_entity) used to act on results, so an agent can place it in the workflow without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to invoke it ("at natural stopping points to capture what the session learned") and what to do next ("confirm salient suggestions via create_entity"). It lacks an explicit when-not/alternative condition, but the workflow guidance is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_knowledgeA
Update an existing knowledge entry's content — use to correct or extend documentation when facts change. The only entity type allowing in-place edits; observations and rules are append-only.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| repo | Yes | Absolute path to the project root containing `.cogz/`. | |
| tags | No | ||
| title | No | ||
| content | Yes | ||
| category | No | ||
| references | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses one genuinely useful behavioral trait — this is the only entity type supporting in-place edits, while observations and rules are append-only — but says nothing about whether content replaces or merges with existing content, whether null optional fields clear values, or permission requirements for a mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and followed by the discriminating constraint. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation with no annotations, no output schema, and near-zero schema coverage, the description is too thin: it omits parameter behavior, mutation semantics, and any failure/return expectations. An agent would have to guess how the optional fields are applied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 14% (only 'repo' is documented), and the description adds no parameter meaning beyond naming 'content'. Nothing explains how the six other fields (id, title, tags, category, references) behave, nor whether a required 'content' implies wholesale replacement.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Update an existing knowledge entry's content') and immediately differentiates it from sibling entity types by noting observations and rules are append-only, so an agent can tell it apart from create_entity and the other entity tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear usage condition ('use to correct or extend documentation when facts change') and implicitly routes append-only cases elsewhere. It stops short of naming an explicit alternative tool for the append-only case, but the context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_knowledgeA
Re-verify a knowledge entity against its referenced code — re-stamps verified_against provenance, clears drift annotations, and reactivates the entity if it was stale. Use after reading drift-flagged or stale knowledge and confirming it is still accurate — not for changing content (use update_knowledge instead).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | UUID of the knowledge entity to verify. | |
| repo | Yes | Absolute path to the project root containing `.cogz/`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the mutation's side effects (provenance re-stamp, drift-annotation clearing, reactivation of stale entities). It does not cover permissions/auth needs, reversibility, or failure behavior (e.g., what happens if the code still drifted), so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences: the effect chain comes first, then the routing guidance. Zero filler and no repetition of the schema or name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation tool with no annotations and no output schema, the description covers purpose, side effects, and selection criteria adequately. It leaves minor gaps around return values and error/drift-remaining behavior, which keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters (id and repo), so the schema already documents them fully; the description adds no syntax, format, or constraint detail beyond what is structured. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (re-verify) on a specific resource (knowledge entity) and enumerates the concrete effects: re-stamping verified_against provenance, clearing drift annotations, and reactivating stale entities. It is clearly distinguishable from update_knowledge, which it names.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it ('after reading drift-flagged or stale knowledge and confirming it is still accurate') and when not to ('not for changing content'), naming the correct alternative tool. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
15 tool updates
v0.5.6- First observed
capture_event - First observed
consolidate - First observed
create_entity - First observed
find_orphans - First observed
get_callers - First observed
get_context - First observed
get_impact - First observed
get_status - First observed
list_entities - First observed
query_entities - First observed
reject_entity - First observed
search - First observed
suggest_observations - First observed
update_knowledge - First observed
verify_knowledge
TDQS
Scored across 15 tools
Tools are mostly well-differentiated with explicit cross-references ('use update_knowledge instead', 'for code entities use list_entities'). The overlap between list_entities and query_entities is real but the descriptions carefully delineate them (code vs knowledge-layer, IDs only vs content). search vs query_entities vs get_context is a slightly murky trio, but each description states its intended use.
Predominantly verb_noun snake_case (update_knowledge, verify_knowledge, list_entities, query_entities, get_impact, create_entity). A few bare verbs (search, consolidate) and internal/verb-only names (reject_entity, capture_event, suggest_observations) are acceptable and readable. Minor deviation but consistent overall.
15 tools for a knowledge/code-graph memory system with retrieval, lifecycle, write, and maintenance concerns is reasonable. One or two (capture_event explicitly marked 'not intended for direct use') could be hidden from the agent-facing surface, but nothing feels excessive.
Covers the full knowledge lifecycle: create, update, verify, reject, consolidate, plus retrieval (search, query, context, impact, callers, orphans) and status. Gaps are minor — no explicit delete of knowledge-layer entities (rejection/supersede likely intended instead) and no direct supersede tool despite it being referenced.
Maintenance
Related MCP Connectors
Shared memory for coding agents. Stop re-explaining your codebase every session.
Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.
Shared project memory for AI coding agents: decisions, lessons, risks and tasks in one graph.
Local-first long-term memory for AI agents, with byte-recomputable signed verification receipts.
Related MCP Servers
- AlicenseBqualityAmaintenanceBasic Memory is a knowledge management system that allows you to build a persistent semantic graph from conversations with AI assistants. All knowledge is stored in standard Markdown files on your computer, giving you full control and ownership of your data. Integrates directly with Obsidan.md177,385 PyPI4,080AGPL 3.0
- AlicenseAqualityDmaintenanceA coding agent toolkit that provides IDE-like semantic code retrieval and editing tools, enabling LLMs to efficiently navigate and modify codebases at the symbol level rather than working with entire files.29MIT
- AlicenseAqualityAmaintenanceLocal-first memory layer for AI coding agents — captures issues, attempts, fixes, and decisions, and warns at git commit before you repeat a mistake.17173 PyPI850MIT
- AlicenseNot gradedqualityAmaintenanceProvides persistent, local-first memory with knowledge graph and hybrid search for AI coding agents, reducing token usage by storing decisions, patterns, and codebase context.39 PyPI9MIT