CogZ
# CogZ
[](https://github.com/balaianu/CogZ/actions/workflows/ci.yml)
[](https://opensource.org/licenses/MIT)
[](https://www.rust-lang.org/)
[](https://github.com/balaianu/CogZ/releases)
[](https://scorecard.dev/viewer/?uri=github.com/balaianu/CogZ)
[](https://www.bestpractices.dev/projects/15167)
[](https://buymeacoffee.com/balaianu)
Local-first, code-aware engineering cognition for AI coding agents.
CogZ gives a coding agent persistent memory, contextual retrieval, and continuous cognition about a software repository — all running locally on your machine, no cloud services required.
Works with Claude Code, Cursor, Codex, Gemini CLI, GitHub Copilot, Devin, and any MCP-compatible agent.
## What it looks like
Real output from CogZ running on its own codebase:
```
$ cogz context --mode task "token budget estimation and context pack compression"
Context pack (mode: task)
Query: token budget estimation and context pack compression
Search mode: hybrid
Sections: 91
Token estimate: 8188
Dropped: 6 sections over token budget
---
## 1. [rule] New expansion channels: emit early, filter before seen-mark, sort deterministically, never displace directs (relevance: 0.5148)
Conventions proven across the sibling and co-change channels:
1. Emit before the generic expansion loops — candidates emitted
later get claimed-and-floored by graph traversal …
2. Apply entity-type/test filters BEFORE `seen.insert` …
…
## 2. [rule] cfg-gated code must be typechecked per-target before release (relevance: 0.4690)
Code behind #[cfg(unix)]/cfg(target_os = ...) is invisible to host
builds, tests, and clippy — a compile error in a cfg'd branch ships
silently until a real target build sees it. The v0.5.0 Windows leg
failure is the canonical example.
## 3. [rule] Degradation must be loud, never silent (relevance: 0.3680)
Every degraded or failed code path must surface a signal …
## 4. [identity] CogZ (relevance: —)
Project: CogZ
## 5. [file] assemble.rs (relevance: 0.6993)
//! Context pack assembly — the tiered-push pipeline.
//! Tier 0 (baseline: identity + top rules) always ships for task and
//! escalation packs …
… 86 more sections …
```
That's not a text chunk from a vector search. The pack leads with validated rules — one learned from a release failure on this very project — plus the identity baseline and the actual source file, all ranked, traceable, and budgeted.
This repository already contains real dogfooding knowledge — CogZ has been used on its own codebase throughout development. You can clone it, install CogZ, and try the commands above against it directly.
## What it does
CogZ maintains a project-specific knowledge layer that connects what an agent learns to the code it is working with.
**Memory**
CogZ stores three kinds of project knowledge:
- **Observations** — things an agent has learned or noticed. Raw, unvalidated experience: bugs found, decisions made, patterns noticed.
- **Rules** — validated knowledge that should influence future work. Coding standards, design decisions, confirmed patterns.
- **Knowledge** — structured information about the codebase. Architecture explanations, module responsibilities, trade-off rationale.
These are stored as Markdown files with YAML frontmatter, linked to each other and to code entities in the repository. The files are the canonical source of truth — SQLite is a derived index, disposable and rebuildable. Your knowledge is portable, version-controlled, and editable by hand.
**Context**
Instead of giving an agent everything it knows, CogZ builds scoped context packs for the current situation. A context pack combines relevant rules, observations, knowledge, and code structures — ranked by relevance, traceable through the code graph, and limited by a token budget so the agent gets what matters for the task rather than the entire project history.
**Cognition**
CogZ periodically consolidates what has been learned: deduplicates entries, detects contradictions, promotes well-supported observations to rules, merges superseded entries, and flags knowledge as stale when the code it references changes.
## Quick start
**Linux / macOS / Windows (Git Bash):**
```bash
# Install
curl -fsSL https://raw.githubusercontent.com/balaianu/CogZ/master/install.sh | bash
# Initialize in a repo (add --configure auto to wire MCP + hooks for detected agents)
cd ~/your-project
cogz init
# Index (downloads models on first run, or use --no-download for FTS-only)
cogz index
# Verify it's working — entity counts, model status, DB stats
cogz status
```
**Windows (PowerShell):**
```powershell
# Install
irm https://raw.githubusercontent.com/balaianu/CogZ/master/install.ps1 | iex
# Initialize in a repo
cd your-project
cogz init
cogz index
```
See [Getting Started](docs/getting-started.md) for the mental model and a complete walkthrough.
## MCP integration
CogZ runs as a stateless MCP server over stdio. Every tool call specifies which repo it targets via a required `repo` parameter — no Roots, no session state, no fallbacks.
```json
{
"mcpServers": {
"cogz": {
"command": "cogz",
"args": ["mcp-stdio"]
}
}
}
```
The server exposes 15 tools: `create_entity`, `update_knowledge`, `verify_knowledge`, `reject_entity`, `query_entities`, `search`, `get_context`, `get_status`, `list_entities`, `consolidate`, `capture_event`, `get_callers`, `get_impact`, `find_orphans`, `suggest_observations`.
See [MCP Tools](docs/integration/mcp-tools.md) for full parameter reference and example responses. See [Agent Setup](docs/integration/agent-setup.md) for per-agent config files, hook formats, and verified capability notes for all six supported agents — or just run `cogz configure auto`.
## Hook integration
Hooks capture lifecycle events and inject context packs into agent sessions. CogZ's binary is the hook handler — no wrapper scripts needed.
```json
{
"hooks": {
"SessionStart": [{
"matcher": "",
"hooks": [{
"type": "command",
"command": "cogz capture-event session_start --hook-json",
"timeout": 15
}]
}]
}
}
```
See [Hooks](docs/integration/hooks.md) for all 7 event types and per-agent wiring guides.
## CLI commands
Normal operation is automatic: hooks fire on lifecycle events, the agent drives CogZ through MCP. The CLI is not needed for day-to-day use — it's available for setup, manual exploration, and automation if you want or need it.
| Command | Description |
|---|---|
| `cogz init` | Initialize `.cogz/` in a repository |
| `cogz configure <harnesses>` | Write agent MCP + hook config (`auto` detects installed agents) |
| `cogz index [--no-download]` | Sync files to DB + index source code |
| `cogz reindex` | Incremental reindex (changed files only) |
| `cogz search <query>` | Hybrid FTS + vector + graph search |
| `cogz context --mode <mode> [query]` | Assemble context pack |
| `cogz status` | DB stats, entity counts, model status |
| `cogz consolidate [--dry-run]` | Run promotion and merge |
| `cogz suggest [--days N]` | List mined observation candidates |
| `cogz verify <entity-id>` | Re-stamp a drifted entity's provenance |
| `cogz reject <entity-id>` | Mark an entity rejected (`--reason` stored) |
| `cogz capture-event <type>` | Capture lifecycle event from hooks |
| `cogz models <download\|list\|clean>` | Model management |
| `cogz doctor [--prune-observations]` | Health check, policy violations, usage metrics |
| `cogz update [--check]` | Self-update from GitHub releases |
| `cogz reset [--purge]` | Drop DB (optionally purge observations) |
| `cogz mcp-stdio` | Run MCP server over stdio |
See [CLI Reference](docs/cli-reference.md) for all flags and options.
## Requirements
### Minimum (FTS-only mode)
| Resource | Requirement |
|---|---|
| RAM | 256 MB free |
| Disk | 50 MB (binary + DB, no models) |
| CPU | any x86_64 or ARM64 |
Works without ONNX Runtime or model downloads. All hooks, FTS search, context packs, consolidation, doctor, and prune are functional. Vector search, embedding-based dedup, and contradiction detection are not available.
### Recommended (hybrid search mode)
| Resource | Requirement |
|---|---|
| RAM | 2 GB free |
| Disk | 550 MB (binary + ONNX Runtime + 3 models + DB) |
| CPU | any x86_64 or ARM64, 4+ cores speeds up batch embedding |
Full functionality including vector search, semantic dedup, and NLI contradiction detection. Models auto-download on first use and auto-unload after 5 min idle (RAM drops back to ~11 MB). See [Evaluations](docs/evaluations/) for the full resource consumption profile.
## Benchmarks
CogZ ships a reproducible suite (`benchmark/`) run on pinned public corpora — httpx, cobra, clap, each injected with memory seeds mined from its real git history — plus this repository's own `.cogz` corpus. Seeded ground truth:
| Corpus | P@5 | MRR | Recall@20 |
|---|---|---|---|
| cobra | 0.200 | 0.531 | 0.967 |
| httpx | 0.173 | 0.358 | 0.917 |
| clap | 0.185 | 0.278 | 0.839 |
Channel ablations on commit queries: removing graph expansion costs 10–16pt recall@20 on every corpus; FTS-only mode retains ~75–85% of hybrid recall with ~745 MB less RSS. Context packs keep 0.70–0.90 expected-entity recall at the default 8K budget. Reruns are byte-identical. Full methodology, per-phase numbers, and the raw artifacts: [benchmark/README.md](benchmark/README.md).
**What using it buys (measured):** in a 14-task agent replay, the seeded-knowledge arm finished ~2x faster than bare (871s vs 1748s average) and completed more runs (14/14 vs 10/14) at equal correctness. Consolidation machinery is precise: dedup precision/recall 1.0, NLI contradiction detection 4/4 with zero false alarms, drift marking exact.
**Honest limits:** top-5 precision is weak on mixed corpora (P@5 <= 0.20; code entities outrank knowledge at the top of the ranking), commit-intent queries reach 0.36–0.56 recall@20, adjacent-domain negative queries leak confident hits (silence-gate clean rate 0–0.4 across corpora), and at n=14 tasks there is no measurable task-correctness lift yet.
## Architecture
- **Single Rust binary** — no runtime dependencies except optional ONNX models for vector search.
- **Files are canonical** — all entities are Markdown files. The SQLite DB is a derived index, disposable and rebuildable.
- **Code-aware** — tree-sitter indexes source code as first-class graph entities. Supported languages: Rust, Python, Go, JavaScript, TypeScript, TSX, Bash.
- **Graceful degradation** — works without ML models in FTS-only mode.
- **Local-first** — no cloud, no telemetry, no accounts. The only network access is optional model downloads.
See [Architecture](docs/design/architecture.md) for the full system design.
## Compatibility
| Platform | Support | Embeddings | FTS-only | Install |
|---|---|---|---|---|
| Linux x86_64 | Full | Auto-download | Yes | `install.sh` |
| Linux aarch64 | Full | Auto-download | Yes | `install.sh` |
| macOS arm64 (Apple Silicon) | Full | Auto-download | Yes | `install.sh` |
| macOS x86_64 (Intel) | Not supported | — | — | — |
| Windows x86_64 | Full | Auto-download | Yes | `install.ps1` or `install.sh` (Git Bash) |
**macOS Intel** is not supported because Microsoft dropped ONNX Runtime macOS Intel binaries after v1.22. Intel Mac users can run the arm64 binary under Rosetta 2 (with a compatible ORT build) or use `cargo install cogz` for FTS-only mode.
**Windows 10+** is required (bsdtar is bundled since build 17063, needed for ONNX Runtime auto-extraction).
Cross-platform team collaboration is supported: code entity UUIDs use forward-slash path normalization so the same source file produces the same entity ID on all platforms.
## Documentation
**User guides:**
- [Getting Started](docs/getting-started.md) — mental model and walkthrough
- [Configuration](docs/configuration.md) — full `config.toml` reference
- [CLI Reference](docs/cli-reference.md) — every command and flag
**Integration:**
- [MCP Tools](docs/integration/mcp-tools.md) — all 15 tool signatures and response shapes
- [Hooks](docs/integration/hooks.md) — lifecycle events and output format
- [Agent Setup](docs/integration/agent-setup.md) — all six agents + generic MCP, with per-agent effect coverage
**Design:**
- [Architecture](docs/design/architecture.md) — system overview and module map
- [Entity Model](docs/design/entity-model.md) — entity types, frontmatter, state machine
- [Search](docs/design/search.md) — hybrid FTS + vector, RRF, graph expansion
- [Consolidation](docs/design/consolidation.md) — dedup, contradiction, promotion, merge
- [Degradation](docs/design/degradation.md) — FTS-only mode and fallback behavior
**Contributing:**
- [Building](docs/dev/building.md) — build, release, cross-compile
- [Testing](docs/dev/testing.md) — test categories and mock models
- [Conventions](docs/dev/conventions.md) — code patterns and invariants
- [Dependencies](docs/dev/dependencies.md) — pinned versions and supply-chain policy
- [Schema](docs/dev/schema.md) — DB schema and migrations
## Contributing
See [CONTRIBUTING.md](CONTRIBUTING.md) for build, test, and PR guidelines.
## License
MIT — see [LICENSE](LICENSE).
## Support
If you find this tool useful, consider buying me a coffee:
[](https://buymeacoffee.com/balaianu)
TDQS
Scored across 15 tools
Tools are mostly well-differentiated with explicit cross-references ('use update_knowledge instead', 'for code entities use list_entities'). The overlap between list_entities and query_entities is real but the descriptions carefully delineate them (code vs knowledge-layer, IDs only vs content). search vs query_entities vs get_context is a slightly murky trio, but each description states its intended use.
Predominantly verb_noun snake_case (update_knowledge, verify_knowledge, list_entities, query_entities, get_impact, create_entity). A few bare verbs (search, consolidate) and internal/verb-only names (reject_entity, capture_event, suggest_observations) are acceptable and readable. Minor deviation but consistent overall.
15 tools for a knowledge/code-graph memory system with retrieval, lifecycle, write, and maintenance concerns is reasonable. One or two (capture_event explicitly marked 'not intended for direct use') could be hidden from the agent-facing surface, but nothing feels excessive.
Covers the full knowledge lifecycle: create, update, verify, reject, consolidate, plus retrieval (search, query, context, impact, callers, orphans) and status. Gaps are minor — no explicit delete of knowledge-layer entities (rejection/supersede likely intended instead) and no direct supersede tool despite it being referenced.