Skip to main content
Glama

brain-v42

Persistent memory for coding agents, served over MCP.

brain-v42 gives Claude Code, Codex and any other MCP client a durable second brain: decisions, learnings, code snippets, runbooks, ADRs, tickets and project roadmaps — stored in PostgreSQL, retrieved by full-text + semantic search with reranking, and consolidated every night by an agent pipeline.

  • Typed knowledge, not a notes dump — a decision records its WHY and alternatives; a snippet records its intent; a runbook records executable steps. Each type has its own lifecycle (supersession chains, ADR acceptance, learning validation).

  • Explicit session lifecycle — the user owns every session boundary. Sessions capture the artifacts they produced, and closing is fail-closed: a session ends with either captured knowledge or an explicit "nothing to capture" reason, never silence.

  • Search that ranks — pgvector semantic search + PostgreSQL FTS, fused and re-ranked by a cross-encoder.

  • Nightly consolidation ("dream") — an agent pipeline cleans orphan links, merges duplicates, synthesises learnings and proposes promotions, behind per-phase killswitches that all ship closed.

  • Multi-project — per-project focus with compare-and-swap revisions, roadmaps, cross-project tickets.

  • Observable delivery — versioned delivery contracts bind addressed work to a pull request, persisted CI evidence, integration, and policy-governed fulfillment.

  • Measured facts, not remembered ones — a closed catalogue of named probes reads live state (schema head, running release, declared killswitches) against a source identity the operator declared independently, so a briefing states what is rather than what someone last wrote down.

Architecture

Claude Code / Codex (MCP client)
       │ HTTP loopback :8765/mcp (production) · stdio (dev/fallback)
  brain-v42 (FastMCP)
       ├── SQLAlchemy async ─▶ PostgreSQL 16 + pgvector   (source of truth)
       ├── HTTP ─────────────▶ embedding endpoint :8003   (optional, pluggable)
       ├── HTTP ─────────────▶ :8003/rerank               (optional reranker)
       └── bolt ─────────────▶ Neo4j 5 Community          (relationship index, optional)

MCP transport: production = HTTP loopback http://127.0.0.1:8765/mcp; configuration default and dev/fallback = stdio.

PostgreSQL is the single source of truth. Neo4j is a disposable projection fed by a relational ledger/outbox — it can always be rebuilt from PostgreSQL, never the other way around. The canonical path is active in production since 22 July 2026; design and evidence live in docs/ARCHITECTURE.md and the graph ledger runbook.

Embeddings are optional and pluggable, and they degrade gracefully when the endpoint is away — brain_search falls back to full-text search, writes persist with a NULL embedding and are backfilled later. An install with no embedding endpoint at all works.

Two wire formats ship, selected by BRAIN_EMBEDDING_BACKEND:

Backend

Wire

Use it for

shim (default)

POST /embed, POST /embed/query, GET /healthz

The bundled reference stack (services/), serving Qodo-Embed-1-1.5B as GGUF via llama.cpp on a local GPU

openai

POST /v1/embeddings

Any OpenAI-compatible endpoint — Ollama, vLLM, llama.cpp server, LM Studio, TEI, Jina, Mistral, Voyage, OpenAI

So a machine without a GPU needs no bundled stack. Point it at whatever serves embeddings, for example a local Ollama:

BRAIN_EMBEDDING_BACKEND=openai
BRAIN_EMBEDDING_SERVICE_URL=http://localhost:11434
BRAIN_EMBEDDING_MODEL=nomic-embed-text
EMBEDDING_DIMENSION=768   # unprefixed on purpose — see Dimension below

Reranking is separately pluggable via BRAIN_RERANK_BACKEND (shim, or cohere for the POST /v1/rerank shape implemented by TEI, Jina and vLLM), and stays best-effort: an unavailable reranker falls back to RRF ordering rather than failing a search.

Instruction prefixes

Asymmetric models (Qodo, E5, BGE) expect queries and documents to be marked differently. Both prefixes default to empty, which is correct for symmetric models and reproduces the unprefixed behaviour exactly:

BRAIN_EMBEDDING_QUERY_PREFIX="query: "        # applied to searches only
BRAIN_EMBEDDING_DOCUMENT_PREFIX="passage: "   # applied to everything written

Keep the trailing space if the model's card shows one — it is part of the prefix. Changing the query prefix is free. Changing the document prefix on a populated corpus requires a full scripts/regen_embeddings.py pass, or the column ends up holding two incompatible vector populations with nothing to flag it.

Dimension

EMBEDDING_DIMENSION is chosen at install time and must be ≤ 2000, the ceiling pgvector's HNSW index accepts. Switching models later means re-embedding the corpus (scripts/regen_embeddings.py).

Set it unprefixed. Every other setting here also answers to a BRAIN_-prefixed name, but the ORM column widths are read straight from EMBEDDING_DIMENSION in db/tables.py, so BRAIN_EMBEDDING_DIMENSION=768 alone would leave the tables at 1536 while the rest of the process believed 768.

One caveat to know before a non-default install: the ORM honours this setting, but four migrations (002, 005, 009, 014) hardcode vector(1536), so a fresh alembic upgrade head creates 1536-wide columns whatever the setting says. Until that is fixed, a non-1536 install needs the columns retyped and their HNSW indexes rebuilt by hand after migrating. If you are running the reference stack at 1536, this does not affect you.

Related MCP server: Muninn

Quick start

git clone https://github.com/hawkixs/brain-v42 && cd brain-v42
uv sync --extra dev --python 3.12      # creates .venv; see "Development" for why not pip
source .venv/bin/activate
cp .env.example .env                   # set POSTGRES_PASSWORD, and the same value in POSTGRES_URL

# 1. Secrets docker-compose.yml expects on the host, and the network the
#    embedding services attach to (compose refuses to start without either).
docker network create brain-net
install -d -m 0700 .secrets
read -rsp "Neo4j password (written to .secrets/neo4j-auth, the file compose mounts): " PW
(umask 0022; printf 'neo4j/%s\n' "$PW" > .secrets/neo4j-auth); unset PW
(umask 0177; openssl rand -hex 32 > .secrets/embedding-shim-bearer)

# 2. Databases (PostgreSQL 16 + pgvector, Neo4j). embedding-llama and
#    embedding-shim additionally need an NVIDIA GPU and the GGUF model file
#    (see "Embeddings" above) — on a machine without one, start only the two
#    services `pytest tests/integration` and the migrations below need:
#    `docker compose up -d postgres neo4j`.
docker compose up -d

# 3. Migrations
export POSTGRES_URL="$(grep -E '^POSTGRES_URL=' .env | cut -d= -f2-)"   # alembic reads the environment, not .env
BRAIN_ALEMBIC_ALLOW_PROD=1 alembic upgrade head

# 4. Run the MCP server (stdio)
python -m brain_v42.mcp.server

Wire it into Claude Code — .mcp.json at the repo root already targets the production HTTP loopback endpoint; for a plain stdio dev setup:

claude mcp add brain-v42 -- python -m brain_v42.mcp.server

BRAIN_ALEMBIC_ALLOW_PROD is required only when the database name is exactly brain; keep it a one-command opt-in, never exported persistently. Alembic rejects DSN query parameters; use the plain form above with host, port, username and password all present.

MCP tools

Domain

Tools

Search & list

brain_search, brain_list, brain_get, brain_update, brain_delete

Graph traversal

brain_get_neighbors, brain_graph_path

Session lifecycle

brain_session_start, brain_session_list, brain_session_resume, brain_session_capture, brain_session_heartbeat, brain_session_checkpoint, brain_session_end, brain_session_abandon

Project context

brain_set_project_context, brain_update_project_focus, brain_focus_history, brain_list_projects, brain_list_project_groups, brain_project_archive, brain_project_unarchive

Decisions

brain_log_decision, brain_supersede_decision, brain_get_supersession_chain

Learnings

brain_learn, brain_validate_learning

Snippets

brain_save_snippet, brain_use_snippet

Runbooks

brain_create_runbook, brain_promote_runbook, brain_get_runbook, brain_execute_runbook

ADRs

brain_propose_adr, brain_promote_adr, brain_accept_adr, brain_deprecate_adr

Coordination

brain_ticket_create, brain_ticket_reply, brain_ticket_transition, brain_ticket_list, brain_ticket_get

Observable delivery

brain_delivery_contract_set, brain_delivery_bind_pr, brain_delivery_get, brain_delivery_list, brain_delivery_refresh, brain_delivery_claim, brain_delivery_claim_renew, brain_delivery_claim_release, brain_delivery_accept, brain_delivery_attest, brain_delivery_attestation_list

Dream / graph

brain_get_clusters, brain_backfill_links_batch, brain_consolidation_candidates, brain_merge_entities, brain_refresh_entity, brain_reindex_plans, brain_list_orphans_for_classification, brain_assign_domain, brain_list_curation_proposals, brain_reject_curation_proposals, brain_apply_curation_proposal

Roadmap & decay

brain_get_roadmap, brain_feature_create, brain_feature_update, brain_decay_status

Workflow guidance

brain_workflow_guide

Measured facts

brain_fact_list, brain_fact_get

Claims

brain_claim_verify, brain_claim_list, brain_claim_history

Full catalog with signatures: docs/MCP_TOOLS.md.

The default catalog profile is compact: the seven session lifecycle tools stay visible, and every other tool is reached through two gateways — brain_find_tool to discover, brain_call_tool to invoke. Set BRAIN_MCP_PROFILE=native to expose every tool directly.

Observable delivery

One delivery path is contract → PR binding → persisted PR/CI observation → integration receipt → explicit requester acceptance. A merge receipt proves the contracted revision reached its target branch. Contracts with acceptance_mode=explicit require the requester to accept the current attempt and delivery digest; automatic contracts can produce fulfillment without that decision.

Brain stores and evaluates this evidence. It does not launch agents, choose work, push commits, merge pull requests, or deploy releases. External orchestrators keep those responsibilities. Delivery reads use persisted observations and make no GitHub calls. With BRAIN_DELIVERY_ENABLED=false, mutations pause while reads and the completion guard for existing contracts remain active.

Ledger and policy: the boundary with red-rail

Brain is the ledger; red-rail is policy. The split is deliberate and it is the reason attestations exist as a separate table (ticket 04bc1f4a).

brain_delivery_attest records one issuer-declared fact about a ticket's delivery workflow — released, deployed, rolled_back, incident_detected, gate_passed and their kin. Brain validates the form and nothing else: the kind must match ^[a-z][a-z0-9_]{0,63}$, the payload must be a bounded JSON object, and the server — never the issuer — computes the digest. Brain never judges what a kind means, never derives a completion refusal from an attestation, and never updates or deletes one. The well-known kinds above are documentation, not an allowlist.

That restraint is what makes the rows usable: red-rail reads them to compute DORA metrics under its own policy, which can change without a schema migration here. Attestations also arrive legitimately after a ticket has closed — an incident or a rollback does not wait for a workflow state — so no ticket-status restriction applies.

For consumers that must not import brain_v42, the whole API is published as data in docs/contracts/delivery_attestations.json — kinds, bounds, the digest recipe with its test vectors, the stable error codes and the list scopes — and a unit test keeps that file equal to the code.

The observer is a separate process with separate credentials. Its dedicated ~/.config/brain-v42/delivery-observer.env must be an owned, regular, non-symlink file with mode 0600. It carries BRAIN_DELIVERY_ENABLED=true, an explicit BRAIN_DELIVERY_POSTGRES_URL, the repository registry, and either a dedicated GitHub token or a complete GitHub App credential set. Keep the observer's GitHub credentials out of the shared application environment; the application keeps its own existing PostgreSQL configuration. The immutable service command is:

<release>/venv/bin/python -m brain_v42.delivery_observer --env-file ~/.config/brain-v42/delivery-observer.env

Each host release lives under ~/.local/share/brain-v42/releases/<full-source-sha>/ and retains the same-SHA source archive, wheel, lock, copied Python 3.12 environment, and hashed manifest. The archive supplies Dream and root scripts that the wheel does not install. Build and installed-wheel checks do not attest a rollout. Follow the immutable delivery release and canary runbook (docs/runbooks/2026-09-07-observable-delivery-workflows.md in the private brain-v42-internal repository) for preflight, activation, evidence capture, and compatible forward rollback.

Sessions

The user controls every session boundary: start, resume, end and abandon are explicit commands, never inferred by a hook, an agent or a client. Sessions capture the durable artifacts they produced into an exclusive ledger, and closing is fail-closed: captured knowledge or an explicit "nothing to capture" reason, never silence.

After 24 hours without a heartbeat, an open session exposes is_stale=true; the marker is derived, the persistent status stays open, and only the 7-day server-side sweep ever abandons a session without an explicit user command.

The full lifecycle contract (capture rules, focus semantics, briefing) lives in docs/MCP_TOOLS.md; the contract is v4 and still evolving.

Configuration (.env)

# Required
POSTGRES_URL=postgresql+asyncpg://brain:change-me-locally@localhost:5433/brain

# Optional — semantic search and reranking
EMBEDDING_SERVICE_URL=http://localhost:8003
EMBEDDING_DIMENSION=1536              # <= 2000 (pgvector HNSW ceiling)
RERANKER_URL=http://localhost:8003

# Optional — point at any OpenAI-compatible endpoint instead of the bundled stack
BRAIN_EMBEDDING_BACKEND=shim          # shim (default) or openai
BRAIN_EMBEDDING_MODEL=qodo            # model name sent by the openai backend
BRAIN_EMBEDDING_QUERY_PREFIX=         # e.g. "query: " for asymmetric models
BRAIN_EMBEDDING_DOCUMENT_PREFIX=      # changing this needs a full re-embed
BRAIN_RERANK_BACKEND=shim             # shim (default) or cohere

# Optional — relationship graph (safe defaults for a fresh environment)
GRAPH_ENABLED=false
GRAPH_LEDGER_WRITE_ENABLED=false

# Tool catalog profile
BRAIN_MCP_PROFILE=compact   # compact (default) or native

# Observable delivery MCP mutations; reads and existing completion guards remain active
BRAIN_DELIVERY_ENABLED=false

LOG_LEVEL=INFO

Never place MCP_HTTP_TOKEN or MCP_HTTP_DREAM_TOKENS in the shared .env: bearer tokens live in a private 0600 file (~/.config/brain-v42/mcp-token.env), and the graph projector credential in its own (~/.config/brain-v42/graph-projector.env). BRAIN_EMBEDDING_API_KEY and BRAIN_RERANK_API_KEY follow the same rule. Note that pointing either endpoint at a hosted provider sends the text being embedded off this machine — the rest of this deployment is loopback-bound, that step is not. Full reference — every variable, the private secret files, preflights and rollout gates: docs/OPERATIONS.md.

Network trust model

The deployment targets personal agents on a trusted LAN. MCP, PostgreSQL and Neo4j bind to loopback; metrics and automation default to loopback.

Embedding topology: production/default = local unified endpoint http://localhost:8003; the personal dev-pc deployment is a superseded rollback/reference path, now private.

The reranker shares the unified embedding endpoint :8003/rerank. Treat :8003 as LAN-exposed until you have proved the live bind yourself, and never expose it — or the MCP port — to the Internet. Repository code alone does not prove a live firewall state.

Dream mode

Nightly agent pipeline (scripts/dream.sh: scan → clean → connect → synth → promote → reorg) plus server-side ticket-extraction, roadmap-curation and session-sweep jobs. Every mutating phase sits behind a killswitch and every killswitch ships closed; dry-run is the shipped default. Each phase runs under an exact MCP tool allowlist, and each phase holds a capability bearer scoped to its (project, phase) pair, so a phase sees only the project it was started for.

Phases run on an ordered chain of agent providers (BRAIN_DREAM_AGENT_PROVIDERS), each preflighted before the night starts. When a link is exhausted or unreachable the run falls through to the next rather than failing the phase, and a link that dies before any Brain call is replayable — it is retired for the night and not charged to the retry budget, because a provider that never reached Brain performed no work to redo. Every phase writes which link served it, and what it fell through, to logs/dream/<date>_<project>_<phase>.chain.json. Read that file rather than assuming the first link ran: a night can be entirely green on one provider and prove nothing about the fallthrough. Details: docs/ARCHITECTURE.md and docs/OPERATIONS.md.

Production state

The repository migration target is migration 057. No page in this repository proves a live schema head — measure it, do not read it here:

docker exec brain_v42_postgres psql -U brain -d brain -Atc "select version_num from alembic_version;"

The running build names itself: GET /health returns version (the installed distribution) and alembic_head (the revision shipped with it), both measured, never written by hand.

Measured facts

The sentence above — measure it, do not read it here — is the rule. The facts registry is the mechanism that enforces it, so the session briefing can state the live schema head, the running release and the declared Dream killswitches without anyone retyping them into a document.

A fact is a named, versioned reader. Each one declares its target, its TTL, its timeout, its policies and the exact shape of the value it returns, and the catalogue is closed: probes are registered once at composition and then frozen, so no runtime caller can install a reader of its own.

What makes a reading trustworthy is not the probe but the source identity, and the identity a probe is checked against is declared independently by the operator — never derived from the connection the probe uses. A PostgreSQL reading must match a cluster system identifier, database, address and port the operator wrote down; a live_release reading must match the release SHA and package version the release tooling rendered; a host reading must match a declared hostname. Undeclared means no fact: the probe is refused at registration and the briefing says which one is missing and why, rather than leaving a silent gap. This closes the obvious hole — a probe that reports its own DSN back to you proves nothing about which database it reached.

A measurement has exactly two shapes and no third: Measured, carrying a bounded canonical JSON value and a digest over it, or Unreadable, carrying a closed error_code. A timeout, an unexpected identity or an over-large value each produce an Unreadable that renders as such — never a stale value dressed up as current.

Three targets ship today (production, live_release, host) and the catalogue is declared in src/brain_v42/facts/composition.py. Read it with brain_fact_list and brain_fact_get; facts declared briefing=true also render as lines in the session briefing. Design: docs/superpowers/specs/2026-09-19-measured-facts-and-claims-design.md in the private brain-v42-internal repository.

Development

pytest tests/unit -v                          # no PostgreSQL required
pytest --cov=brain_v42 --cov-report=term-missing
ruff check src/ tests/ && ruff format --check src/ tests/
mypy src/
  • pytest tests/unit needs no env var at all on a fresh clone: a tests/unit fixture hands Settings() a syntactically valid, unreachable database URL whenever neither POSTGRES_URL nor BRAIN_POSTGRES_URL is set. A handful of tests opt into a REAL PostgreSQL and skip loudly ("BRAIN_V42_TEST_DB_URL not set — skipping...") unless BRAIN_V42_TEST_DB_URL points at an isolated test database — see CONTRIBUTING "Running the tests".

  • Stack: Python 3.12+, FastMCP 3.x, SQLAlchemy 2.0 async + asyncpg, Alembic, Pydantic 2, structlog.

  • TDD is mandatory — red, green, refactor; tests are never edited to make code pass.

  • Coverage floor: 60% (CI blocks below).

  • Install with uv sync --extra dev --python 3.12, not with pip. pip install -e ".[dev]" fails on this layout and always has: headless-agents is a uv workspace member ([tool.uv.workspace] + [tool.uv.sources] in pyproject.toml), not a published distribution, so pip looks for it on PyPI and stops with No matching distribution found for headless-agents. The dev toolchain is pinned exactly in uv.lock, so a synced environment resolves to the versions CI runs.

  • Pin --python 3.12 explicitly. requires-python is >=3.12, so a bare uv sync on a fresh clone picks the newest interpreter it can find — measured 3.14 — while every CI job, the release job and [tool.mypy] target 3.12. Matching CI is the whole point of the lock; an unpinned interpreter quietly gives it up.

Project layout

brain-v42/
├── src/brain_v42/
│   ├── config.py              # pydantic-settings — single config surface
│   ├── db/                    # SQLAlchemy engine + tables
│   ├── models/                # Pydantic models
│   ├── repositories/          # CRUD + FTS + pgvector + graph adapters
│   ├── services/              # business logic, embedding, reranker, dream, dedup
│   ├── metrics/               # sidecar + collector + cockpit endpoint
│   ├── automation/            # independent webhook/dedup runtime (:9201)
│   └── mcp/                   # FastMCP server + brain_*/dream_* tool handlers
├── tests/{unit,integration}
├── alembic/versions/          # migrations (shipped inside the wheel)
├── scripts/                   # operational CLIs (dream.sh, canaries, repair)
├── services/                  # GPU embedding service + shim + supervisor
├── deploy/                    # systemd units, per-host compose, install.sh
└── docs/                      # ARCHITECTURE, SCHEMA, MCP_TOOLS, OPERATIONS, runbooks

The top-level module graph is enforced acyclic in CI (scripts/check_module_layering.py): any module can still be extracted into a standalone service without dragging a cycle with it.

CI/CD

Stages: lint → test → security → build. Security gates: pip-audit, bandit, gitleaks, container-image pin checks. Docker images are built and pushed on main; there is no deploy stage — rollout to a host is always a manual, out-of-band step. Releases are tag-driven: the release rail builds the wheel + sdist, proves the wheel ships its migrations, and attaches both to the GitHub release.

Versioning

  • The shipped version is 0.6.1, and it stays 0.x on purpose: a 1.0.0 would promise a stable interface and a way back, and this project has neither yet.

  • No lossless downgrade is promised, at any version. Several migrations protect stored history: 037 refuses when a session capture would be lost, 039 requires an explicit operator opt-in, 053 refuses once delivery workflow history exists, 054 refuses once a delivery attestation exists, 055 refuses once a claim, a verdict or a fact definition exists, and 056 refuses while a project is archived; those two accept a named operator opt-in.

  • Follow the release's operator runbook for recovery. For 0.6.1 (as for 0.6.0 and 0.5.0), use the compatible forward rollback section of that runbook (docs/runbooks/2026-09-07-observable- delivery-workflows.md in the private brain-v42-internal repository) and keep the repository's migration target in place: the head the release ships, never a lower one. Rollback means selecting a release that supports that head or deploying a forward fix; it never means alembic downgrade, and never means restoring an older dump over a live database.

  • 0.6.0 ships the headless-agents workspace member (packages/headless-agents/, version 0.1.0) as a second distribution that brain_v42 depends on; its own version moves independently of this one.

License

Source code: Apache-2.0.

Model weights are not covered by that license, and this is not a formality. The production embedding model, Qodo/Qodo-Embed-1-1.5B, is published under QodoAI-Open-RAIL-M — a license carrying use-based restrictions, not a permissive one. No weights are stored in or distributed by this repository: every model is downloaded from its upstream host at build time, by the operator, who accepts each model's terms directly from its publisher. See NOTICE before redistributing anything.

Available Tools

9 tools
brain_call_toolA

Call a tool by name with the given arguments.

Use this to execute tools discovered via search_tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe name of the tool to call
argumentsNoArguments to pass to the tool

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full responsibility for behavioral disclosure. It only says 'Call a tool by name' and does not mention potential side effects, error behavior, dynamic execution risk, or that it can invoke arbitrary tools with arbitrary consequences. This is a significant gap for an execution tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, action stated first. The second sentence provides essential context (use after search_tools) without redundancy. Excellent front-loading.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 2-parameter tool, the description covers the core action and usage context. However, without an output schema or annotations, it omits return value behavior and error handling, which an agent might need when invoking arbitrary tools. It is adequate but not comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents both parameters. The description adds marginal value by linking parameters to 'given arguments' from search_tools, but does not enrich semantics beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Call') and resource ('a tool by name with the given arguments'), and differentiates itself from discovery tools by positioning it as the execution step after search_tools. This makes it unambiguous against siblings like brain_find_tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Use this to execute tools discovered via search_tools' gives clear context of when to invoke it (after discovery) and implies it is not for searching. It does not explicitly exclude session-management siblings, but the usage context is sufficient for a generic dispatcher.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brain_find_toolA

Search for tools using natural language.

Returns matching tool definitions ranked by relevance, in the same format as list_tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesNatural language query to search for tools

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden. It discloses that the tool returns definitions rather than executing tools, indicates relevance ranking, and references the list_tools output format. It does not explicitly state lack of side effects, but the search/return framing makes the read-only behavior clear enough.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The action is front-loaded and the return behavior is described in the second sentence. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, output-schema-bearing search tool, the description is complete: it states the query source, the result type, ranking behavior, and output format. No critical invocation information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'query' parameter is already described as a natural language query. The description adds no additional semantic detail beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('Search for tools') and resource ('tool definitions'), and adds that results are ranked by relevance. This differentiates it from session lifecycle tools and brain_call_tool without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies its use case: natural-language discovery of tools, as opposed to session manipulation or direct tool invocation. However, it does not explicitly name alternatives or state when not to use it, such as preferring list_tools for browsing all tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brain_session_abandonB
DestructiveIdempotent

Abandon a session only from an explicit user command.

An agent tracer is the only session the server opens or closes on its own; no hook and no auto-close may invoke this lifecycle boundary.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYes
session_idYes
expected_client_keyYesClient identity expected for the addressed session UUID. The pair must match before any session mutation. This is an isolation guard, not authentication.

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, which cover the destructive nature. The description adds the explicit-user-command constraint and notes that only an agent tracer session is server-managed. However, it does not describe consequences (e.g., irreversibility, impact on session data) beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with the primary action and constraint. The second sentence adds relevant context about the agent tracer without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with three required parameters and no output schema, the description is sparse. It does not explain what 'abandon' entails versus brain_session_end, what the response contains, or prerequisites beyond the schema. The annotation and schema cover only part of the operational context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (only expected_client_key has a description). The tool description provides no parameter information, failing to compensate for the low coverage. An agent receives no guidance on session_id or reason beyond the schema's minimal type constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action ('abandon a session') and resource, and adds a critical constraint ('only from an explicit user command'). It does not explicitly differentiate from the sibling brain_session_end, but the user-command qualifier hints at a distinct lifecycle role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance ('only from an explicit user command') and when-not-to-use ('no hook and no auto-close may invoke this lifecycle boundary'). It does not name alternative tools directly, but the constraint is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brain_session_captureA
Idempotent

Attach durable artifacts from an explicit user command — to repair or refine.

No longer a compulsory ritual. With derived capture armed, artifacts created during the session are attributed without this call, and end no longer demands a receipt. What this tool is for now: attaching something the derivation could not see — created outside the window, on another connection, or in another project — and saying so on the record.

It never steals: an artifact already attributed stays where it is.

An agent tracer is the only session the server opens or closes on its own; no hook and no auto-close may invoke this lifecycle boundary.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes
knowledge_idsYes
expected_client_keyYesClient identity expected for the addressed session UUID. The pair must match before any session mutation. This is an isolation guard, not authentication.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover idempotency and non-destructiveness; the description adds that already-attributed artifacts are left in place and that this is not a compulsory step. It does not disclose return behavior, permissions, or failure modes, so with annotations lowering the bar, this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and usage are front-loaded in the first paragraphs, but the explanation includes redundant historical context ('No longer a compulsory ritual' followed by the derived-capture explanation) and a final lifecycle sentence that is tangential. It is readable but not every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a side-effecting 3-parameter tool with no output schema and with annotations, the description provides enough usage context to call correctly, including scope and idempotent ownership behavior. It lacks explicit statements about return values or errors and uses jargon like 'derived capture' and 'receipt' without definition, so completeness is only moderate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema describes only expected_client_key; the description adds meaning for knowledge_ids as durable artifacts and specifies their suitability (outside window/connection/project), plus the 'never steals' rule for already-attributed IDs. It does not discuss session_id or expected_client_key beyond schema, so it only partially compensates for the 33% schema description coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence identifies a specific action ('Attach durable artifacts') and trigger ('from an explicit user command'). Later text narrows the scope to artifacts the derivation could not see, which differentiates it from automatic capture, though it does not name sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: for artifacts created outside the session window, on another connection, or in another project. It also says when it is not needed, because derived capture attributes normal session artifacts and `end` no longer requires a receipt. It stops short of naming sibling alternatives, but the when/when-not guidance is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brain_session_endA
DestructiveIdempotent

End a session from an explicit user command, on judgement alone.

nothing_to_capture_reason is now OPTIONAL. The old rule — a non-empty ledger XOR a written reason — measured whether the client had DECLARED its work. Derived capture feeds that signal from the server, and a check is hollow the moment the thing it checks can influence its own signal; worse, it made a session whose ledger the server had filled impossible to close. Give a reason if you have one to give; a blank one is still refused, because saying nothing and saying " " are not the same act.

What end still requires is what the server cannot produce for you: a summary and a next focus. It REPORTS unattributed_in_window — artifacts of this project created during the session that belong to no ledger — as a measure, never a gate. It cannot refuse a close, and a session cannot improve it by doing nothing.

An agent tracer is the only session the server opens or closes on its own; no hook and no auto-close may invoke this lifecycle boundary.

ParametersJSON Schema
NameRequiredDescriptionDefault
summaryYes
next_focusYesJugement et engagements pour la session suivante — ce qui n'est pas dérivable automatiquement (ex : ne pas publier ce brouillon et pourquoi, une échéance et sa raison, une décision opérateur à respecter). N'y recopie pas l'état mesurable (révision de schéma, HEAD git, arbre propre, résultat de suite de tests) : il est déjà recalculé à chaque briefing depuis la source réelle, et une copie manuelle se périme en silence.
session_idYes
expected_client_keyYesClient identity expected for the addressed session UUID. The pair must match before any session mutation. This is an isolation guard, not authentication.
expected_focus_revisionYes
nothing_to_capture_reasonNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
sessionYesPersistent state of one explicitly controlled concurrent session.
replayedYes
focus_diffNo
focus_at_endYes
current_focusYes
focus_outcomeYesPersisted outcome of the one focus update attempted while ending.
focus_revision_at_endNo
current_focus_revisionYes
unattributed_in_windowYes
remaining_open_session_countYes

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by explaining that nothing_to_capture_reason is optional, that blank reasons are still refused, that unattributed_in_window is reported as a measure rather than a gate, and that the server cannot provide the required summary or next_focus. It also clarifies the only server-initiated lifecycle case. No statement contradicts the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the purpose and organized into clear blocks: purpose, parameter behavior, required inputs, and lifecycle boundary. The historical rationale about the old XOR check is somewhat longer than strictly necessary, but it supports the blank-reason rule. It is dense and structured, though trimming the design history would make it leaner.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive lifecycle operation with an output schema and annotations, the description covers the essential decision factors: when it may be invoked, what must be supplied, and what is reported rather than enforced. The remaining gaps are the lack of an explicit distinction from brain_session_abandon and the undocumented expected_focus_revision parameter. Still, an agent has enough context to call the tool correctly in most cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 33% schema coverage, the description compensates meaningfully by explaining the optionality and refusal semantics of nothing_to_capture_reason and by emphasizing that summary and next_focus are required because the server cannot derive them. It does not explain expected_focus_revision or add much beyond what the schema already provides for the other parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'End a session from an explicit user command, on judgement alone.' This makes the core operation clear. However, it does not explicitly contrast itself with the sibling tool brain_session_abandon, so the agent must infer how ending differs from abandoning a session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear invocation context: only an explicit user command, never a hook or auto-close, may invoke this lifecycle boundary. It also states that callers must supply a summary and next_focus. It does not name alternatives or exclusion conditions relative to sibling session tools, which prevents a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brain_session_heartbeatA

Refresh an open session from an explicit user command.

An agent tracer is the only session the server opens or closes on its own; no hook and no auto-close may invoke this lifecycle boundary.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes
expected_client_keyYesClient identity expected for the addressed session UUID. The pair must match before any session mutation. This is an isolation guard, not authentication.

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, so the mutation is known. The description adds that this is a lifecycle boundary and only for explicit commands, but it does not explain what 'refresh' does to the session (e.g., extends timeout) or any side effects. It adds some context but not deep behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero wasted words. The main purpose is front-loaded, and the critical usage constraint is stated immediately after. Perfectly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation with no output schema, the description is somewhat bare. It does not specify behavior for a closed session, failure modes, or the return value. The agent can infer that the session must be open from the phrase 'open session', but it is not explicitly guaranteed. Given the simplicity, this is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: session_id has only a format, while expected_client_key has a detailed description. The tool description itself does not elaborate on either parameter. Since the missing session_id is self-explanatory (a UUID), the description does not need to compensate much, but it still adds no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('refresh') with a clear resource ('an open session') and immediately distinguishes itself from siblings by stating it is the only lifecycle boundary invoked by explicit user command, not by hooks or auto-close. This fully differentiates it from brain_session_start/end/resume.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use the tool ('from an explicit user command') and, more importantly, when not to use it ('no hook and no auto-close may invoke this lifecycle boundary'). This gives the agent a clear decision rule, even though alternatives are not named directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brain_session_listA
Read-onlyIdempotent

List sessions only in response to an explicit user command.

An agent tracer is the only session the server opens or closes on its own; no hook and no auto-close may invoke this lifecycle boundary.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
statusNoopen
project_keyNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, so the safe read-only nature is established. The description adds meaningful behavioral context beyond those annotations: the server only auto-opens/closes an agent tracer session, and hooks/auto-close must never invoke this boundary. This is valuable context, though it does not describe list ordering, pagination, or return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The primary action is front-loaded, and the critical invocation restriction is placed in the supporting sentence. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description fully covers the unusual invocation policy, which is likely the highest-risk aspect of this tool. However, with no output schema and no parameter descriptions, the agent still lacks context about what a session entry contains, which status filters mean, and what the list response looks like. Adequate for basic invocation but with meaningful gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not compensate for any of the four parameters. The schema provides names, types, defaults, and the status enum, but meanings like 'stale', 'abandoned', and 'closed_inactive' are left unexplained, and project_key semantics are not addressed. With zero coverage, the description should add parameter context but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource ('List sessions') and adds a distinctive policy constraint: it runs only on explicit user command. It does not explain what kinds of sessions are returned or how this differs from sibling read tools, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit when-to-use condition ('only in response to an explicit user command') and an explicit when-not-to-use condition ('no hook and no auto-close may invoke this lifecycle boundary'). It does not mention sibling alternatives, but the trigger and exclusion are clear enough for an agent to decide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brain_session_resumeB
Read-onlyIdempotent

Resume an open session from an explicit user command.

An agent tracer is the only session the server opens or closes on its own; no hook and no auto-close may invoke this lifecycle boundary.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes
expected_client_keyYesClient identity expected for the addressed session UUID. The pair must match before any session mutation. This is an isolation guard, not authentication.

Output Schema

ParametersJSON Schema
NameRequiredDescription
sessionYesPersistent state of one explicitly controlled concurrent session.
briefingNo
current_focusYes
open_session_countYes
current_focus_revisionYes

TDQS

B3.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

This is an annotation contradiction. The description frames the operation as a lifecycle mutation ('resume', 'lifecycle boundary'), and the input schema mentions guarding 'any session mutation', while annotations declare readOnlyHint=true. An agent cannot trust whether this tool changes session state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, with the primary action and trigger front-loaded and the important system-level restriction stated immediately after. There is no filler; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The trigger restriction is clear, but the tool is incomplete for a lifecycle operation: the reader is not told what resuming actually changes, whether the session must already be open, or how expected_client_key mismatch is handled. The annotation contradiction further undermines the safety profile, so the presence of an output schema does not make this definition sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds minimal semantic value by indicating session_id should refer to an open session, while expected_client_key is already well described in the schema. It does not explain what happens on a client-key mismatch or where the expected client key comes from, but the parameter-level meaning is at least minimally viable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (resume), object (an open session), and trigger (an explicit user command). It distinguishes this from lifecycle siblings by emphasizing the user-command boundary, but it does not explicitly name an alternative tool or detail what resuming entails, so it falls just short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says this is for an explicit user command and that hooks and auto-close must not invoke this lifecycle boundary, giving clear when and when-not guidance. It does not mention how to handle a non-open session or when to prefer a sibling tool such as brain_session_start, so it is not fully complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brain_session_startA
Idempotent

Start or replay a concurrent session from an explicit user command.

An agent tracer is the only session the server opens or closes on its own; no hook and no auto-close may invoke this lifecycle boundary.

ParametersJSON Schema
NameRequiredDescriptionDefault
client_keyYesStable identity for one intended session. Reuse it for every retry of that session; use a distinct stable key for each parallel session.
project_keyYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
sessionYesPersistent state of one explicitly controlled concurrent session.
briefingNo
replayedYes
open_session_countYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover idempotency, destructiveness, and read-only hints. The description adds value by disclosing that this is a controlled lifecycle boundary and that the server only opens/closes agent tracer sessions on its own. This is behavioral context beyond what annotations provide, with no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The main action and trigger are front-loaded, and the second sentence earns its place by preventing incorrect automatic invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema and annotations cover return values and safety traits, so those gaps are acceptable. However, the required project_key parameter is left unexplained, and the relationship between 'start/replay' and the sibling brain_session_resume is not clarified. This is a meaningful gap for an agent deciding which session tool to call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds no parameter-level meaning. The schema describes client_key well, but project_key has no description and is required; the description does not compensate for that gap. With 50% schema coverage, the description should have provided at least some project/session context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Start or replay') and the resource ('a concurrent session'), plus the specific trigger ('from an explicit user command'). It does not explicitly distinguish itself from sibling brain_session_resume, and 'replay' vs 'resume' creates some potential overlap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use context: only from an explicit user command. It also gives a clear exclusion: no hook and no auto-close may invoke this lifecycle boundary. It does not name alternatives like brain_session_resume, but the trigger guidance is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.6.1
    • First observedbrain_call_tool
    • First observedbrain_find_tool
    • First observedbrain_session_abandon
    • First observedbrain_session_capture
    • First observedbrain_session_end
    • First observedbrain_session_heartbeat
    • First observedbrain_session_list
    • First observedbrain_session_resume
    • First observedbrain_session_start

TDQS

A3.7/5.0

Scored across 9 tools

Disambiguation3/5

Most session tools map to distinct lifecycle states, but resume and heartbeat both target an open session with no clear boundary between them, and start's 'replay' capability also blurs with resume. The find_tool/call_tool pair is clearly distinct, which keeps the set usable.

Naming Consistency4/5

All tools share the brain_ prefix and snake_case, and the seven session tools follow a consistent brain_session_<verb> pattern. The two meta-tools break that pattern with brain_find_tool and brain_call_tool, so the set is mostly consistent but not fully uniform.

Tool Count5/5

Nine tools is well within the ideal range, and each tool maps to a distinct session lifecycle action or a general tool-search/call capability. Nothing feels redundant or bloated.

Completeness5/5

The session lifecycle is well covered: start, resume, heartbeat, capture, end, abandon, and list, with artifact attribution handled. Adding the meta find/call tools rounds out the surface with no obvious dead ends.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Provides persistent memory for AI coding agents via MCP, enabling teams to share and recall facts across sessions. Automatically captures, classifies, and curates knowledge from supported transcript sources.
    18
    18 npm
    1
    AGPL 3.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Provides persistent, searchable memory across coding projects and machines, letting agents record and retrieve projects, reusable assets, sessions, decisions, commits, and handoffs via MCP.
    2
    Apache 2.0