ellmos-homebase-mcp
OfficialLocal-first MCP server for LLM agent memory, knowledge, state, routing, and orchestration planning — all persisted in local SQLite, no cloud egress.
Memory (
hb_mem_*) — store facts/lessons/working memories with confidence andagent_id, keyword search, build compact prompt-injection context, merge duplicates, and decay/prune low-confidence entries (dry-run by default).Knowledge base (
hb_kb_*) — ingest text/URL/file content, full-text search, fetch individual entries with metadata, list tags.Persistent state & tasks (
hb_state_*) — key/value state memory, create/list/update tasks with priority and status, plus connector-dispatch status.Model routing (
hb_route_*) — recommend the best model/provider for a prompt (credential-free), rate responses to feed an epsilon-greedy learning loop, view routing stats.Swarm patterns (
hb_swarm_*) — plan parallel chunking, consensus voting, boss-worker hierarchies, and stigmergy/blackboard coordination.Garden store (
hb_garden_*) — search/get/put small entries and run a stored command only if execution is explicitly enabled.API probing (
hb_api_*) — probe URLs via OpenAPI/wordlist/pattern/HATEOAS strategies, auto-detect schemas, export results as Markdown/JSON, browse probe history.Self-tests (
hb_test_*) — list, run, and review built-in metadata/smoke test batteries.Connectors (
hb_conn_*) — list connectors, queue outbound messages and read inbox messages locally without any network delivery.Automation (
hb_auto_*) — list local chain definitions, queue plan-only runs, check run status and results without external execution.Plugins (
hb_plug_*) — discover local plugins, inspect metadata, record dry-runs without executing plugin code.Canonical read-only seams (
hb_policy_*,hb_ticket_*,hb_lock_*) — resolve/list policy rules, list/show tickets by lifecycle category, and check active locks before writing to a project directory (fail closed if the canonical engine is unreachable).
Primary target of the server: it is designed for local, offline LLM orchestration against locally-hosted models such as Ollama (also Qwen/Llama-class models), providing a model router, credential-free API probing, swarm patterns, team memory with agent provenance, and knowledge/state tools that operate entirely locally with zero cloud egress.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ellmos-homebase-mcpRemember that I prefer dark mode for all UIs."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ellmos-homebase-mcp
Alpha MCP server for local-first LLM orchestration: memory, knowledge, routing, swarm patterns, API probing, persistent state, tests, automation planning, and plugin discovery in one stdio server.
Homebase is designed primarily for local LLMs (Ollama, Qwen, Llama, or any locally-hosted model via a MCP-capable harness). All persistent storage uses SQLite with no cloud dependency. External LLM providers (Claude, Codex, Gemini, OpenAI) can also connect as MCP clients, but local, offline-capable setups are the primary target.
German README: README_de.md
Part of the ellmos-ai family under the open-bricks umbrella.
Discoverability: Published on npm as ellmos-homebase-mcp and maintained in the ellmos-ai organization.
For AI Assistants & LLM Agents: Machine-readable architecture summary, index, and tool capabilities are published in llms.txt. MCP registry metadata is available in server.json.
Quick Navigation / Schnellnavigation
System Architecture (
#sec-01)Sequence Flow & Lifecycle (
#sec-02)Core Capabilities & Security Invariants (
#sec-03)Governance & Runtime Invariants (
#sec-04)Target Personas & Discoverability (
#sec-05)Comparative Matrix vs. Alternatives (
#sec-06)Start Here (
#sec-07)Status (
#sec-08)Install (
#sec-09)MCP Client Configuration (
#sec-10)Server Configuration (
#sec-11)Tools (
#sec-12)Discovery Context (
#sec-13)ellmos-ai Ecosystem (
#sec-14)Security & Vulnerability Reporting (
#sec-16)Development (
#sec-17)License & Statutory Liability Disclaimer (§ 521 BGB) (
#sec-18)Marketing Log (MARKETING-LOG.txt) | Changelog (CHANGELOG.md) | Legal Attribution (NOTICE) | Deutsche Version (README_de.md)
Related MCP server: nuzo-memory
System Architecture
flowchart TD
subgraph Clients ["MCP Clients (Local / Remote)"]
Ollama["Local LLMs (Ollama, Qwen, Llama)"]
Claude["Claude Code / Desktop"]
Codex["Codex / Antigravity"]
end
subgraph Transport ["Transport Layer"]
Stdio["stdio (Python MCP SDK)"]
end
subgraph Core ["ellmos-homebase-mcp Core Engine"]
Server["homebase.server"]
Config["homebase.config"]
end
subgraph ToolGroups ["51 MCP Tools across 14 Functional Modules"]
Mem["hb_mem_* (SQLite Memory)"]
KB["hb_kb_* (Knowledge Digest)"]
State["hb_state_* (State & Tasks)"]
Route["hb_route_* (Model Router)"]
Swarm["hb_swarm_* (Swarm Patterns)"]
Api["hb_api_* (API Probing)"]
Conn["hb_conn_* (Connectors Queue)"]
Auto["hb_auto_* (Automation Chains)"]
Plug["hb_plug_* (Plugin Discovery)"]
Garden["hb_garden_* (Garden Store)"]
Test["hb_test_* (Self Tests)"]
Policy["hb_policy_* (Policy Registry, read-only)"]
Ticket["hb_ticket_* (Ticket Master, read-only)"]
Lock["hb_lock_* (Lock Master, read-only)"]
end
subgraph Storage ["Local Storage (Offline-First)"]
DB[(SQLite Storage ~/.homebase/)]
end
Clients --> Stdio
Stdio --> Server
Server --> Config
Server --> ToolGroups
ToolGroups --> DBFour-View Architectural Topology Projection
+-------------------------------------------------------------------------------+
| VIEW 1: CALLER RUNTIMES, AGENT CLIENTS & ENTRYPOINTS |
| - Local LLM Engines: Ollama (Qwen, Llama, Mistral, DeepSeek), Local Harnesses|
| - Multi-Agent Orchestrators: Claude Code, Claude Desktop, OpenAI Codex, AGY |
| - IDE & Extension Interfaces: Cursor, VS Code MCP Extension, Windsurf |
| - stdio Protocol Transport: JSON-RPC 2.0 via Python Model Context Protocol |
+---------------------------------------+---------------------------------------+
| JSON-RPC 2.0 stdio (tools/list, tools/call)
v
+-------------------------------------------------------------------------------+
| VIEW 2: HOMEBASE MCP SOVEREIGN CORE & DISPATCH ORCHESTRATOR |
| - Core Server & Life Cycle: homebase.server (stdio loop, signal handling) |
| - Module Registry & Dispatch: homebase.registry (i18n schemas, locale norm) |
| - 51 Sovereign MCP Tools across 14 Specialized Functional Modules: |
| * Memory & Knowledge: hb_mem_* (SQLite memory), hb_kb_* (FTS5 search) |
| * State & Planning: hb_state_* (tasks & KV), hb_garden_* (garden store) |
| * Routing & Swarms: hb_route_* (offline routing), hb_swarm_* (blueprints) |
| * Exploration & Testing: hb_api_* (schema probe), hb_test_* (diagnostics) |
| * Integration & Staging: hb_conn_* (safe queues), hb_auto_* (chain plans) |
| * Extensibility: hb_plug_* (dry-run discovery, no remote code execution) |
| * Canonical Seams: hb_policy_*, hb_ticket_*, hb_lock_* (read-only views) |
+-------------------+-----------------------------------+-----------------------+
| |
v (bundled SQLite mode) v (canonical engine mode)
+---------------------------------------+ +-------------------------------------+
| VIEW 3: RUNTIME PERSISTENCE & | | VIEW 3-ALT: CANONICAL SEAMS |
| SQLITE STORAGE ENGINE | | (MODE-CONTRACT.md) |
| - Database: ~/.homebase/homebase.db | | - Policy Registry (hb_policy_*) |
| - Concurrency: Write-Ahead Log (WAL) | | - Ticket Master (hb_ticket_*) |
| - Multi-Agent Provenance: agent_id | | - Lock Master (hb_lock_*) |
| - Search Engine: SQLite FTS5 index | | - Fail-Closed Discipline: |
| - Integrity: Busy timeouts & rollback| | Raises CanonicalEngineUnavailable|
+-------------------+-------------------+ +-----------------+-------------------+
| |
+-------------------+-------------------+
|
v
+-------------------------------------------------------------------------------+
| VIEW 4: AIR-GAP DEFENSE PERIMETER, RUNASINVOKER & ZERO-EGRESS BOUNDARY |
| - 100% Local-First & Zero Egress: INV-LOCAL-01 (0 cloud calls, 0 telemetry) |
| - Unprivileged Execution: INV-PERM-08 (RunAsInvoker non-elevation principle) |
| - Engine Seam Integrity: INV-ENGINE-02 & INV-SEAM-03 (strict fail-closed) |
| - Deterministic Provenance: INV-PROV-04 (agent_id attribution on all state) |
| - Credential-Free Operation: INV-CRED-05 & INV-STAGE-06 (no tokens/secrets) |
| - Multi-Host Lock Defense: INV-SYNC-09 (.gitignore sync/lock immunity) |
| - Permissive Licensing: Zero-Copyleft stack (MIT, PSFL-2.0, Apache-2.0) |
| - Statutory SLA & Disclaimer: INV-SLA-10 (48h response SLA, § 521 BGB) |
+-------------------------------------------------------------------------------+Sequence Flow & Lifecycle
sequenceDiagram
autonumber
participant Client as MCP Client (Local LLM / Claude / Codex)
participant Stdio as Transport Layer (stdio)
participant Server as Server & Registry (homebase)
participant Module as Functional Module (hb_mem / hb_state / hb_route)
participant Engine as Engine Seam (Bundled vs Canonical)
participant DB as SQLite Storage (~/.homebase/)
Client->>Stdio: JSON-RPC 2.0 Request (tools/call: hb_mem_store, agent_id="agent-01")
Stdio->>Server: Decode & dispatch tool call
Server->>Module: Validate arguments & inject agent provenance
alt Bundled Engine Mode (Default)
Module->>DB: Execute SQLite query (WAL mode, busy timeout)
DB-->>Module: Return structured records / mutation status
else Canonical Engine Mode ([engines].mode = "canonical")
Module->>Engine: Seam check (Gardener / TASKPLAN / USMC)
alt Engine Available
Engine-->>Module: Delegate to canonical subsystem
else Engine Unreachable
Engine-->>Module: Raise CanonicalEngineUnavailable (Fail-Closed)
end
end
Module-->>Server: Format response in requested language (i18n: en/de/es/zh/ja/ru)
Server-->>Stdio: Encode JSON-RPC 2.0 Response
Stdio-->>Client: Result payload (Zero cloud egress, 100% local)Core Capabilities & Security Invariants
Capability / Invariant | Guarantee | Technical Implementation |
100% Local-First & Zero-Egress | Complete privacy and offline operation; no unexpected cloud communication or telemetry. | All persistent memory, knowledge, and state are saved in local SQLite ( |
Strict Engine Seams & Fail-Closed | No silent fallback into disconnected databases when requesting canonical systems. |
|
Team-Memory Provenance ( | Deterministic audit trail and filterable ownership for multi-agent workflows. | Native |
Credential-Free Discovery & Planning | Zero secret exposure during local routing recommendations and API probing. |
|
Safe Plan-and-Queue Adapters | Safe queueing and chain staging without arbitrary remote code execution. |
|
Full Native i18n Localization | Seamless multilingual developer and agent interaction. | Localized tool descriptions and JSON schemas for |
Non-Elevation & Secret Hygiene | Unprivileged execution and strict credential exclusion from distribution. | Non-root compatibility; live configs/secrets ignored in |
Multi-OS CI Smoke Integrity | Verified cross-platform reliability on all major operating systems. | Multi-version CI matrix covering Python 3.10–3.13 and Node.js 20–24 on Linux/Windows/macOS. |
Governance & Runtime Invariants
Invariant ID | Title & Scope | Guarantee & Technical Enforcement | Verification Seam |
| 100% Local-First & Zero-Egress | All persistent memory, knowledge entries, and task states are stored locally in SQLite ( |
|
| Strict Engine Seams & Fail-Closed |
|
|
| Canonical-Only Isolation |
|
|
| Deterministic Provenance & Team-Memory | Multi-agent coordination requires strict isolation. All memories, knowledge facts, and task transitions record |
|
| Credential-Free Discovery & Probing | Model routing suggestions ( |
|
| Plan-Only Staging & Bounded Offline Queues | Connector queues ( |
|
| Native Multilingual Schema Parity | All 51 tool definitions, input schemas, and validation errors maintain 100% complete localization across 6 supported languages ( |
|
| Non-Elevation & RunAsInvoker Principle | Homebase runs strictly in unprivileged user space. It requires no administrator or root privileges and ignores sensitive local dotfiles and system credentials. |
|
| Multi-Host Lock & Conflict Discipline | Strict exclusion of conflict copies ( |
|
| 48h Response, 5-Day Triage & 30-Day Remediation SLA | Security disclosures sent to |
|
Target Personas & Discoverability
Homebase is purpose-built to solve architectural and operational challenges across four core technical audiences:
[PERSONA-01] Local LLM & Edge AI Developers
Profile & Objective: AI engineers building offline or edge applications with Ollama, Qwen, or Llama models who need a robust orchestration harness.
Pain Points: Cloud memory APIs introduce unwanted latency, privacy leaks, subscription billing, and network failure modes.
Homebase Solution: Zero-cloud dependency, local SQLite WAL persistence (
~/.homebase/), and 51 standard stdio tools providing memory, FTS5 knowledge search, and task tracking.Reference Workflow:
{"tool": "hb_mem_store", "arguments": {"fact": "User prefers compact JSON output", "agent_id": "ollama-coder"}} {"tool": "hb_kb_search", "arguments": {"query": "API routing rules", "fts": true}}
[PERSONA-02] Multi-Agent Swarm Orchestrators & Swarm Architects
Profile & Objective: System architects orchestrating multi-agent collectives (Claude Code, Codex, Antigravity, local agents) operating concurrently on shared codebases.
Pain Points: State collisions, lack of origin tracking, race conditions in shared memory, and uncoordinated task delegation.
Homebase Solution: Native
agent_idprovenance across all facts, memories, and task states; built-in swarm templates (boss/worker, chunked parallel, consensus voting viahb_swarm_*).Reference Workflow:
{"tool": "hb_swarm_plan", "arguments": {"goal": "Audit security seams", "pattern": "consensus"}} {"tool": "hb_state_task_create", "arguments": {"title": "Verify fail-closed mode", "agent_id": "worker-audit-01"}}
[PERSONA-03] Enterprise Security & Data Governance Officers
Profile & Objective: CISOs, SecOps teams, and compliance auditors in regulated industries (healthcare, finance, defense) evaluating developer agent toolchains.
Pain Points: Silent cloud telemetry, unvetted remote side-effects, privilege escalation risks, and missing SLA assurances.
Homebase Solution: Strict zero-egress architecture, fail-closed canonical engine seams (
MODE-CONTRACT.md), unprivilegedRunAsInvokeroperation, and formal 48h Security Response SLA (SECURITY.md).Reference Workflow:
{"tool": "hb_policy_list_rules", "arguments": {}}Guaranteed fail-closed behavior: raises
CanonicalEngineUnavailableinstead of silently falling back to insecure stubs.
[PERSONA-04] Cross-Framework AI Assistants & Pair Programmers
Profile & Objective: Developers utilizing multiple AI coding assistants (Claude Desktop, Codex, Cursor, Gemini) seeking uniform context and tool parity across environments.
Pain Points: Incompatible custom tool APIs, fragmented scratchpads, and lack of multilingual developer schemas.
Homebase Solution: Standard stdio MCP transport, machine-readable project metadata (
llms.txt,server.json,glama.json), and 100% complete schema localization across 6 languages (en,de,es,zh,ja,ru).Reference Workflow:
{"tool": "hb_ticket_list", "arguments": {"folder": "ACTIVE"}}
High-Intent Search & SEO Keywords
English Intent:
local-first LLM orchestration MCP server,offline agent memory SQLite WAL,stdio Model Context Protocol Ollama Qwen,multi-agent swarm planning persistent state,zero-egress MCP server enterprise AI,fail-closed engine seams MODE-CONTRACT,team-memory agent_id provenance.German Intent:
Local-First LLM-Orchestrierung MCP-Server,Offline Agenten-Memory SQLite WAL,Model Context Protocol Stdio-Server Ollama,Multi-Agenten Schwarmplanung persistenter Zustand,Zero-Egress MCP-Server Unternehmens-KI,Fail-Closed Schnittstellen MODE-CONTRACT,Team-Memory Agenten-Provenienz.
Comparative Matrix vs. Alternatives
Homebase provides a uniquely comprehensive, local-first MCP capability stack compared to specialized or cloud-bound alternatives:
Architectural & Runtime Dimension |
| Cloud Memory SaaS (Letta, Pinecone, LangSmith) | Generic Memory MCPs (mcp-server-memory, sqlite) | Heavy Agent Frameworks (CrewAI, AutoGen, LangGraph) | Ad-Hoc Scripts / Custom SQLite |
1. 100% Local-First & Zero Egress ( | Yes (100% local SQLite WAL, zero telemetry) | No (Cloud-hosted, mandatory egress, PII risk) | Partial (Local file, but no strict egress contracts) | Variable (Often requires cloud API keys / SaaS) | Yes (Local, but no protocol guarantees) |
2. Engine Seams & Fail-Closed ( | Yes (Strict | No (Opaque cloud failovers) | No (Single hardcoded backend) | No (Unchecked exceptions / silent fallbacks) | No (Ad-hoc failure handling) |
3. Canonical-Only Seams ( | Yes ( | No (No canonical system awareness) | No (No policy or lock integration) | No (No governance seam layer) | No (Manual coordination) |
4. Team-Memory & Attribution ( | Yes (Native | Partial (User-level only, lacks multi-agent filters) | No (Single global unpartitioned graph) | Partial (In-memory agent state, lost on restart) | No (Manual schema management) |
5. Credential-Free Discovery ( | Yes (Offline routing & swarm planning without tokens) | No (Requires active paid cloud credentials) | No (No model routing or swarm tools) | No (Requires API keys for LLM planners) | No (No structured planning) |
6. Plan-Only Staging Queues ( | Yes (Safe connector queues & dry-run automation) | No (Direct execution or none) | No (No connector or automation support) | No (Direct runtime side-effects) | No (Unsafe arbitrary execution) |
7. Tool Breadth & Surface | 51 Tools across 14 Modules in single stdio server | 1-5 API endpoints | 2-5 basic tools | Framework-level Python library (not MCP native) | Fragmented CLI utilities |
8. Multilingual Schema Parity ( | Yes (Full en, de, es, zh, ja, ru schema coverage) | English only | English only | English only | English only / None |
9. Non-Elevation Security ( | Yes (Unprivileged RunAsInvoker, dotfile defense) | Cloud SaaS (Tenant-isolation trust model) | Variable (Local file permissions) | Variable (Often runs in root containers) | Variable (User scripts) |
10. Security Response SLA ( | Yes (Formal 48h Response, 5d Triage & 30d Remediation in | Commercial SLA (Paid tiers only) | None / Best-effort community | None / Best-effort community | None |
Start Here
Need | Entry point |
Install the alpha MCP server |
|
Run from a source checkout |
|
Configure a local LLM harness, Claude Code, Codex, or any MCP client | |
Inspect the machine-readable project summary | |
Check registry metadata |
Status
Transport: stdio via the Python MCP SDK
Package status: public alpha package under
ellmos-aiRelease metadata: MIT
LICENSE,NOTICE,CHANGELOG.md,llms.txt, and MCP Registry metadata inserver.jsonTest gate: GitHub Actions covers Python 3.10/3.11/3.12/3.13 plus Node.js 20/22/24 smoke and npm package checks
Current core: module discovery, MCP tool listing, MCP tool dispatch, config fallbacks, local planning/probing/queue/dry-run adapters
Real local SQLite modules:
hb_mem_*,hb_kb_*,hb_garden_*,hb_state_*Engine seams:
hb_garden_*,hb_state_task_*andhb_mem_*can delegate to the real canonical Gardener/Rinnsal/USMC engines instead of the bundled SQLite copies via[engines].mode = "canonical"(default remains"bundled"for a zero-dependency install). No silent fallback: if you requestcanonicaland the engine is unreachable, those tools return an error rather than quietly using the bundled DB — the server still starts and lists its tools. Binding rule and migration notes: MODE-CONTRACT.md; mechanism: KONZEPT.md.Canonical-only seams (no bundled alternative at all):
hb_policy_*(policy-registry),hb_ticket_*(ticket-master),hb_lock_*(lock-master) — all read-only in v1. A locally faked copy of live policy/ticket/lock state would mislead rather than help, so these three always attempt the canonical module and fail closed unconditionally if it is unreachable.Team-memory basics:
agent_idprovenance and filters for memory, knowledge, state memory, and tasks; SQLite uses WAL plus a busy timeout for safer concurrent agentsCredential-free alpha adapters:
hb_route_*,hb_swarm_*,hb_api_*,hb_test_*,hb_conn_*,hb_auto_*,hb_plug_*i18n: fully localized MCP tool descriptions, input-schema field descriptions, and unknown-tool errors for
en,de,es,zh,ja,ru(English fallback for any unset key)Roadmap: optional real LLM/API integrations and explicit execution backends
Install
The npm package contains a Node wrapper that starts the Python server. You still need Python 3.10+ and the Python package mcp>=1.0.0.
Option 1: Install From npm
npm install -g ellmos-homebase-mcp@alpha
ellmos-homebaseOption 2: Install From Source
git clone https://github.com/ellmos-ai/ellmos-homebase-mcp.git
cd ellmos-homebase-mcp
$env:PYTHONIOENCODING = "utf-8"
python -m pip install -e ".[dev]"
python -m pytest -ra -vAvoid creating a .venv inside cloud-synced folders if your sync client locks files. If you need an isolated environment, create it outside that folder.
Start From Source
$env:PYTHONPATH = "src"
python -m homebase.serverMCP Client Configuration
Homebase uses the standard stdio mcpServers configuration format. The same snippet works in any MCP-capable client or harness: BACH/Buddha (local Ollama), Claude Code, Codex, Cursor, or any other MCP host.
Note on local LLMs: A bare Ollama instance does not speak MCP natively — you need a MCP-capable harness on top of it (e.g., BACH, an open-source MCP proxy, or another orchestration layer). Configure that harness to include Homebase as an MCP server using the snippet below.
Global npm Install
{
"mcpServers": {
"homebase": {
"command": "ellmos-homebase"
}
}
}Source Checkout
{
"mcpServers": {
"homebase": {
"command": "python",
"args": ["-m", "homebase.server"],
"env": {
"PYTHONPATH": "/absolute/path/to/ellmos-homebase-mcp/src"
}
}
}
}Replace /absolute/path/to/ellmos-homebase-mcp with your local checkout path.
Server Configuration
Example: config/homebase.example.toml
Machine-readable project context: llms.txt
MCP Registry metadata: server.json
Default paths:
%USERPROFILE%\.homebase\homebase.toml%USERPROFILE%\.config\homebase\homebase.tomloverride with
HOMEBASE_CONFIG
Language can be configured with [server].language, HOMEBASE_LANG, or HOMEBASE_LOCALE.
The writing agent can be passed per tool call as agent_id; otherwise modules use
HOMEBASE_AGENT_ID, AGENT_ID, a module-level agent_id, or unknown.
[server]
name = "ellmos-homebase"
language = "en" # en, de, es, zh, ja, ru
[modules]
enabled = ["mem", "route", "kb", "swarm", "state", "garden", "api", "test", "conn", "auto", "plug"]Modules with missing optional dependencies are skipped without blocking server startup.
Tools
Important tool groups:
hb_mem_*for SQLite-backed memoryhb_kb_*for SQLite-backed knowledge entrieshb_state_*for persistent SQLite state and taskshb_garden_*for a small SQLite garden storehb_route_*for credential-free model-routing recommendations and feedback statshb_swarm_*for credential-free swarm planning patternshb_api_*for passive HTTP API discovery with SQLite historyhb_test_*for built-in metadata and smoke self-testshb_conn_*for a local connector registry plus SQLite-backed inbox/outbox queues without network sendshb_auto_*for local automation chain definitions and queued plan-only runs without backend executionhb_plug_*for local plugin discovery and dry-run records without executing plugin codehb_policy_*(read-only, canonical-only) for resolving/listing policy-registry ruleshb_ticket_*(read-only, canonical-only) for listing/showing ticket-master tickets by lifecycle folderhb_lock_*(read-only, canonical-only) for checking/listing active lock-master locks
Discovery Context
Use ellmos-homebase-mcp when searching for a local-first, offline-capable MCP server that gives local LLMs (Ollama, Qwen, Llama, or similar) persistent memory, knowledge management, routing, and orchestration — without requiring any cloud dependency. External LLM providers can also use it as an MCP server, but local-first setups are the primary design target.
Good search phrases:
ellmos Homebase MCP serverlocal-first LLM orchestration MCPMCP server SQLite memory knowledge routingoffline agent orchestration MCP serverMCP swarm planning persistent state API discovery
Not the same as Elmo/ELMO voice tools, AllenAI ELMo embeddings, Eclipse LMOS, generic cloud agent platforms, or single-purpose MCP memory servers.
ellmos-ai Ecosystem
This MCP server is part of the ellmos-ai ecosystem — AI infrastructure, MCP servers, and intelligent tools.
MCP Server Family
Server | Tools | Focus | npm |
47 | Filesystem, process management, interactive sessions, cloud-lock-safe operations | ||
23 | Code analysis, JSON repair, imports, diffs, regex | ||
12 | File repair, format conversion, batch operations | ||
19 | n8n workflow management via AI assistants | ||
20 | MCP stack discovery, profile management, control plane | ||
51 | Local-first LLM memory, knowledge, state, routing, swarm orchestration |
| |
8 | Server operations: health checks, log analysis, deploy dry-runs, mail diagnostics |
| |
3 | Headless Blender asset QA and FBX reimport verification |
| |
10 | Model-agnostic computer use: capture, safety-gated actions, Windows UIA |
|
AI Infrastructure
Project | Description |
Local-first text-based OS for LLM agents — 113+ handlers, 550+ tools, SQLite memory | |
Model-agnostic computer-use core powering Open Compute MCP | |
Provider-neutral LLM orchestration with auto-routing and budget tracking | |
Lightweight agent memory, connectors, and automation infrastructure | |
Self-hosted AI research stack (Ollama + n8n + Rinnsal + KnowledgeDigest) | |
Autonomous agent chain framework for Claude Code | |
Minimalist database-driven LLM OS prototype (4 functions, 1 table) | |
Testing framework for LLM operating systems (7 dimensions) |
Desktop Software & Sibling Ecosystem
Our partner umbrella organization open-bricks and sister organizations maintain local-first, privacy-centric desktop software and developer tools:
Application / Tool | Organization | Focus & Integration |
| Local-first desktop file organizer and PII-safe workspace exchange | |
| Distraction-free Markdown & PDF documentation manager | |
| Local-first PDF OCR and text layer embedding | |
| Offline document summarization and embedding engine | |
| Developer workspace hub and multi-repository management | |
| Isolated sandbox runner and local code execution assistant | |
| Hook-based LLM memory provenance and session injection gate | |
| Zero-dependency SQLite schema migration & replication layer |
Third-Party Licenses & Level 1 SBOM
ellmos-homebase-mcp is verified to contain 0% copyleft dependencies. All runtime dependencies are permissively licensed (MIT, BSD-2-Clause, Apache-2.0, PSFL).
Full inventory, Level 1 SBOM Invariant Cross-Reference Matrix, and non-elevation certifications are documented in THIRD_PARTY_LICENSES.md and plain-text companion THIRD_PARTY_LICENSES.txt. Canonical copyright and author attribution is maintained in NOTICE.
Security & Vulnerability Reporting
ellmos-homebase-mcp strictly adheres to local-first, zero-egress, and non-elevation security principles. Full policies, SLAs, and security guarantees are documented in SECURITY.md:
Supported Versions:
0.1.0-alpha.xResponse SLA: Initial acknowledgment and triage within 48 hours. Detailed triage within 5 business days; remediation within 30 calendar days.
Security Contacts:
security@ellmos.ai,support@lukasgeiger.com, andsecurity@open-bricks.org.Private Advisory: GitHub Security Advisories.
Development
$env:PYTHONIOENCODING = "utf-8"
$env:PYTHONDONTWRITEBYTECODE = "1"
python -m pytest -ra -v
npm run smoke
npm pack --dry-run --jsonNext useful step: add optional execution backends behind explicit configuration.
License & Statutory Liability Disclaimer (§ 521 BGB)
Software License
ellmos-homebase-mcp is open-source software licensed under the MIT License.
Canonical attribution to Lukas Geiger, the ellmos-ai family, and the open-bricks ecosystem is formally preserved in NOTICE.
Third-party component licenses are cataloged in THIRD_PARTY_LICENSES.md.
Statutory Notice & Liability Limitation (§ 521 BGB - German Law)
This software is made available free of charge as an open-source project. Under German statutory law governing gratuitous software provision (§ 521 BGB Gefälligkeitsrecht):
Liability Limitation: The author and contributors are liable only in cases of intentional misconduct (Vorsatz) or gross negligence (grobe Fahrlässigkeit).
Warranty Limitation: In accordance with §§ 523, 524 BGB, warranty claims for material and legal defects (Sach- und Rechtsmängel) are excluded, except in cases where defects have been fraudulently concealed (arglistiges Verschweigen).
Local-First & Non-Elevation Principle:
ellmos-homebase-mcpis provided on an "as is" and "as available" basis without any express or implied warranty. Operators run Homebase in unprivileged user mode (RunAsInvoker) at their own discretion.
Coordinated Security Response SLA
For vulnerability reporting or security inquiries, our coordinated disclosure policy guarantees an initial response within 48 hours and triage within 5 business days:
Security Contact:
security@ellmos.ai|support@lukasgeiger.com|security@open-bricks.orgAdvisory Portal: GitHub Security Advisories
Policy Documentation:
SECURITY.md
Bundles and partners
Homebase MCP remains a standalone local-first MCP server. In the V4
composition it is an optional MCP access surface of the
ellmos-memory-human-context-bundle: a configured system may use it to reach
memory and human-context capabilities. This access role does not make Homebase
the canonical owner of every memory, knowledge, state, routing or automation
function; the selected host and system manifests retain those bindings.
Canonical or bundled engines are integration partners selected by explicit configuration, not implicit replacements for this server. Authoritative bundle membership, versions, profiles and private composition recipes remain in the corresponding bundle manifests. This public section is discovery-only.
Available Tools
51 toolshb_api_discoverC
Auto-detect API schema from a base URL
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to inspect. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Auto-detect' implies an outbound network call to a user-supplied URL, but the description does not disclose whether it performs a live request, what it does on failure, whether authentication is needed, or what format the detected schema takes. With zero annotation coverage this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, appropriate for a one-parameter tool. It is efficient, though the terseness leaves it short of the 'complete' end of the scale.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and many similarly named siblings, the description is too thin. It does not explain return shape, error behavior, or how it differs from hb_plug_discover, leaving key agent decisions ungrounded.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (the single url parameter is documented), so the baseline is 3. The description adds no meaning beyond the schema, e.g. whether the URL must be reachable, include a scheme, or point to an OpenAPI/GraphQL endpoint.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Auto-detect API schema') and its input ('from a base URL'), which is clear enough on its own. However, it does not distinguish this tool from the closely named sibling hb_plug_discover or from hb_api_probe/hb_api_export, all of which appear to touch API-schema concepts. Without that differentiation an agent must infer which discovery tool to use.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to use hb_api_discover versus hb_plug_discover, hb_api_probe or hb_api_export. No prerequisites, exclusions, or context are given, so the agent gets no routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_api_exportC
Export probe results as Markdown or JSON
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Output format. | |
| probe_id | Yes | Probe result ID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full disclosure burden. 'Export' implies a read operation, but it never says whether output is returned inline or written to a file, whether permissions are needed, or how large a result set is handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the format options come last where they belong. It is perhaps terse to the point of under-specification, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter export tool with a fully documented schema and no output schema, the definition is minimally viable. It leaves open the practical question of where the exported content goes, which matters for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with only 2 params, one enum-constrained, so the schema already documents both fields adequately. The description adds nothing beyond restating the format options, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Export) and resource (probe results) plus the output formats. It is clearly distinguishable from related siblings like hb_api_probe or hb_api_history, though it does not name them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus hb_api_history or hb_test_results, nor any prerequisites such as needing a prior probe run. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_api_historyC
List previous probe results
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing about ordering (newest first?), retention window, scope (per endpoint, per session, global), or that it is a safe read-only listing. Only the 'previous' qualifier hints at recency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single four-word sentence, front-loaded with the verb and with zero filler. It is efficient, though the extreme brevity leaves the definition thin rather than structurally rich.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only list tool with no output schema and no annotations, this is the minimum viable description. An agent still lacks return-shape, ordering, and scoping information that cannot be recovered from the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'limit' parameter is already fully documented in the schema, so the baseline of 3 applies. The description adds no extra meaning such as default behavior when results exceed the limit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource ('List previous probe results') that an agent can distinguish from the sibling hb_api_probe, which presumably executes probes rather than retrieving history. However, it never explicitly contrasts itself with that sibling the way a 5-rated definition would.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no mention of prerequisites (e.g. that results come from hb_api_probe runs), and no alternatives named. The connection to probing is only inferable from the shared 'api' prefix.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_api_probeC
Probe a URL using all strategies (OpenAPI, wordlist, pattern, HATEOAS)
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to inspect. | |
| strategies | No | Discovery strategies to use. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it only says 'Probe'. It does not disclose whether this makes network requests, requires auth, is read-only or destructive, rate limits, or what it returns. 'Using all strategies' hints at exhaustive behavior but nothing concrete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste. It is efficient, though it is arguably too terse for an unannotated tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a network-probing tool with zero annotations, no output schema, and an unclear relationship to hb_api_discover, the description is far too thin. It omits safety profile, auth needs, return shape, and when it differs from its sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both 'url' and 'strategies'. The description adds the names of the four default strategies, which is slightly useful context, but no format or interaction details beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Probe) and resource (URL), and enumerates the strategies applied. It is reasonably distinct from siblings like hb_api_discover, but the description never explains how 'probe' relates to 'discover', leaving the agent to guess whether they overlap or differ in scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of the clearly related sibling hb_api_discover. The agent has no basis for choosing this tool over hb_api_discover or the other hb_api_* tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_auto_list_chainsB
List available local automation chain definitions
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no read-only confirmation, no return shape, no filtering/pagination, no indication of what 'available' or 'local' means at runtime. 'List' weakly implies a non-mutating read, but that is an inference, not a disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler; the verb and the scoping adjective ('local') come first. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool this is close to sufficient, but with no output schema the description should say something about what a caller gets back (chain identifiers, definitions, ordering) to be fully actionable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the baseline there is nothing for the description to disambiguate. Schema coverage is 100% and there are no arguments to misunderstand.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('local automation chain definitions'), making the intent unambiguous. It does not explicitly distinguish itself from siblings like hb_auto_run or hb_auto_status, but the naming plus 'definitions' vs 'run'/'status' makes the distinction inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives such as hb_auto_run, hb_auto_status, or hb_auto_result. 'List available' weakly implies a discovery step before running a chain, but nothing is spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_auto_resultC
Get the recorded local automation run result
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | Automation run ID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It implies retrieval of an already-recorded run ('recorded') but says nothing about what happens with an unknown run_id, whether the result is only available after completion, or what the returned structure contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words. Efficient, though extremely terse for a tool with no supporting annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a simple single-parameter getter, and the name plus description convey its core purpose. However, with no output schema and no annotations, an agent has no signal about the return shape or failure modes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single required run_id documented as 'Automation run ID.' The description adds no format, sourcing, or example detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (recorded local automation run result), making it clearly distinct from sibling result/production tools like hb_auto_run and hb_auto_status. It does not explicitly name its siblings, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to call this versus hb_auto_status or hb_auto_list_chains, and no prerequisites stated. The agent must infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_auto_runB
Queue a local automation chain plan without executing external backends
| Name | Required | Description | Default |
|---|---|---|---|
| chain | Yes | Automation chain name. | |
| input | No | Input payload for the operation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one meaningful trait: the call queues rather than executes external backends, implying no external side effects. It falls short on auth requirements, whether the named chain must pre-exist, and how/when effects surface — real gaps for a queue-style mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though the terse phrasing leaves the queue-vs-execute nuance slightly opaque.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a queueing operation with no output schema and no annotations, the description should at least hint at follow-up (e.g., hb_auto_status or hb_auto_result to observe progress). It also never resolves the name/behavior tension between 'run' and 'queue', leaving the agent unsure whether this starts work or merely registers it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents 'chain' and 'input' adequately. The description adds no syntax, format, or constraint detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Queue') and resource ('local automation chain plan') and clarifies scope ('without executing external backends'). This distinguishes it well from hb_auto_list_chains/hb_auto_status/hb_auto_result within the hb_auto family, though the 'run' in the name sits oddly against 'queue'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this versus siblings like hb_plug_run, hb_garden_run, or the other hb_auto tools. The 'without executing external backends' phrase implies a deferred-execution scenario but does not state when to prefer it or what prerequisites exist.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_auto_statusB
Check local automation run status
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | Automation run ID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden. 'Check' implies a read-only operation, but this is not explicitly stated, and there is no mention of side effects, authentication, rate limits, or return behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded phrase with no wasted words. It is appropriately sized for a simple status-check tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity and the lack of an output schema, the description is minimally adequate. However, it does not clarify what status values are returned, what 'local' means, or confirm the operation is read-only, leaving some gaps for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single parameter (run_id) is already documented in the schema. The description adds no additional meaning, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Check') and resource ('local automation run status'), making it clear what the tool does. It does not explicitly differentiate from siblings like hb_auto_result or hb_auto_run, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives such as hb_auto_result or hb_auto_list_chains. Usage is only implied by the name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_conn_listB
List configured connectors
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no output schema, the description carries the full disclosure burden but says nothing about return shape, pagination, filtering scope, or whether listing triggers any side effects. A read-only enumeration is implied by 'List' but never confirmed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single three-word phrase with no waste is well front-loaded, though it is arguably under-specified rather than concise. It earns its place but adds little beyond the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool with no output schema the description is minimally viable, but an agent still lacks any sense of what a connector record contains or how results are scoped. It does not misuse complexity, but it leaves obvious questions open.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no semantics to explain beyond what the schema provides. Baseline of 4 applies for a parameterless definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (List) and resource (configured connectors), so an agent can tell it apart from the action-oriented siblings hb_conn_send/hb_conn_receive and even the state-oriented hb_conn_status. It stops short of clarifying what a 'connector' is or what fields are returned, so it is clear but not fully disambiguating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus hb_conn_status or the other conn_* siblings. The agent must infer that this is the enumeration entry point for connectors purely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_conn_receiveB
Get recent local inbox messages for a connector
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return. | |
| connector | Yes | Connector name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It says 'Get' and 'recent local inbox messages' but does not clarify whether receiving consumes or deletes messages, whether it marks them as read, whether the connector must exist, or any auth or rate-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It states the action, scope, and target directly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and the schema covers both parameters, but with no annotations and no output schema, the description should say more about what 'receive' means behaviorally. It is adequate for basic selection but leaves gaps around return format and side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both the connector and limit parameters. The description adds the scoping phrase 'for a connector' and 'recent' implies limiting results, but it does not add syntax, format, or default details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (recent local inbox messages) with scope (for a connector). It is clearly different from sending or listing connectors, but it does not name or distinguish itself from siblings like hb_conn_list or hb_conn_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies retrieval usage but gives no explicit when-to-use guidance, no when-not-to-use conditions, and no alternatives. It does not mention hb_conn_send or other connector tools that an agent might confuse it with.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_conn_sendB
Queue a message for a connector without network delivery
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Connector target, channel, or chat ID. | |
| message | Yes | Message text. | |
| connector | Yes | Connector name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the core side effect (queued, not sent over the network), but says nothing about persistence, auth requirements, whether queued messages are later delivered, or what happens on duplicates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler. The key behavioral constraint (queue, no network) appears immediately and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description should carry more behavioral weight for a mutation-style tool. It omits what the queued message becomes, the distinction between connector and target, and any return/confirmation semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents connector, message, and target. The description adds no syntax or format detail beyond what the schema supplies, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb (queue) and resource (a message for a connector) plus a distinguishing scope constraint: no network delivery. That separates it from siblings like hb_conn_receive or hb_conn_status, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without network delivery' hints at the use case, but there is no explicit when-to-use, when-not-to-use, or named alternative (e.g., vs hb_conn_receive). The agent must infer the routing decision on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_conn_statusB
Check local connector health and queue counts
| Name | Required | Description | Default |
|---|---|---|---|
| connector | Yes | Connector name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It reveals that the tool reports health and queue counts, implying a read-only operation, but does not state whether it requires authentication, has side effects, rate limits, or how errors are handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no filler and the core purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter status tool, the description is minimally adequate. However, it does not explain what 'health' entails or what the queue counts represent, and it lacks usage guidance, leaving some contextual gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single required 'connector' parameter, so the schema already provides adequate parameter semantics. The description adds no additional meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Check') and two specific resources ('local connector health' and 'queue counts'), making the tool's purpose clear. It does not explicitly differentiate from siblings like hb_conn_list or hb_conn_receive, but the status/health focus is distinctive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives. The implied usage is a status check, but no conditions, prerequisites, or alternative tools are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_garden_findC
Search the garden
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return. | |
| query | Yes | Search query. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the entire behavioral burden, and it discloses nothing: not whether this is a read-only operation, whether results are ranked, how pagination or the "limit" default behaves, or what permissions are needed. For a search tool with zero annotation coverage this is a complete gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three words with zero padding or repetition, so there is no bloat to penalize. However, it is a bare fragment with no structure, and the single phrase does not earn its place because it conveys nothing the tool name did not already convey.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description should at least characterize the search domain and result shape, and it does neither. The only fully specified surface is the two-parameter input schema, leaving the agent without enough context to invoke this confidently in a 50-tool environment.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — both "query" and "limit" (including its default of 10) are documented in the schema itself. The description adds no syntax, matching, or format guidance beyond that, so the baseline 3 applies where the schema does all the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Search the garden" does state a verb and a nominal resource, but "garden" is never defined and the phrase effectively restates the tool name hb_garden_find. With no differentiation from siblings like hb_kb_search, hb_garden_get, or hb_garden_run, an agent cannot tell what domain this searches or how it differs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative-tool guidance anywhere in the description. The verb "Search" faintly implies a retrieval use case, but nothing tells the agent when to prefer this over hb_kb_search or hb_garden_get.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_garden_getC
Get a specific entry by key
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Entry key. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not say whether the read is safe, what happens when the key is absent (error vs null), or whether the lookup is case-sensitive or namespaced — all meaningful for a retrieval tool with a free-form string key.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with zero filler — verb and lookup mechanism come first. It is efficient but so terse that there was no opportunity to front-load anything more useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, nothing in the structured data tells the agent what an entry contains or how failures surface. For a tool whose only behavior is returning a keyed object, the description should at least sketch the return shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is documented as "Entry key." The description adds no format, prefix, or case rules beyond that, so it sits at the baseline appropriate when the schema does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and the keyed lookup, which does contrast with hb_garden_find's implied search. However, the resource is only described as "a specific entry" with no indication of what a garden entry is or which namespace it lives in, so the agent gets the mechanic but not the domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus hb_garden_find (search) or hb_garden_put (write). The description implies read-by-key but never states the condition that selects this tool over its siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_garden_putC
Store or update an entry
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Entry key. | |
| value | Yes | Entry value. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does imply upsert semantics ('store or update'), which is a genuine behavioral trait, but says nothing about overwrite behavior on existing keys, permissions, durability, or what is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence that is front-loaded and waste-free, but it is thin rather than tight — the brevity comes from under-specification rather than disciplined editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description should at minimum clarify overwrite behavior, scope of the store, and error cases. None of that is present, leaving real gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the two parameters (key, value) are already documented in the schema. The description adds no format, length, or naming-convention details beyond that, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb phrase 'store or update' plus 'an entry' does communicate a key-value upsert, which implicitly distinguishes it from read-only siblings like hb_garden_get and hb_garden_find. However, 'entry' and the 'garden' namespace are never grounded in a domain, so the purpose is only vaguely conveyed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the many overlapping write/state siblings (hb_mem_store, hb_state_mem_set, hb_kb_ingest). No prerequisites, no exclusions, no mention of when an update is preferable to a create.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_garden_runC
Run a stored command if execution is explicitly enabled
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Entry key. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It hints that execution may be gated by a flag, but does not clarify what 'explicitly enabled' means, who can enable it, what happens when disabled (error vs no-op), or any side effects. For an execution tool this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence that front-loads the action and appends the condition. It is efficient, though the vagueness of 'stored command' stems from under-specification rather than verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given it is an execution tool with no annotations, no output schema, and only a one-word parameter description, the definition is under-specified. An agent lacks enough to know when to invoke it, what authority is required, or what a disabled state yields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% but the only parameter description is 'Entry key.', which is terse. The description adds no meaning about what the key references or how it maps to a stored command. Baseline 3 applies since schema covers the structure, but no extra value is added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a verb ('Run') and a resource ('stored command'), which is clearer than the bare name. But 'stored command' is vague and the description does not distinguish this tool from siblings like hb_plug_run, hb_test_run, or hb_auto_run. An agent cannot confidently tell when to select this over those.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is a conditional clause ('if execution is explicitly enabled'), but no guidance on which sibling to use, no prerequisites explained, and no explicit when-to-use versus when-not. The condition is stated without clarifying how or where execution is enabled.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_kb_getB
Get a single knowledge entry with metadata
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Numeric entry ID. | |
| agent_id | No | Optional agent filter |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. 'Get' clearly implies a non-mutating read, and 'with metadata' hints at the return shape, but there is no disclosure of auth requirements, behavior for missing IDs, or what 'metadata' includes. For a simple read the exposure is modest, but the gap is real.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words; the action and resource come first. It is efficient rather than padded, though it errs slightly toward under-specification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with no output schema and no annotations, the description is minimally sufficient but thin: the 'with metadata' phrase is the only clue about the response, and nothing covers the optional agent_id's effect or failure modes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with both 'id' and the optional 'agent_id' documented in the schema, so the baseline of 3 applies. The description adds no meaning about the ID format or what 'agent filter' actually scopes, so it neither helps nor hurts.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (a single knowledge entry), and the word 'single' implicitly distinguishes it from the sibling list/search tools hb_kb_list and hb_kb_search. It stops short of naming those siblings explicitly, so an agent must infer the retrieval-by-ID vs. query distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: fetching one entry by ID. There is no statement of when to prefer this over hb_kb_list or hb_kb_search, nor any prerequisite or error condition (e.g. non-existent ID). Adequate but leaves routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_kb_ingestC
Add new knowledge (text, URL, or file content)
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | List of tags. | |
| source | No | URL or file path | |
| content | Yes | Content text. | |
| agent_id | No | Agent identifier for shared-knowledge provenance |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full disclosure burden. It implies a persistent mutation but says nothing about indexing, deduplication, chunking, provenance handling, permissions, or whether the operation is reversible — all material for an ingest tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the core verb front-loaded and no wasted words. It is efficient, though the brevity edges toward under-specification rather than optimal density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With four parameters, no annotations, and no output schema, the description should at minimum indicate what ingestion produces (e.g., a knowledge ID) and where the content lands. It instead covers only the action and input formats, leaving the agent short of what it needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters; baseline is 3. The description's mention of 'URL or file content' loosely maps to the source/content fields but adds no format, size, or tagging syntax beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb and resource ('Add new knowledge') and enumerates the accepted content forms (text, URL, file content), so the action is unambiguous. It does not distinguish itself from adjacent write-oriented siblings such as hb_mem_store or hb_garden_put, which prevents a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to ingest into the knowledge base versus storing to memory (hb_mem_store) or the garden (hb_garden_put), nor any prerequisite or ordering advice. The agent is left to infer the appropriate context entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_kb_listC
List tags in the knowledge base
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Optional agent filter |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no return format, no pagination, no statement that it is a safe read, and no behavior when agent_id is omitted or matches nothing. Only a bare "list" implies read-only.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short, front-loaded sentence with no waste. It is efficient, though its brevity edges toward under-specification rather than genuine economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description leaves critical questions unanswered: what entity is actually listed, what the response contains, and whether results are paginated. An agent could invoke it but cannot predict its output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with a single optional agent_id documented as "Optional agent filter", so the schema does the work; the description adds no clarification of what the filter means (filter tags by agent?) or what happens without it. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a verb+resource ("List tags"), but the resource is ambiguous against the tool name hb_kb_list and the sibling set (hb_kb_search, hb_kb_get, hb_kb_ingest) – an agent cannot tell whether this lists tags, KB entries, or tag-filtered entries. It also doesn't distinguish itself from hb_kb_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no exclusions, and no mention of alternatives such as hb_kb_search for content lookup or hb_kb_get for a specific entry. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_kb_searchB
Full-text search in knowledge database
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return. | |
| query | Yes | Search query. | |
| agent_id | No | Optional agent filter | |
| category | No | Category filter or category to store. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It indicates a read-style search operation, but says nothing about result ranking, pagination, permissions, rate limits, or whether it returns snippets versus full documents.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. It immediately communicates the core operation and resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The schema fully covers the input parameters, and the tool is a relatively simple search operation. However, with no output schema and no annotations, the description does not explain the return format or result behavior, leaving some context for the agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters including query, limit, agent_id, and category. The description adds no parameter-level meaning beyond what is already in the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Full-text search') and resource ('knowledge database'), so the agent knows this is a search operation. It does not explicitly differentiate from sibling tools such as hb_kb_list or hb_kb_get, but the word 'search' distinguishes it enough to infer its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives like hb_kb_list for enumeration or hb_kb_get for direct retrieval. The agent must infer that it is for query-based search over the knowledge base.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_lock_checkA
Check whether a project directory currently has an active lock -- call this BEFORE any write action there (read-only).
| Name | Required | Description | Default |
|---|---|---|---|
| project_dir | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses that the operation is read-only and should precede writes, but omits authentication needs, return semantics, and failure behavior, leaving clear gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that states the action, scope, timing prerequisite, and read-only nature without any wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, one-parameter read tool with no output schema and no annotations, the description covers purpose, timing, and safety sufficiently. It could still say more about the return value or lock semantics, but the essential information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter project_dir is only named in the schema. The description adds that it refers to a project directory, which clarifies meaning, but does not specify path format, relative vs. absolute, or any other constraint.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: checking whether a project directory has an active lock. It does not explicitly distinguish itself from the sibling tool hb_lock_list, so it falls short of a 5 despite being clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit timing guidance: call this BEFORE any write action there. However, it does not mention alternatives or when-not to use it, such as hb_lock_list, so it is not fully complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_lock_listA
List all currently active locks across configured roots (read-only).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden: it usefully discloses that the operation is read-only and scoped to configured roots, which prevents an agent from expecting arbitrary paths. It says nothing about return shape, ordering, or whether stale locks are included.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence covering verb, resource, scope, and safety trait with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only listing tool with no output schema, the description covers what is listed, where it is scoped, and that it is safe. Only minor gaps remain (e.g., what fields each lock entry contains), which is acceptable at this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so no parameter-level explanation is required; baseline is 4. The description's mention of 'configured roots' correctly signals that scope is implicit rather than caller-supplied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (currently active locks) with scope (across configured roots). It does not explicitly differentiate from the sibling hb_lock_check, but the 'all currently active' phrasing largely conveys the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Use is implied by the listing semantics, but there is no explicit when-to-use guidance or routing hint against hb_lock_check (which presumably inspects a single lock). No exclusions or prerequisites are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_mem_consolidateA
Decay memory confidence and prune low-confidence entries (dry_run previews; dry_run=false applies)
| Name | Required | Description | Default |
|---|---|---|---|
| decay | No | Absolute confidence reduction per entry. | |
| dry_run | No | Preview without applying. Set false to apply. | |
| agent_id | No | Optional agent filter | |
| min_confidence | No | Entries below this (after decay) are pruned. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It does disclose the key behavioral trait: a mutation happens only when dry_run=false, otherwise it is a preview, and that entries get pruned. It stops short of stating irreversibility, whether pruning is permanent, or any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words; the dry_run safety note is appended efficiently. The nested parenthetical is slightly dense but earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-param mutation tool with no output schema and no annotations, the description covers purpose and the dry-run safety valve, and the schema fully documents parameters. Missing only irreversibility/scope caveats about what pruning destroys.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents decay, dry_run, agent_id, and min_confidence with bounds and defaults. The description only restates the dry_run semantics, adding little beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Decay memory confidence and prune low-confidence entries.' An agent can tell this is a memory-maintenance operation distinct from hb_mem_store or hb_mem_merge. However, it does not explicitly differentiate itself from the closest sibling (hb_mem_merge), which also mutates memory state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(dry_run previews; dry_run=false applies)' tells the agent how to invoke it safely, implying usage context. But it never states when this tool should be chosen over hb_mem_merge or other memory siblings, nor any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_mem_contextC
Generate compact context string for prompt injection
| Name | Required | Description | Default |
|---|---|---|---|
| focus | No | Optional topic focus | |
| agent_id | No | Optional agent filter | |
| max_tokens | No | Approximate token budget |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not state what data source the context is drawn from, whether it is read-only, whether repeated calls mutate or consume memory, or what happens when the token budget is exceeded.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no filler, and the artifact and its purpose are front-loaded. It is appropriately sized but arguably under-specified rather than over-worded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-parameter tool with no output schema, the description should explain the returned string's shape and the source of context. It conveys the output type but leaves provenance and budget behavior unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so focus, agent_id, and max_tokens are already documented in the schema. The description adds nothing about semantics such as how the token budget is honored or what focus filtering actually selects.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb (generate) plus specific artifact (compact context string) and intended use (prompt injection). It is distinguishable from hb_mem_store/hb_mem_query by producing a consumable string rather than storing or filtering records, though it never names those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to call this versus hb_mem_query or hb_state_mem_get, nor any prerequisite (e.g., existing stored memories) or ordering advice. The agent must infer the trigger condition from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_mem_mergeB
Confidence-based merge of duplicate memories (dry_run previews; dry_run=false applies the merge)
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | Preview merge without applying. Set false to apply. | |
| agent_id | No | Optional agent filter |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that dry_run previews and dry_run=false applies the merge, which is meaningful mutation context. However, it does not state irreversibility, what happens to merged memories, permission requirements, or conflict resolution beyond 'confidence-based'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with a parenthetical that covers both purpose and dry_run behavior. Every word earns its place and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter memory-mutation tool with no annotations and no output schema, the description is minimally viable. It covers the core action and dry_run semantics but omits sibling routing guidance and behavioral caveats such as irreversibility.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters fully. The description adds the 'confidence-based' framing but does not mention agent_id or add syntax/format detail beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and resource: 'merge of duplicate memories', and adds 'confidence-based' to clarify the mechanism. It does not distinguish this tool from the sibling hb_mem_consolidate, which also deals with memory consolidation, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains dry_run behavior but gives no guidance on when to choose this merge tool over alternatives like hb_mem_consolidate or hb_mem_query. The agent must infer the use case from the phrase 'duplicate memories' alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_mem_queryC
Search memory by keyword
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return. | |
| query | Yes | Search query | |
| agent_id | No | Optional agent filter | |
| category | No | Category filter or category to store. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it discloses almost nothing. "Search" implies a read-only operation, but there is no statement about result ordering, pagination via limit, category semantics, or what happens when no match exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single phrase is front-loaded and waste-free, which is good, but the brevity here reads as under-specification rather than discipline. A slightly longer statement could route usage without bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter tool with no annotations and no output schema, one line is insufficient. An agent lacks guidance on filtering by category or agent_id, how limit interacts with result ordering, and how this search relates to the other mem and kb siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented in the schema, making 3 the baseline. The word "keyword" hints at query matching style but adds no format or syntax detail beyond the schema's "Search query".
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Search memory by keyword" names a verb (search) and resource (memory), so the core purpose is legible. But with siblings like hb_mem_context, hb_mem_merge, hb_mem_consolidate, and hb_kb_search, a one-phrase description does nothing to distinguish which memory surface this targets or how it differs from context retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative routing is given. An agent cannot tell from this description whether to reach for hb_mem_query, hb_mem_context, or hb_kb_search first, and no prerequisites or result expectations are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_mem_storeC
Store a fact, lesson, or working memory entry
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | The content to store | |
| agent_id | No | Agent identifier for shared-memory provenance | |
| category | Yes | Memory category | |
| confidence | No | Confidence score (0-1) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it says nothing about persistence semantics, deduplication, whether entries are append-only or overwritable, or what provenance/agent_id actually does. 'Store' implies a mutation but no safety, auth, or reversibility context is given.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler or redundancy. It is efficient, though the terseness is partly what leaves the behavioral and usage gaps unfilled.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with four parameters, zero annotations, and no output schema, the description does not cover enough ground. It never addresses what happens on store, whether duplicates are handled, or what the caller gets back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented in the schema, establishing the baseline of 3. The description only re-lists the category values that the schema's enum already covers, adding no new meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Store') and the resource is clearly scoped by the enumeration of what may be stored ('fact, lesson, or working memory entry'). The verb alone distinguishes it from read-oriented siblings like hb_mem_query and hb_mem_context, though it never names those siblings explicitly to confirm the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to store versus query, merge, consolidate, or route memory, and no preconditions. An agent must infer the write-vs-read split purely from the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_plug_discoverC
Scan a local directory for plugin metadata
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Local filesystem path. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. 'Scan' implies a read-only local operation, but the description does not state permissions, side effects, recursion behavior, error handling, or output shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. It is appropriately concise, though extremely terse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple discovery tool with no output schema and no annotations, the description leaves key context unstated: default path behavior, what 'plugin metadata' includes, and whether the operation is read-only. These gaps matter for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single path parameter, which is fully documented as a local filesystem path. The description adds no further parameter meaning, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (Scan), resource (local directory), and target (plugin metadata). It is clear what the tool does, though it does not explicitly differentiate itself from siblings like hb_plug_list or hb_plug_info.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives, nor on prerequisites or exclusions. The description only states what it does, leaving usage entirely implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_plug_infoC
Get local plugin metadata
| Name | Required | Description | Default |
|---|---|---|---|
| plugin | Yes | Plugin name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read-only operation but does not disclose return format, permissions required, caching behavior, or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no wasted words. It is appropriately concise for a simple retrieve operation, though its terseness contributes to gaps elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description should explain what metadata is returned or at least characterize the return value. It also lacks basic usage context and behavioral detail, leaving the agent with minimal information beyond the tool name and schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single 'plugin' parameter is already well documented in the schema. The description adds no additional meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('local plugin metadata'), distinguishing it from siblings like hb_plug_list and hb_plug_run by its metadata focus. However, it does not explicitly name or contrast with any sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no guidance on when to use this tool versus hb_plug_list, hb_plug_discover, or hb_plug_run. Usage is only implied by the phrase 'Get local plugin metadata', with no exclusions or alternatives mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_plug_listB
List discovered local plugins
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It does not state whether the operation is read-only, whether it triggers or refreshes plugin discovery, whether results are cached, or what permissions are required. The minimal phrasing leaves key behavioral traits undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or repetition. Every word contributes to stating the tool's purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple listing tool, the description is minimally adequate but leaves important contextual gaps. Most notably, it does not clarify how it relates to hb_plug_discover or whether listing implies prior discovery, and absent an output schema it gives no hint of what is returned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so there are no parameter semantics to document. The baseline score of 4 applies for a zero-parameter tool whose schema also has nothing to clarify.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'List' plus 'local plugins'. It also scopes the resource to 'discovered' plugins, which helps distinguish it from raw discovery. However, it does not explicitly differentiate itself from siblings like hb_plug_discover or hb_plug_info.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus hb_plug_discover, hb_plug_info, or hb_plug_run. The intended usage is only implied by the word 'List' and the 'discovered' qualifier.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_plug_runB
Record a plugin dry-run without executing plugin code
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Optional argument object. | |
| plugin | Yes | Plugin name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that plugin code is not executed, which is important for safety and intent. It does not disclose persistence side effects, permissions, idempotency, or what 'record' entails beyond the non-execution guarantee.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It states the action and the key constraint immediately. It is appropriately sized for a two-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description must carry more context than usual. It covers the critical non-execution behavior but does not explain what the recorded dry-run returns, whether it persists state, or how failures are reported. It is adequate but leaves meaningful gaps for an agent invoking a write-like operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the input schema. The description adds no additional meaning about the plugin name or args object beyond what the schema provides. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb plus resource: 'Record a plugin dry-run.' It also clarifies the critical boundary that plugin code is not executed. It does not explicitly distinguish itself from siblings like hb_plug_list, hb_plug_info, or hb_garden_run, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without executing plugin code' implies this is for dry-run recording rather than real plugin execution. However, there is no explicit when-to-use guidance, no when-not-to-use guidance, and no alternative tool is named. The agent must infer the usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_policy_listC
List policy-registry entries (read-only, canonical-only), optionally filtered.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| query | No | Search query. | |
| scope | No | ||
| consumer | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses two behavioral traits—read-only and canonical-only (excluding non-canonical/overridden entries)—but says nothing about permissions, pagination, result count limits, or ordering for what is presumably a large registry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the core action front-loaded and the key constraint parenthesized. Nothing is wasted, though it is arguably too terse for a four-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With four parameters, 25% schema coverage, no annotations, and no output schema, the definition omits far too much: the meaning of kind/scope/consumer, the shape of returned entries, and any pagination behavior. An agent can invoke it only by guessing at filter semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (only 'query' is documented), so three of four parameters (kind, scope, consumer) are opaque. The description says entries can be filtered but never explains what kinds, scopes, or consumers are valid or how they combine, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List policy-registry entries'), so an agent knows this is a read/listing operation over the policy registry. It does not name or contrast with the sibling hb_policy_resolve, so sibling differentiation is missing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'optionally filtered' hints that filters are supported but gives no when-to-use guidance, no conditions, and no mention of the alternative hb_policy_resolve. The agent must infer that resolution (not listing) is the alternative path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_policy_resolveC
Resolve the authoritative policy/rule/decision for a scope via the real policy-registry engine (read-only, canonical-only).
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Search query. | |
| scope | Yes | ||
| consumer | No | ||
| required_kind | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It discloses that the operation is 'read-only' and 'canonical-only', which are useful safety and filtering traits, but it omits auth requirements, error behavior, and what 'resolve' actually returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is front-loaded with the core action and resource, and the parenthetical qualifiers are compact. It contains no filler, though its extreme brevity leaves much unsaid.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given four parameters, one required field, no annotations, no output schema, and low schema coverage, the description is far too sparse. It does not explain how to use scope, consumer, required_kind, or query, nor what a resolved result looks like.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (just the 'query' parameter). The description mentions 'for a scope' but adds no real meaning beyond the parameter name; 'consumer' and 'required_kind' are left completely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Resolve') and resource ('authoritative policy/rule/decision for a scope'), making the tool's purpose clear. It implicitly distinguishes itself from sibling 'hb_policy_list' by indicating resolution rather than listing, but does not explicitly name that alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as hb_policy_list. The description only states what it does, leaving usage context entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_route_evaluateC
Rate a response for routing feedback loop (epsilon-greedy learning)
| Name | Required | Description | Default |
|---|---|---|---|
| quality | Yes | Quality rating from 0 to 1. | |
| route_id | Yes | Routing decision ID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose a meaningful behavioral trait: this feeds a routing feedback loop / learning process, so an agent can infer it mutates routing state and affects future selections. It does not say whether the effect is immediate, persistent, or reversible, which leaves a gap for a write-like operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single, front-loaded sentence with no filler or repetition. It is efficient, though the brevity is partly why other dimensions are thin rather than a sign of strong structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-mutating learning tool with no annotations and no output schema, the description omits what the rating does, what is returned, and its prerequisite relationship to hb_route_select. The schema covers the inputs, but the behavioral contract an agent needs before invoking it is largely missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already defines both parameters (route_id, quality with a 0-1 range). The description adds no syntax, format, or constraint details beyond the schema. This meets the baseline 3 when structured fields do the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a verb ("Rate") and a target ("a response for routing feedback loop"), so the general purpose is inferable. However, "a response" is ambiguous and it does not explicitly distinguish itself from close siblings like hb_route_select or hb_route_stats. The purpose is vague-but-directional rather than specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit statement of when to call this tool, what must precede it (e.g., a prior hb_route_select), or which sibling to use instead. The "epsilon-greedy learning" parenthetical is context, not guidance. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_route_selectC
Analyze a prompt and recommend the best model/provider
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | The prompt to route | |
| constraints | No | Optional constraints (max_tokens, speed, cost) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not disclose whether the call is a read-only analysis, how constraints influence the recommendation, whether there are latency or rate-limit implications, or what triggers a reroute.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler or repetition. Appropriately sized for the tool's simplicity, though minimal enough that no structure beyond one clause exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter routing tool with no output schema and no annotations, the description should at least hint at the shape of the recommendation or how constraints affect it. The schema covers inputs, but the behavioral and output picture remains thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented, including the nested constraints object shape (max_tokens, speed, cost). The description adds nothing beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (analyze/recommend) and resource (model/provider routing) for an input prompt. It is clear what the tool does, but it does not distinguish itself from the adjacent hb_route_evaluate or hb_route_stats siblings in the same family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus hb_route_evaluate, hb_route_stats, or any other sibling. No prerequisites, no exclusions, and nothing about when routing is or is not appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_route_statsC
Get routing statistics and learning progress
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no read-only assurance, no auth requirements, no indication of whether stats are cumulative or per-session, and no return shape. For a no-annotation tool this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single short sentence with no padding, which is good, but the brevity comes at the cost of under-specification rather than genuine economy. Nothing is front-loaded because there is only one clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description is the only source of information about what this tool returns, and it does not describe the statistics, their format, or their freshness. An agent could call it, but could not anticipate the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to add beyond the empty schema. Baseline of 4 applies; nothing is missing or misleading on this dimension.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a verb (Get) and a resource area (routing statistics and learning progress), which loosely separates it from hb_route_select and hb_route_evaluate. But 'routing statistics' and 'learning progress' are vague — it never says what is measured, over what window, or in what scope, so an agent cannot tell exactly what it retrieves.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus hb_route_select, hb_route_evaluate, or any sibling. The agent must infer that a stats getter is for observation rather than route selection, with no explicit trigger condition or exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_state_dispatchC
Return connector-dispatch status
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Channel/chat ID | |
| message | Yes | Message text. | |
| connector | Yes | Connector name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Return' implies a read, but it does not disclose whether this requires auth, whether it mutates dispatch state, why a 'message' is required to query status, or what the response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single five-word sentence, fully front-loaded with no filler. It is efficient, though its brevity borders on under-specification rather than tight conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and a confusing required 'message' parameter for what is called a status query, the description leaves too much unexplained. An agent cannot confidently determine what this returns or why it must supply message text.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (target, message, connector) are already documented in the schema. The description adds no meaning beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a verb ('Return') and a resource ('connector-dispatch status'), so the purpose is minimally identifiable. However, 'connector-dispatch' is undefined jargon and it does not distinguish itself from close siblings like hb_conn_status, which also appears to report connector state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is given. There is no mention of prerequisites, when to prefer this over hb_conn_status or the other hb_state_* tools, or what a 'dispatch' refers to.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_state_mem_getC
Get state memory entries
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Memory type. | |
| query | No | Search query. | |
| agent_id | No | Optional agent filter |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies a read operation via 'Get' but says nothing about permissions, side effects, pagination, or return behavior, leaving disclosure minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded phrase with no wasted words. However, it is extremely terse for a tool with three optional filters and no annotations, so it is concise without being adequately structured for selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description is too thin. It does not explain what state memory entries are, how type/query/agent_id interact, or what is returned, leaving significant ambiguity against sibling memory tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all three parameters are documented there, including the 'type' enum. The description adds no syntax or filter semantics beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb ('Get') and resource ('state memory entries'), enough to distinguish from hb_state_mem_set, but it does not differentiate from other memory query siblings such as hb_mem_query or hb_mem_context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no alternatives named, and no prerequisites are stated. It is not misleading, but an agent gets no help choosing this tool over hb_mem_query, hb_mem_context, or related state tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_state_mem_setC
Store a state memory entry
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Entry key. | |
| type | No | Memory type. | |
| value | Yes | Entry value. | |
| agent_id | No | Agent identifier for state-memory provenance |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, yet it only says 'Store' – nothing about whether an existing key is overwritten, whether this requires provenance (the agent_id param hints at it), idempotency, or return behavior. A mutation tool with zero annotation coverage needs more.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence is front-loaded and waste-free, but it is under-specified rather than concise – the terseness leaves obvious gaps for a mutating, four-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A mutation tool with four parameters, no annotations, and no output schema demands explanation of overwrite behavior, provenance, and the fact/lesson distinction, none of which is present. The definition is incomplete for its complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are documented in the schema (key, type enum, value, agent_id). The description adds no syntax, format, or provenance semantics beyond what the schema already provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a verb ('Store') and resource ('state memory entry'), but the phrasing is nearly a restatement of the tool name and does not distinguish it from siblings like hb_mem_store or hb_state_mem_get. It conveys the basic action but nothing about scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus hb_mem_store (general memory) or hb_state_mem_get (retrieval). The sibling namespace makes routing ambiguous, and the description offers no help resolving it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_state_task_createC
Create a new task
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Task title. | |
| agent_id | No | Agent identifier for task provenance | |
| priority | No | Task priority. | |
| description | No | Optional description. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden, and it discloses nothing: not whether the task is persisted, whether it is dispatchable, whether agent_id is required for provenance, or what side effects creation has. For a mutation tool with zero annotation coverage, this is a critical gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is front-loaded and wastes no words, but it is under-specification rather than genuine conciseness — the brevity comes at the cost of all useful content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter mutation tool with no annotations and no output schema, the description is far too thin: it omits provenance semantics for agent_id, priority defaults, and any indication of the creation result. An agent could not call this confidently without reading the schema alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (title, agent_id, priority, description) are already documented in the schema. The description adds no format, constraint, or relationship detail beyond that, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Create a new task" essentially restates the tool name hb_state_task_create with no added specificity. It gives a verb and resource but does nothing to distinguish the task-creation semantics from siblings like hb_state_task_list or hb_state_task_update, and it never explains what a "task" is in this system.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance at all: no prerequisites, no mention of the sibling hb_state_task_update or hb_state_dispatch as alternatives, and no note on when creating a task is appropriate versus dispatching one. The agent gets no routing help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_state_task_listC
List tasks with optional status filter
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Status filter. | |
| agent_id | No | Optional agent filter |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the entire behavioral burden. It does not disclose default behavior when status is omitted (does it default to 'all'?), pagination, ordering, or return shape. For a tool with zero annotation coverage this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single terse phrase, front-loaded with the verb and resource and nothing wasted. It is efficient, though the fragmentary style leaves room for a useful clause or two about defaults.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only list tool with a fully documented schema and no output schema, the minimum viable information is present. It still lacks the default-status behavior and any hint of the result set, which an agent calling it blind would want to know.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema and the baseline is 3. The description restates the status filter (already in the schema with a full enum) and omits mention of the agent_id filter entirely, adding no meaning beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List tasks') plus a scope qualifier, so an agent immediately knows this is a read-only enumeration of tasks. It does not, however, differentiate itself from the sibling task tools (hb_state_task_create, hb_state_task_update, hb_state_dispatch), which the description never mentions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'with optional status filter' implies the filters exist but gives no when-to-use guidance, no prerequisites, and no alternative tools to consider. Nothing tells the agent when this tool is the right choice over hb_state_task_create or hb_state_task_update.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_state_task_updateC
Update a task (status, description, priority)
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Status filter. | |
| task_id | Yes | Task ID. | |
| agent_id | No | Optional agent filter | |
| priority | No | Task priority. | |
| description | No | Optional description. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. The key question for an update tool — whether omitted fields are left unchanged or cleared — is not answered, nor is there any mention of permissions, reversibility, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the action and affected fields front-loaded; nothing is wasted. It is arguably under-specified rather than padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter mutation with no annotations and no output schema, the description is thin. Partial-update semantics, the effect of each enum transition (open/in_progress/done), and the role of the filter-like fields are all unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description lists three of the five fields but says nothing that adds meaning beyond the schema, and it ignores agent_id, which the schema oddly labels 'Optional agent filter' for an update operation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Update a task') and names the three mutable fields, so the agent knows it is a mutation of an existing task. It does not differentiate itself from siblings like hb_state_task_create or hb_state_task_list, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus hb_state_task_create or hb_state_task_list, nor any stated prerequisites. Usage is only implied by the verb 'Update'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_swarm_consensusC
Get majority vote from multiple independent agents
| Name | Required | Description | Default |
|---|---|---|---|
| voters | No | Number of independent voters to plan. | |
| question | Yes | Question to evaluate by consensus. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not disclose cost, latency, determinism, tie-breaking behavior, or whether the operation is read-only or spawns real external agents.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It is appropriately sized for the amount of information it chooses to convey.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a multi-agent consensus tool with no output schema and no annotations, the one-line description is too sparse. It omits result format, tie behavior, and how the independent voters are realized.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented. The description adds only the concept of majority vote and does not extend parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Get) and resource (majority vote from multiple independent agents), so the basic action is clear. However, it does not differentiate this tool from sibling swarm tools like hb_swarm_parallel, hb_swarm_hierarchy, or hb_swarm_stigmergy.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the other swarm coordination tools. The description gives no context, prerequisites, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_swarm_hierarchyD
Boss-worker delegation pattern
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | Task description. | |
| subtasks | No | Optional subtasks for hierarchy planning. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations and no output schema, so the description carries the full burden and delivers nothing. It does not disclose whether tasks are actually executed, whether worker agents are spawned, what the hierarchy depth means, or whether the call is synchronous — all critical for an orchestration tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four words is not conciseness but under-specification; there is no front-loaded statement of what the tool does or returns, so nothing earns its place because nothing is there.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a multi-agent orchestration tool with zero annotations, no output schema, and only a fragment of description text. Everything an agent needs to invoke it correctly (execution semantics, agent spawning, result shape, relationship to the other swarm tools) is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with two well-labelled parameters (`task`, `subtasks`), so the schema does the documentation work. The description adds no meaning beyond it, which is the baseline-3 case for a fully covered schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase "Boss-worker delegation pattern" essentially restates the tool name (swarm_hierarchy) in different words rather than stating a verb and resource. It hints at an orchestration strategy but never says whether the tool plans, executes, or just configures the hierarchy, and it does not distinguish itself from siblings like hb_swarm_parallel or hb_swarm_consensus.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no routing to the sibling swarm patterns. An agent cannot tell from this text why it would pick hierarchy over parallel or consensus.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_swarm_parallelC
Split task into chunks and process in parallel
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | Task description. | |
| chunks | Yes | Task chunks to assign across workers. | |
| workers | No | Number of parallel workers to plan. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it delivers only a one-line mechanism. It says nothing about side effects, whether workers are real concurrency or just a planning hint, failure/partial-completion behavior, or result aggregation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no wasted words and the core action front-loaded. It is concise, but the brevity comes at the cost of under-specification rather than disciplined economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A parallel fan-out tool with no annotations, no output schema, and only a one-line description leaves critical questions unanswered: what gets returned, whether the split is automatic or manual, and how workers interact with the chunks. The description is too thin for the tool's apparent complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so task, chunks, and workers are already documented in the schema. The description restates the chunking concept but adds no syntax, format, or constraint detail (e.g., chunk-to-worker mapping), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a verb-and-resource action (split a task and process chunks in parallel), which is clearer than a tautology. However, it gives no differentiation from the sibling swarm tools (hb_swarm_consensus, hb_swarm_hierarchy, hb_swarm_stigmergy), so an agent cannot tell when this parallel primitive is the right choice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative-selection guidance is present. The agent must infer the context entirely from the name and the sibling list, which is exactly the situation the guidelines dimension is meant to penalize.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_swarm_stigmergyC
Indirect coordination via shared state (blackboard pattern)
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | Task description. | |
| iterations | No | Number of coordination iterations to plan. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the underlying coordination pattern (shared state/blackboard), which is one behavioral trait, but says nothing about side effects, state mutation, permissions, return behavior, or reversibility. For a coordination tool with zero annotation coverage, this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise phrase with no filler, but it is under-specified for the tool's complexity. It is front-loaded only in the sense that everything is one fragment; there is no structure to guide the agent through purpose, usage, or behavior. Concise but not appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and a coordination task that likely returns a plan or state, the description should explain what the tool returns and how it operates. It only provides a conceptual label. It is not completely empty, but it is far from complete enough to invoke the tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters. The description adds no additional parameter meaning, syntax, or constraints. Baseline 3 applies when the schema does the heavy lifting, and the description does not enhance it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the coordination mechanism ('indirect coordination via shared state') and differentiates it from other swarm strategies, but it never states the actual operation the tool performs. It is a noun phrase, not a verb+resource, so an agent must infer whether the tool plans, runs, or configures stigmergy coordination. Vague but not tautological.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no conditions, and no alternatives named. The description implies a context (indirect coordination), but it does not help the agent choose between this tool and siblings like hb_swarm_consensus or hb_swarm_parallel. Matches the 'no guidance' level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_test_listB
List available test batteries
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing beyond the name. It does not say whether results are paginated, scoped, filtered, or whether the list is static or environment-dependent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded phrase with zero filler. Every word earns its place and the purpose is conveyed immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only list tool with no output schema and no annotations, the description is barely adequate. It leaves unanswered what a 'test battery' is in this system and how its output connects to hb_test_run, which is the key thing an agent needs for multi-step test workflows.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to convey; the baseline for a parameterless tool is 4. The description correctly implies a no-argument enumeration call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List ... test batteries'), so the agent knows this enumerates test batteries. It does not distinguish itself from siblings hb_test_run and hb_test_results, which also concern tests, so the agent must infer that this is the discovery/enumeration step rather than execution or result retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance: nothing says to call this before hb_test_run to discover battery identifiers, and no alternatives or exclusions are named. The agent must guess the ordering relative to the closely related test siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_test_resultsC
Get results of last test run
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Output format. | |
| battery | No | Test battery name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not say whether results are read-only, how long results are retained, what happens if no run has occurred, or whether a run is required beforehand — significant gaps for a results-retrieval tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short, front-loaded sentence with no waste. It is efficient, though the brevity edges toward under-specification rather than true conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must explain the return shape and the 'last run' scoping semantics, but it does neither. An agent cannot tell what fields come back or what determines the 'last' run.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (format enum, battery) are already documented in the schema. The description adds no syntax, defaults, or format meaning beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb ('Get') and resource ('results of last test run'), so an agent knows the operation. However, it does not differentiate itself from siblings hb_test_list or hb_test_run, and 'last test run' is ambiguous given an optional battery parameter.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus hb_test_run (which produces results) or hb_test_list. The implicit ordering dependency — that a run must exist first — is left for the agent to infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_test_runC
Run a test battery or single test
| Name | Required | Description | Default |
|---|---|---|---|
| test | No | Optional single test name. | |
| battery | Yes | Test battery name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, yet it says nothing about side effects, whether execution is blocking, timeouts, permission needs, or how failures surface. For an execution tool this is a substantial gap; 'Run' alone does not tell an agent what will happen.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It is efficient, though the brevity contributes to the missing behavioral and usage context rather than compensating for it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations mean the description must explain what execution produces and where results are retrieved (likely hb_test_results), but it does neither. For a test-execution tool with side effects and no structured coverage, the description is materially incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both 'battery' and 'test' are already documented in the schema. The description only hints that a single test may be run instead of a battery, which is already implied by the optional 'test' parameter, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Run') and resource ('test battery or single test'), which distinguishes it from the read-only siblings hb_test_list and hb_test_results. It does not explicitly name those siblings or explain the boundary, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus hb_test_list or hb_test_results, no prerequisites (must the battery already exist?), and no indication of when a single test is preferable to a battery. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_ticket_listA
List tickets in one lifecycle category (INBOX/ACTIONABLE/QUEUED/BLOCKED/WAITING/USER/PARKED/SOLVED) -- header fields only, read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return. | |
| category | Yes | Category filter or category to store. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It usefully discloses that only header fields are returned and that the operation is read-only, but says nothing about pagination behavior (limit defaults to 50, what happens beyond it), ordering, or whether empty categories error out.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the read-only and header-only caveats front-loaded after the core action. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description does well to state that only header fields come back, but an agent still lacks ordering, pagination, and error behavior for the required category parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented in the schema, giving a baseline of 3. The description adds the notion of 'lifecycle category' but does not clarify limit semantics or what 'header fields' means for the returned objects.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (tickets) scoped to one lifecycle category, and enumerates the valid categories inline. It implicitly distinguishes itself from hb_ticket_show by declaring 'header fields only', though it never names that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the category enumeration and the 'header fields only' qualifier, but the description never states when to prefer this over hb_ticket_show or what categories mean operationally. No explicit when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hb_ticket_showB
Show one ticket's header fields by ID (e.g. T-20260825-196589547), read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| ticket_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It declares read-only, which is a useful behavioral trait, but says nothing about permissions, error behavior for missing tickets, or what 'header fields' omits (e.g. comments, attachments). For a retrieval tool with zero annotations this is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the verb and resource, no waste. The example ID and read-only hint are packed efficiently rather than padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with no annotations and no output schema, the description should explain what 'header fields' includes, what a not-found response looks like, and how it differs from hb_ticket_list. It leaves these gaps unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It does provide a concrete ID format example (T-20260825-196589547), which adds meaning beyond the bare 'ticket_id' string type. Still, only one param is documented this way and no format constraints beyond the example are given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Show) and resource (one ticket's header fields) with an ID example. It's clear and distinguishable from sibling hb_ticket_list, though it doesn't explicitly name that sibling to differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'by ID' and the example format, and the read-only note signals a safe inspection call. But there's no explicit when-to-use vs hb_ticket_list, no prerequisites or error conditions for invalid IDs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
51 tool updates
v0.1.0-alpha.29- First observed
hb_api_discover - First observed
hb_api_export - First observed
hb_api_history - First observed
hb_api_probe - First observed
hb_auto_list_chains - First observed
hb_auto_result - First observed
hb_auto_run - First observed
hb_auto_status - First observed
hb_conn_list - First observed
hb_conn_receive - First observed
hb_conn_send - First observed
hb_conn_status - First observed
hb_garden_find - First observed
hb_garden_get - First observed
hb_garden_put - First observed
hb_garden_run - First observed
hb_kb_get - First observed
hb_kb_ingest - First observed
hb_kb_list - First observed
hb_kb_search - First observed
hb_lock_check - First observed
hb_lock_list - First observed
hb_mem_consolidate - First observed
hb_mem_context - First observed
hb_mem_merge - First observed
hb_mem_query - First observed
hb_mem_store - First observed
hb_plug_discover - First observed
hb_plug_info - First observed
hb_plug_list - First observed
hb_plug_run - First observed
hb_policy_list - First observed
hb_policy_resolve - First observed
hb_route_evaluate - First observed
hb_route_select - First observed
hb_route_stats - First observed
hb_state_dispatch - First observed
hb_state_mem_get - First observed
hb_state_mem_set - First observed
hb_state_task_create - First observed
hb_state_task_list - First observed
hb_state_task_update - First observed
hb_swarm_consensus - First observed
hb_swarm_hierarchy - First observed
hb_swarm_parallel - First observed
hb_swarm_stigmergy - First observed
hb_test_list - First observed
hb_test_results - First observed
hb_test_run - First observed
hb_ticket_list - First observed
hb_ticket_show
TDQS
Scored across 51 tools
Most tools are cleanly separated by domain prefix and verb. Minor overlaps exist: hb_api_probe vs hb_api_discover (probe-all-strategies vs schema auto-detect), and hb_mem_* vs hb_state_mem_* memory families could be confused by an agent. The many *_run/*_list/*_get tools are differentiated by their domain prefix, so ambiguity is limited.
Strong, predictable hb_<domain>_<verb> pattern across nearly all tools (hb_kb_search, hb_mem_store, hb_lock_check). A few deviations: hb_auto_list_chains inverts the order used by hb_auto_run/status/result, and accessors are inconsistently named hb_garden_get / hb_kb_get vs hb_ticket_show.
51 tools is well above the heavy threshold and far beyond what an agent can reliably select from in one session. The surface bundles at least a dozen separate domains (plugins, memory, routing, KB, swarm, state, garden, API, tests, connectors, automation, policy, tickets, locks), which should be split into focused servers.
Coverage is broad per domain but shallow: plugins have no install/enable, KB and garden lack update/delete, tasks have create/update but no delete, and policy/ticket tools are read-only by design. Read-only design is defensible, but the missing mutating operations leave lifecycle dead ends in several domains.
Maintenance
Related MCP Connectors
Private-by-default, local-first memory/context/task orchestrator for MCP apps and agents.
Persistent memory, hybrid search and a goal graph for AI agents, over stdio or remote HTTP.
Agent-native notes, tasks, dev-docs, vaults, sync & handoffs. MCP + OpenAPI dual surface.
Hosted runtime for persistent agent teams, durable workflows, memory, schedules, and goals.
Related MCP Servers
- FlicenseAqualityDmaintenanceA durable multi-agent orchestrator for software development with explicit run graphs, checkpoint/resume capabilities, and project memory exposed through MCP resources and tools. It enables coordinated agent workflows for coding, review, repair, CI, and approval with SQLite-backed memory retrieval and pluggable research backends.10-
- AlicenseNot gradedqualityFmaintenanceLocal-first, auditable memory for AI agents. Provides durable context for MCP hosts with SQLite storage, CLI, and MCP tools for memory management.2Apache 2.0
- AlicenseAqualityFmaintenanceLocal-first MCP memory server that gives AI coding agents long-term memory via SQLite and sqlite-vec, with optional LLM-powered layering. No gateway or API key required.7131 npm4MIT
- FlicenseNot gradedqualityAmaintenanceA local-first MCP server and CLI that gives coding agents structured project memory, task contracts, context packs, backlog workflows, and verification evidence, storing data in reviewable Markdown/YAML with a fast SQLite index.4-