Skip to main content
Glama
ellmos-ai

ellmos-homebase-mcp

Official

ellmos-homebase-mcp

Alpha MCP server for local-first LLM orchestration: memory, knowledge, routing, swarm patterns, API probing, persistent state, tests, automation planning, and plugin discovery in one stdio server.

Homebase is designed primarily for local LLMs (Ollama, Qwen, Llama, or any locally-hosted model via a MCP-capable harness). All persistent storage uses SQLite with no cloud dependency. External LLM providers (Claude, Codex, Gemini, OpenAI) can also connect as MCP clients, but local, offline-capable setups are the primary target.

German README: README_de.md

Part of the ellmos-ai family under the open-bricks umbrella.

Ecosystem: open-bricks Organization: ellmos-ai License: MIT Attribution: NOTICE npm version Python Matrix Node.js Platforms Privacy Storage MCP Status: alpha Tests Security SLA Security: RunAsInvoker Level 1 SBOM: Plain Text Audit Last-Checked Code style: ruff LLMs-Ready Homebase tests

Discoverability: Published on npm as ellmos-homebase-mcp and maintained in the ellmos-ai organization.

NOTE

For AI Assistants & LLM Agents: Machine-readable architecture summary, index, and tool capabilities are published in llms.txt. MCP registry metadata is available in server.json.

Quick Navigation / Schnellnavigation


Related MCP server: nuzo-memory

System Architecture

flowchart TD
    subgraph Clients ["MCP Clients (Local / Remote)"]
        Ollama["Local LLMs (Ollama, Qwen, Llama)"]
        Claude["Claude Code / Desktop"]
        Codex["Codex / Antigravity"]
    end

    subgraph Transport ["Transport Layer"]
        Stdio["stdio (Python MCP SDK)"]
    end

    subgraph Core ["ellmos-homebase-mcp Core Engine"]
        Server["homebase.server"]
        Config["homebase.config"]
    end

    subgraph ToolGroups ["51 MCP Tools across 14 Functional Modules"]
        Mem["hb_mem_* (SQLite Memory)"]
        KB["hb_kb_* (Knowledge Digest)"]
        State["hb_state_* (State & Tasks)"]
        Route["hb_route_* (Model Router)"]
        Swarm["hb_swarm_* (Swarm Patterns)"]
        Api["hb_api_* (API Probing)"]
        Conn["hb_conn_* (Connectors Queue)"]
        Auto["hb_auto_* (Automation Chains)"]
        Plug["hb_plug_* (Plugin Discovery)"]
        Garden["hb_garden_* (Garden Store)"]
        Test["hb_test_* (Self Tests)"]
        Policy["hb_policy_* (Policy Registry, read-only)"]
        Ticket["hb_ticket_* (Ticket Master, read-only)"]
        Lock["hb_lock_* (Lock Master, read-only)"]
    end

    subgraph Storage ["Local Storage (Offline-First)"]
        DB[(SQLite Storage ~/.homebase/)]
    end

    Clients --> Stdio
    Stdio --> Server
    Server --> Config
    Server --> ToolGroups
    ToolGroups --> DB

Four-View Architectural Topology Projection

+-------------------------------------------------------------------------------+
|  VIEW 1: CALLER RUNTIMES, AGENT CLIENTS & ENTRYPOINTS                         |
|  - Local LLM Engines: Ollama (Qwen, Llama, Mistral, DeepSeek), Local Harnesses|
|  - Multi-Agent Orchestrators: Claude Code, Claude Desktop, OpenAI Codex, AGY  |
|  - IDE & Extension Interfaces: Cursor, VS Code MCP Extension, Windsurf        |
|  - stdio Protocol Transport: JSON-RPC 2.0 via Python Model Context Protocol   |
+---------------------------------------+---------------------------------------+
                                        | JSON-RPC 2.0 stdio (tools/list, tools/call)
                                        v
+-------------------------------------------------------------------------------+
|  VIEW 2: HOMEBASE MCP SOVEREIGN CORE & DISPATCH ORCHESTRATOR                  |
|  - Core Server & Life Cycle: homebase.server (stdio loop, signal handling)    |
|  - Module Registry & Dispatch: homebase.registry (i18n schemas, locale norm)  |
|  - 51 Sovereign MCP Tools across 14 Specialized Functional Modules:          |
|    * Memory & Knowledge: hb_mem_* (SQLite memory), hb_kb_* (FTS5 search)      |
|    * State & Planning: hb_state_* (tasks & KV), hb_garden_* (garden store)    |
|    * Routing & Swarms: hb_route_* (offline routing), hb_swarm_* (blueprints)  |
|    * Exploration & Testing: hb_api_* (schema probe), hb_test_* (diagnostics)  |
|    * Integration & Staging: hb_conn_* (safe queues), hb_auto_* (chain plans)  |
|    * Extensibility: hb_plug_* (dry-run discovery, no remote code execution)   |
|    * Canonical Seams: hb_policy_*, hb_ticket_*, hb_lock_* (read-only views)   |
+-------------------+-----------------------------------+-----------------------+
                    |                                   |
                    v (bundled SQLite mode)             v (canonical engine mode)
+---------------------------------------+ +-------------------------------------+
|  VIEW 3: RUNTIME PERSISTENCE &        | |  VIEW 3-ALT: CANONICAL SEAMS        |
|  SQLITE STORAGE ENGINE                | |  (MODE-CONTRACT.md)                 |
|  - Database: ~/.homebase/homebase.db  | |  - Policy Registry (hb_policy_*)    |
|  - Concurrency: Write-Ahead Log (WAL) | |  - Ticket Master (hb_ticket_*)      |
|  - Multi-Agent Provenance: agent_id   | |  - Lock Master (hb_lock_*)          |
|  - Search Engine: SQLite FTS5 index   | |  - Fail-Closed Discipline:          |
|  - Integrity: Busy timeouts & rollback| |    Raises CanonicalEngineUnavailable|
+-------------------+-------------------+ +-----------------+-------------------+
                    |                                       |
                    +-------------------+-------------------+
                                        |
                                        v
+-------------------------------------------------------------------------------+
|  VIEW 4: AIR-GAP DEFENSE PERIMETER, RUNASINVOKER & ZERO-EGRESS BOUNDARY        |
|  - 100% Local-First & Zero Egress: INV-LOCAL-01 (0 cloud calls, 0 telemetry)  |
|  - Unprivileged Execution: INV-PERM-08 (RunAsInvoker non-elevation principle) |
|  - Engine Seam Integrity: INV-ENGINE-02 & INV-SEAM-03 (strict fail-closed)   |
|  - Deterministic Provenance: INV-PROV-04 (agent_id attribution on all state)  |
|  - Credential-Free Operation: INV-CRED-05 & INV-STAGE-06 (no tokens/secrets)  |
|  - Multi-Host Lock Defense: INV-SYNC-09 (.gitignore sync/lock immunity)       |
|  - Permissive Licensing: Zero-Copyleft stack (MIT, PSFL-2.0, Apache-2.0)     |
|  - Statutory SLA & Disclaimer: INV-SLA-10 (48h response SLA, § 521 BGB)       |
+-------------------------------------------------------------------------------+

Sequence Flow & Lifecycle

sequenceDiagram
    autonumber
    participant Client as MCP Client (Local LLM / Claude / Codex)
    participant Stdio as Transport Layer (stdio)
    participant Server as Server & Registry (homebase)
    participant Module as Functional Module (hb_mem / hb_state / hb_route)
    participant Engine as Engine Seam (Bundled vs Canonical)
    participant DB as SQLite Storage (~/.homebase/)

    Client->>Stdio: JSON-RPC 2.0 Request (tools/call: hb_mem_store, agent_id="agent-01")
    Stdio->>Server: Decode & dispatch tool call
    Server->>Module: Validate arguments & inject agent provenance
    alt Bundled Engine Mode (Default)
        Module->>DB: Execute SQLite query (WAL mode, busy timeout)
        DB-->>Module: Return structured records / mutation status
    else Canonical Engine Mode ([engines].mode = "canonical")
        Module->>Engine: Seam check (Gardener / TASKPLAN / USMC)
        alt Engine Available
            Engine-->>Module: Delegate to canonical subsystem
        else Engine Unreachable
            Engine-->>Module: Raise CanonicalEngineUnavailable (Fail-Closed)
        end
    end
    Module-->>Server: Format response in requested language (i18n: en/de/es/zh/ja/ru)
    Server-->>Stdio: Encode JSON-RPC 2.0 Response
    Stdio-->>Client: Result payload (Zero cloud egress, 100% local)

Core Capabilities & Security Invariants

Capability / Invariant

Guarantee

Technical Implementation

100% Local-First & Zero-Egress

Complete privacy and offline operation; no unexpected cloud communication or telemetry.

All persistent memory, knowledge, and state are saved in local SQLite (~/.homebase/).

Strict Engine Seams & Fail-Closed

No silent fallback into disconnected databases when requesting canonical systems.

MODE-CONTRACT.md enforcement: raises CanonicalEngineUnavailable if target is unreachable.

Team-Memory Provenance (agent_id)

Deterministic audit trail and filterable ownership for multi-agent workflows.

Native agent_id tracking across memory facts, knowledge entries, and task state.

Credential-Free Discovery & Planning

Zero secret exposure during local routing recommendations and API probing.

hb_route_*, hb_swarm_*, and hb_api_* run without transmitting API keys or private tokens.

Safe Plan-and-Queue Adapters

Safe queueing and chain staging without arbitrary remote code execution.

hb_conn_* and hb_auto_* maintain plan-only queues and offline staging records.

Full Native i18n Localization

Seamless multilingual developer and agent interaction.

Localized tool descriptions and JSON schemas for en, de, es, zh, ja, ru.

Non-Elevation & Secret Hygiene

Unprivileged execution and strict credential exclusion from distribution.

Non-root compatibility; live configs/secrets ignored in .gitignore and .npmignore.

Multi-OS CI Smoke Integrity

Verified cross-platform reliability on all major operating systems.

Multi-version CI matrix covering Python 3.10–3.13 and Node.js 20–24 on Linux/Windows/macOS.


Governance & Runtime Invariants

Invariant ID

Title & Scope

Guarantee & Technical Enforcement

Verification Seam

INV-LOCAL-01

100% Local-First & Zero-Egress

All persistent memory, knowledge entries, and task states are stored locally in SQLite (~/.homebase/). Zero telemetry, analytics, or unrequested outbound cloud network calls.

tests/test_server_transport.py, tests/test_repository_hygiene.py

INV-ENGINE-02

Strict Engine Seams & Fail-Closed

MODE-CONTRACT.md enforcement: switching [engines].mode = "canonical" never silently falls back to bundled storage if the canonical engine is unreachable.

tests/test_engine_seams.py

INV-SEAM-03

Canonical-Only Isolation

hb_policy_*, hb_ticket_*, and hb_lock_* provide read-only views into policy-registry, ticket-master, and lock-master. They possess no bundled imitation and fail closed unconditionally.

tests/test_new_seams.py

INV-PROV-04

Deterministic Provenance & Team-Memory

Multi-agent coordination requires strict isolation. All memories, knowledge facts, and task transitions record agent_id attribution with SQLite WAL concurrency and busy timeouts.

tests/test_module_contracts.py

INV-CRED-05

Credential-Free Discovery & Probing

Model routing suggestions (hb_route_*), swarm pattern blueprints (hb_swarm_*), and API schema probing (hb_api_*) function without private API keys, tokens, or credentials.

tests/test_module_contracts.py

INV-STAGE-06

Plan-Only Staging & Bounded Offline Queues

Connector queues (hb_conn_*) and automation plans (hb_auto_*) record offline blueprints and dry-run staging manifests without executing arbitrary remote code or side-effects.

tests/test_module_contracts.py

INV-I18N-07

Native Multilingual Schema Parity

All 51 tool definitions, input schemas, and validation errors maintain 100% complete localization across 6 supported languages (en, de, es, zh, ja, ru).

tests/test_i18n_completeness.py

INV-PERM-08

Non-Elevation & RunAsInvoker Principle

Homebase runs strictly in unprivileged user space. It requires no administrator or root privileges and ignores sensitive local dotfiles and system credentials.

tests/test_repository_hygiene.py

INV-SYNC-09

Multi-Host Lock & Conflict Discipline

Strict exclusion of conflict copies (*.sync-conflict-*, *-conflict-*) and honor of multi-agent lock mechanisms (LOCK.*, *.lock) to preserve database integrity across hosts.

tests/test_metadata.py

INV-SLA-10

48h Response, 5-Day Triage & 30-Day Remediation SLA

Security disclosures sent to security@ellmos.ai, support@lukasgeiger.com, or security@open-bricks.org receive guaranteed initial response in <=48h, triage within 5 business days, and verified remediation within 30 calendar days.

SECURITY.md, tests/test_metadata.py


Target Personas & Discoverability

Homebase is purpose-built to solve architectural and operational challenges across four core technical audiences:

[PERSONA-01] Local LLM & Edge AI Developers

  • Profile & Objective: AI engineers building offline or edge applications with Ollama, Qwen, or Llama models who need a robust orchestration harness.

  • Pain Points: Cloud memory APIs introduce unwanted latency, privacy leaks, subscription billing, and network failure modes.

  • Homebase Solution: Zero-cloud dependency, local SQLite WAL persistence (~/.homebase/), and 51 standard stdio tools providing memory, FTS5 knowledge search, and task tracking.

  • Reference Workflow:

    {"tool": "hb_mem_store", "arguments": {"fact": "User prefers compact JSON output", "agent_id": "ollama-coder"}}
    {"tool": "hb_kb_search", "arguments": {"query": "API routing rules", "fts": true}}

[PERSONA-02] Multi-Agent Swarm Orchestrators & Swarm Architects

  • Profile & Objective: System architects orchestrating multi-agent collectives (Claude Code, Codex, Antigravity, local agents) operating concurrently on shared codebases.

  • Pain Points: State collisions, lack of origin tracking, race conditions in shared memory, and uncoordinated task delegation.

  • Homebase Solution: Native agent_id provenance across all facts, memories, and task states; built-in swarm templates (boss/worker, chunked parallel, consensus voting via hb_swarm_*).

  • Reference Workflow:

    {"tool": "hb_swarm_plan", "arguments": {"goal": "Audit security seams", "pattern": "consensus"}}
    {"tool": "hb_state_task_create", "arguments": {"title": "Verify fail-closed mode", "agent_id": "worker-audit-01"}}

[PERSONA-03] Enterprise Security & Data Governance Officers

  • Profile & Objective: CISOs, SecOps teams, and compliance auditors in regulated industries (healthcare, finance, defense) evaluating developer agent toolchains.

  • Pain Points: Silent cloud telemetry, unvetted remote side-effects, privilege escalation risks, and missing SLA assurances.

  • Homebase Solution: Strict zero-egress architecture, fail-closed canonical engine seams (MODE-CONTRACT.md), unprivileged RunAsInvoker operation, and formal 48h Security Response SLA (SECURITY.md).

  • Reference Workflow:

    {"tool": "hb_policy_list_rules", "arguments": {}}

    Guaranteed fail-closed behavior: raises CanonicalEngineUnavailable instead of silently falling back to insecure stubs.

[PERSONA-04] Cross-Framework AI Assistants & Pair Programmers

  • Profile & Objective: Developers utilizing multiple AI coding assistants (Claude Desktop, Codex, Cursor, Gemini) seeking uniform context and tool parity across environments.

  • Pain Points: Incompatible custom tool APIs, fragmented scratchpads, and lack of multilingual developer schemas.

  • Homebase Solution: Standard stdio MCP transport, machine-readable project metadata (llms.txt, server.json, glama.json), and 100% complete schema localization across 6 languages (en, de, es, zh, ja, ru).

  • Reference Workflow:

    {"tool": "hb_ticket_list", "arguments": {"folder": "ACTIVE"}}

High-Intent Search & SEO Keywords

  • English Intent: local-first LLM orchestration MCP server, offline agent memory SQLite WAL, stdio Model Context Protocol Ollama Qwen, multi-agent swarm planning persistent state, zero-egress MCP server enterprise AI, fail-closed engine seams MODE-CONTRACT, team-memory agent_id provenance.

  • German Intent: Local-First LLM-Orchestrierung MCP-Server, Offline Agenten-Memory SQLite WAL, Model Context Protocol Stdio-Server Ollama, Multi-Agenten Schwarmplanung persistenter Zustand, Zero-Egress MCP-Server Unternehmens-KI, Fail-Closed Schnittstellen MODE-CONTRACT, Team-Memory Agenten-Provenienz.


Comparative Matrix vs. Alternatives

Homebase provides a uniquely comprehensive, local-first MCP capability stack compared to specialized or cloud-bound alternatives:

Architectural & Runtime Dimension

ellmos-homebase-mcp

Cloud Memory SaaS (Letta, Pinecone, LangSmith)

Generic Memory MCPs (mcp-server-memory, sqlite)

Heavy Agent Frameworks (CrewAI, AutoGen, LangGraph)

Ad-Hoc Scripts / Custom SQLite

1. 100% Local-First & Zero Egress (INV-LOCAL-01)

Yes (100% local SQLite WAL, zero telemetry)

No (Cloud-hosted, mandatory egress, PII risk)

Partial (Local file, but no strict egress contracts)

Variable (Often requires cloud API keys / SaaS)

Yes (Local, but no protocol guarantees)

2. Engine Seams & Fail-Closed (INV-ENGINE-02)

Yes (Strict MODE-CONTRACT.md, raises error on failure)

No (Opaque cloud failovers)

No (Single hardcoded backend)

No (Unchecked exceptions / silent fallbacks)

No (Ad-hoc failure handling)

3. Canonical-Only Seams (INV-SEAM-03)

Yes (hb_policy_*, hb_ticket_*, hb_lock_* fail closed)

No (No canonical system awareness)

No (No policy or lock integration)

No (No governance seam layer)

No (Manual coordination)

4. Team-Memory & Attribution (INV-PROV-04)

Yes (Native agent_id on facts, knowledge, tasks)

Partial (User-level only, lacks multi-agent filters)

No (Single global unpartitioned graph)

Partial (In-memory agent state, lost on restart)

No (Manual schema management)

5. Credential-Free Discovery (INV-CRED-05)

Yes (Offline routing & swarm planning without tokens)

No (Requires active paid cloud credentials)

No (No model routing or swarm tools)

No (Requires API keys for LLM planners)

No (No structured planning)

6. Plan-Only Staging Queues (INV-STAGE-06)

Yes (Safe connector queues & dry-run automation)

No (Direct execution or none)

No (No connector or automation support)

No (Direct runtime side-effects)

No (Unsafe arbitrary execution)

7. Tool Breadth & Surface

51 Tools across 14 Modules in single stdio server

1-5 API endpoints

2-5 basic tools

Framework-level Python library (not MCP native)

Fragmented CLI utilities

8. Multilingual Schema Parity (INV-I18N-07)

Yes (Full en, de, es, zh, ja, ru schema coverage)

English only

English only

English only

English only / None

9. Non-Elevation Security (INV-PERM-08)

Yes (Unprivileged RunAsInvoker, dotfile defense)

Cloud SaaS (Tenant-isolation trust model)

Variable (Local file permissions)

Variable (Often runs in root containers)

Variable (User scripts)

10. Security Response SLA (INV-SLA-10)

Yes (Formal 48h Response, 5d Triage & 30d Remediation in SECURITY.md)

Commercial SLA (Paid tiers only)

None / Best-effort community

None / Best-effort community

None


Start Here

Need

Entry point

Install the alpha MCP server

npm install -g ellmos-homebase-mcp@alpha

Run from a source checkout

python -m homebase.server with PYTHONPATH=src

Configure a local LLM harness, Claude Code, Codex, or any MCP client

MCP Client Configuration

Inspect the machine-readable project summary

llms.txt

Check registry metadata

server.json


Status

  • Transport: stdio via the Python MCP SDK

  • Package status: public alpha package under ellmos-ai

  • Release metadata: MIT LICENSE, NOTICE, CHANGELOG.md, llms.txt, and MCP Registry metadata in server.json

  • Test gate: GitHub Actions covers Python 3.10/3.11/3.12/3.13 plus Node.js 20/22/24 smoke and npm package checks

  • Current core: module discovery, MCP tool listing, MCP tool dispatch, config fallbacks, local planning/probing/queue/dry-run adapters

  • Real local SQLite modules: hb_mem_*, hb_kb_*, hb_garden_*, hb_state_*

  • Engine seams: hb_garden_*, hb_state_task_* and hb_mem_* can delegate to the real canonical Gardener/Rinnsal/USMC engines instead of the bundled SQLite copies via [engines].mode = "canonical" (default remains "bundled" for a zero-dependency install). No silent fallback: if you request canonical and the engine is unreachable, those tools return an error rather than quietly using the bundled DB — the server still starts and lists its tools. Binding rule and migration notes: MODE-CONTRACT.md; mechanism: KONZEPT.md.

  • Canonical-only seams (no bundled alternative at all): hb_policy_* (policy-registry), hb_ticket_* (ticket-master), hb_lock_* (lock-master) — all read-only in v1. A locally faked copy of live policy/ticket/lock state would mislead rather than help, so these three always attempt the canonical module and fail closed unconditionally if it is unreachable.

  • Team-memory basics: agent_id provenance and filters for memory, knowledge, state memory, and tasks; SQLite uses WAL plus a busy timeout for safer concurrent agents

  • Credential-free alpha adapters: hb_route_*, hb_swarm_*, hb_api_*, hb_test_*, hb_conn_*, hb_auto_*, hb_plug_*

  • i18n: fully localized MCP tool descriptions, input-schema field descriptions, and unknown-tool errors for en, de, es, zh, ja, ru (English fallback for any unset key)

  • Roadmap: optional real LLM/API integrations and explicit execution backends


Install

The npm package contains a Node wrapper that starts the Python server. You still need Python 3.10+ and the Python package mcp>=1.0.0.

Option 1: Install From npm

npm install -g ellmos-homebase-mcp@alpha
ellmos-homebase

Option 2: Install From Source

git clone https://github.com/ellmos-ai/ellmos-homebase-mcp.git
cd ellmos-homebase-mcp
$env:PYTHONIOENCODING = "utf-8"
python -m pip install -e ".[dev]"
python -m pytest -ra -v

Avoid creating a .venv inside cloud-synced folders if your sync client locks files. If you need an isolated environment, create it outside that folder.

Start From Source

$env:PYTHONPATH = "src"
python -m homebase.server

MCP Client Configuration

Homebase uses the standard stdio mcpServers configuration format. The same snippet works in any MCP-capable client or harness: BACH/Buddha (local Ollama), Claude Code, Codex, Cursor, or any other MCP host.

Note on local LLMs: A bare Ollama instance does not speak MCP natively — you need a MCP-capable harness on top of it (e.g., BACH, an open-source MCP proxy, or another orchestration layer). Configure that harness to include Homebase as an MCP server using the snippet below.

Global npm Install

{
  "mcpServers": {
    "homebase": {
      "command": "ellmos-homebase"
    }
  }
}

Source Checkout

{
  "mcpServers": {
    "homebase": {
      "command": "python",
      "args": ["-m", "homebase.server"],
      "env": {
        "PYTHONPATH": "/absolute/path/to/ellmos-homebase-mcp/src"
      }
    }
  }
}

Replace /absolute/path/to/ellmos-homebase-mcp with your local checkout path.


Server Configuration

Example: config/homebase.example.toml

Machine-readable project context: llms.txt

MCP Registry metadata: server.json

Default paths:

  • %USERPROFILE%\.homebase\homebase.toml

  • %USERPROFILE%\.config\homebase\homebase.toml

  • override with HOMEBASE_CONFIG

Language can be configured with [server].language, HOMEBASE_LANG, or HOMEBASE_LOCALE. The writing agent can be passed per tool call as agent_id; otherwise modules use HOMEBASE_AGENT_ID, AGENT_ID, a module-level agent_id, or unknown.

[server]
name = "ellmos-homebase"
language = "en" # en, de, es, zh, ja, ru

[modules]
enabled = ["mem", "route", "kb", "swarm", "state", "garden", "api", "test", "conn", "auto", "plug"]

Modules with missing optional dependencies are skipped without blocking server startup.


Tools

Important tool groups:

  • hb_mem_* for SQLite-backed memory

  • hb_kb_* for SQLite-backed knowledge entries

  • hb_state_* for persistent SQLite state and tasks

  • hb_garden_* for a small SQLite garden store

  • hb_route_* for credential-free model-routing recommendations and feedback stats

  • hb_swarm_* for credential-free swarm planning patterns

  • hb_api_* for passive HTTP API discovery with SQLite history

  • hb_test_* for built-in metadata and smoke self-tests

  • hb_conn_* for a local connector registry plus SQLite-backed inbox/outbox queues without network sends

  • hb_auto_* for local automation chain definitions and queued plan-only runs without backend execution

  • hb_plug_* for local plugin discovery and dry-run records without executing plugin code

  • hb_policy_* (read-only, canonical-only) for resolving/listing policy-registry rules

  • hb_ticket_* (read-only, canonical-only) for listing/showing ticket-master tickets by lifecycle folder

  • hb_lock_* (read-only, canonical-only) for checking/listing active lock-master locks


Discovery Context

Use ellmos-homebase-mcp when searching for a local-first, offline-capable MCP server that gives local LLMs (Ollama, Qwen, Llama, or similar) persistent memory, knowledge management, routing, and orchestration — without requiring any cloud dependency. External LLM providers can also use it as an MCP server, but local-first setups are the primary design target.

Good search phrases:

  • ellmos Homebase MCP server

  • local-first LLM orchestration MCP

  • MCP server SQLite memory knowledge routing

  • offline agent orchestration MCP server

  • MCP swarm planning persistent state API discovery

Not the same as Elmo/ELMO voice tools, AllenAI ELMo embeddings, Eclipse LMOS, generic cloud agent platforms, or single-purpose MCP memory servers.


ellmos-ai Ecosystem

This MCP server is part of the ellmos-ai ecosystem — AI infrastructure, MCP servers, and intelligent tools.

MCP Server Family

Server

Tools

Focus

npm

FileCommander

47

Filesystem, process management, interactive sessions, cloud-lock-safe operations

ellmos-filecommander-mcp

CodeCommander

23

Code analysis, JSON repair, imports, diffs, regex

ellmos-codecommander-mcp

Clatcher

12

File repair, format conversion, batch operations

ellmos-clatcher-mcp

n8n Manager

19

n8n workflow management via AI assistants

n8n-manager-mcp

ControlCenter

20

MCP stack discovery, profile management, control plane

ellmos-controlcenter-mcp

Homebase

51

Local-first LLM memory, knowledge, state, routing, swarm orchestration

ellmos-homebase-mcp (alpha)

ServerCommander

8

Server operations: health checks, log analysis, deploy dry-runs, mail diagnostics

ellmos-servercommander-mcp (alpha)

Blender Use

3

Headless Blender asset QA and FBX reimport verification

ellmos-blender-use-mcp (alpha)

Open Compute

10

Model-agnostic computer use: capture, safety-gated actions, Windows UIA

open-compute-mcp (alpha)

AI Infrastructure

Project

Description

BACH

Local-first text-based OS for LLM agents — 113+ handlers, 550+ tools, SQLite memory

open-compute

Model-agnostic computer-use core powering Open Compute MCP

clutch

Provider-neutral LLM orchestration with auto-routing and budget tracking

rinnsal

Lightweight agent memory, connectors, and automation infrastructure

ellmos-stack

Self-hosted AI research stack (Ollama + n8n + Rinnsal + KnowledgeDigest)

MarbleRun

Autonomous agent chain framework for Claude Code

gardener

Minimalist database-driven LLM OS prototype (4 functions, 1 table)

ellmos-tests

Testing framework for LLM operating systems (7 dimensions)

Desktop Software & Sibling Ecosystem

Our partner umbrella organization open-bricks and sister organizations maintain local-first, privacy-centric desktop software and developer tools:

Application / Tool

Organization

Focus & Integration

ProFiler

file-bricks

Local-first desktop file organizer and PII-safe workspace exchange

DokuZen

doc-bricks

Distraction-free Markdown & PDF documentation manager

PDFtoPDFocr

doc-bricks

Local-first PDF OCR and text layer embedding

KnowledgeDigest

doc-bricks

Offline document summarization and embedding engine

DevCenter

dev-bricks

Developer workspace hub and multi-repository management

CodeBox

dev-bricks

Isolated sandbox runner and local code execution assistant

MemoryHooker

ellmos-ai

Hook-based LLM memory provenance and session injection gate

sqlite-transit-sync

ellmos-ai

Zero-dependency SQLite schema migration & replication layer


Third-Party Licenses & Level 1 SBOM

ellmos-homebase-mcp is verified to contain 0% copyleft dependencies. All runtime dependencies are permissively licensed (MIT, BSD-2-Clause, Apache-2.0, PSFL).

Full inventory, Level 1 SBOM Invariant Cross-Reference Matrix, and non-elevation certifications are documented in THIRD_PARTY_LICENSES.md and plain-text companion THIRD_PARTY_LICENSES.txt. Canonical copyright and author attribution is maintained in NOTICE.


Security & Vulnerability Reporting

ellmos-homebase-mcp strictly adheres to local-first, zero-egress, and non-elevation security principles. Full policies, SLAs, and security guarantees are documented in SECURITY.md:

  • Supported Versions: 0.1.0-alpha.x

  • Response SLA: Initial acknowledgment and triage within 48 hours. Detailed triage within 5 business days; remediation within 30 calendar days.

  • Security Contacts: security@ellmos.ai, support@lukasgeiger.com, and security@open-bricks.org.

  • Private Advisory: GitHub Security Advisories.


Development

$env:PYTHONIOENCODING = "utf-8"
$env:PYTHONDONTWRITEBYTECODE = "1"
python -m pytest -ra -v
npm run smoke
npm pack --dry-run --json

Next useful step: add optional execution backends behind explicit configuration.


License & Statutory Liability Disclaimer (§ 521 BGB)

Software License

ellmos-homebase-mcp is open-source software licensed under the MIT License. Canonical attribution to Lukas Geiger, the ellmos-ai family, and the open-bricks ecosystem is formally preserved in NOTICE. Third-party component licenses are cataloged in THIRD_PARTY_LICENSES.md.

Statutory Notice & Liability Limitation (§ 521 BGB - German Law)

This software is made available free of charge as an open-source project. Under German statutory law governing gratuitous software provision (§ 521 BGB Gefälligkeitsrecht):

  1. Liability Limitation: The author and contributors are liable only in cases of intentional misconduct (Vorsatz) or gross negligence (grobe Fahrlässigkeit).

  2. Warranty Limitation: In accordance with §§ 523, 524 BGB, warranty claims for material and legal defects (Sach- und Rechtsmängel) are excluded, except in cases where defects have been fraudulently concealed (arglistiges Verschweigen).

  3. Local-First & Non-Elevation Principle: ellmos-homebase-mcp is provided on an "as is" and "as available" basis without any express or implied warranty. Operators run Homebase in unprivileged user mode (RunAsInvoker) at their own discretion.

Coordinated Security Response SLA

For vulnerability reporting or security inquiries, our coordinated disclosure policy guarantees an initial response within 48 hours and triage within 5 business days:


Bundles and partners

Homebase MCP remains a standalone local-first MCP server. In the V4 composition it is an optional MCP access surface of the ellmos-memory-human-context-bundle: a configured system may use it to reach memory and human-context capabilities. This access role does not make Homebase the canonical owner of every memory, knowledge, state, routing or automation function; the selected host and system manifests retain those bindings.

Canonical or bundled engines are integration partners selected by explicit configuration, not implicit replacements for this server. Authoritative bundle membership, versions, profiles and private composition recipes remain in the corresponding bundle manifests. This public section is discovery-only.

Available Tools

51 tools
hb_api_discoverC

Auto-detect API schema from a base URL

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to inspect.

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Auto-detect' implies an outbound network call to a user-supplied URL, but the description does not disclose whether it performs a live request, what it does on failure, whether authentication is needed, or what format the detected schema takes. With zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, appropriate for a one-parameter tool. It is efficient, though the terseness leaves it short of the 'complete' end of the scale.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations, no output schema, and many similarly named siblings, the description is too thin. It does not explain return shape, error behavior, or how it differs from hb_plug_discover, leaving key agent decisions ungrounded.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (the single url parameter is documented), so the baseline is 3. The description adds no meaning beyond the schema, e.g. whether the URL must be reachable, include a scheme, or point to an OpenAPI/GraphQL endpoint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Auto-detect API schema') and its input ('from a base URL'), which is clear enough on its own. However, it does not distinguish this tool from the closely named sibling hb_plug_discover or from hb_api_probe/hb_api_export, all of which appear to touch API-schema concepts. Without that differentiation an agent must infer which discovery tool to use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use hb_api_discover versus hb_plug_discover, hb_api_probe or hb_api_export. No prerequisites, exclusions, or context are given, so the agent gets no routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_api_exportC

Export probe results as Markdown or JSON

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoOutput format.
probe_idYesProbe result ID.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full disclosure burden. 'Export' implies a read operation, but it never says whether output is returned inline or written to a file, whether permissions are needed, or how large a result set is handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the format options come last where they belong. It is perhaps terse to the point of under-specification, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter export tool with a fully documented schema and no output schema, the definition is minimally viable. It leaves open the practical question of where the exported content goes, which matters for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with only 2 params, one enum-constrained, so the schema already documents both fields adequately. The description adds nothing beyond restating the format options, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Export) and resource (probe results) plus the output formats. It is clearly distinguishable from related siblings like hb_api_probe or hb_api_history, though it does not name them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus hb_api_history or hb_test_results, nor any prerequisites such as needing a prior probe run. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_api_historyC

List previous probe results

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of results to return.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing about ordering (newest first?), retention window, scope (per endpoint, per session, global), or that it is a safe read-only listing. Only the 'previous' qualifier hints at recency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single four-word sentence, front-loaded with the verb and with zero filler. It is efficient, though the extreme brevity leaves the definition thin rather than structurally rich.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, read-only list tool with no output schema and no annotations, this is the minimum viable description. An agent still lacks return-shape, ordering, and scoping information that cannot be recovered from the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'limit' parameter is already fully documented in the schema, so the baseline of 3 applies. The description adds no extra meaning such as default behavior when results exceed the limit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ('List previous probe results') that an agent can distinguish from the sibling hb_api_probe, which presumably executes probes rather than retrieving history. However, it never explicitly contrasts itself with that sibling the way a 5-rated definition would.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no mention of prerequisites (e.g. that results come from hb_api_probe runs), and no alternatives named. The connection to probing is only inferable from the shared 'api' prefix.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_api_probeC

Probe a URL using all strategies (OpenAPI, wordlist, pattern, HATEOAS)

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to inspect.
strategiesNoDiscovery strategies to use.

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it only says 'Probe'. It does not disclose whether this makes network requests, requires auth, is read-only or destructive, rate limits, or what it returns. 'Using all strategies' hints at exhaustive behavior but nothing concrete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no waste. It is efficient, though it is arguably too terse for an unannotated tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a network-probing tool with zero annotations, no output schema, and an unclear relationship to hb_api_discover, the description is far too thin. It omits safety profile, auth needs, return shape, and when it differs from its sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both 'url' and 'strategies'. The description adds the names of the four default strategies, which is slightly useful context, but no format or interaction details beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Probe) and resource (URL), and enumerates the strategies applied. It is reasonably distinct from siblings like hb_api_discover, but the description never explains how 'probe' relates to 'discover', leaving the agent to guess whether they overlap or differ in scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of the clearly related sibling hb_api_discover. The agent has no basis for choosing this tool over hb_api_discover or the other hb_api_* tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_auto_list_chainsB

List available local automation chain definitions

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no read-only confirmation, no return shape, no filtering/pagination, no indication of what 'available' or 'local' means at runtime. 'List' weakly implies a non-mutating read, but that is an inference, not a disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler; the verb and the scoping adjective ('local') come first. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool this is close to sufficient, but with no output schema the description should say something about what a caller gets back (chain identifiers, definitions, ordering) to be fully actionable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the baseline there is nothing for the description to disambiguate. Schema coverage is 100% and there are no arguments to misunderstand.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('local automation chain definitions'), making the intent unambiguous. It does not explicitly distinguish itself from siblings like hb_auto_run or hb_auto_status, but the naming plus 'definitions' vs 'run'/'status' makes the distinction inferable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives such as hb_auto_run, hb_auto_status, or hb_auto_result. 'List available' weakly implies a discovery step before running a chain, but nothing is spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_auto_resultC

Get the recorded local automation run result

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesAutomation run ID.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It implies retrieval of an already-recorded run ('recorded') but says nothing about what happens with an unknown run_id, whether the result is only available after completion, or what the returned structure contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. Efficient, though extremely terse for a tool with no supporting annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a simple single-parameter getter, and the name plus description convey its core purpose. However, with no output schema and no annotations, an agent has no signal about the return shape or failure modes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a single required run_id documented as 'Automation run ID.' The description adds no format, sourcing, or example detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (recorded local automation run result), making it clearly distinct from sibling result/production tools like hb_auto_run and hb_auto_status. It does not explicitly name its siblings, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to call this versus hb_auto_status or hb_auto_list_chains, and no prerequisites stated. The agent must infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_auto_runB

Queue a local automation chain plan without executing external backends

ParametersJSON Schema
NameRequiredDescriptionDefault
chainYesAutomation chain name.
inputNoInput payload for the operation.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose one meaningful trait: the call queues rather than executes external backends, implying no external side effects. It falls short on auth requirements, whether the named chain must pre-exist, and how/when effects surface — real gaps for a queue-style mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It is efficient, though the terse phrasing leaves the queue-vs-execute nuance slightly opaque.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a queueing operation with no output schema and no annotations, the description should at least hint at follow-up (e.g., hb_auto_status or hb_auto_result to observe progress). It also never resolves the name/behavior tension between 'run' and 'queue', leaving the agent unsure whether this starts work or merely registers it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents 'chain' and 'input' adequately. The description adds no syntax, format, or constraint detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Queue') and resource ('local automation chain plan') and clarifies scope ('without executing external backends'). This distinguishes it well from hb_auto_list_chains/hb_auto_status/hb_auto_result within the hb_auto family, though the 'run' in the name sits oddly against 'queue'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this versus siblings like hb_plug_run, hb_garden_run, or the other hb_auto tools. The 'without executing external backends' phrase implies a deferred-execution scenario but does not state when to prefer it or what prerequisites exist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_auto_statusB

Check local automation run status

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesAutomation run ID.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden. 'Check' implies a read-only operation, but this is not explicitly stated, and there is no mention of side effects, authentication, rate limits, or return behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded phrase with no wasted words. It is appropriately sized for a simple status-check tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity and the lack of an output schema, the description is minimally adequate. However, it does not clarify what status values are returned, what 'local' means, or confirm the operation is read-only, leaving some gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single parameter (run_id) is already documented in the schema. The description adds no additional meaning, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Check') and resource ('local automation run status'), making it clear what the tool does. It does not explicitly differentiate from siblings like hb_auto_result or hb_auto_run, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives such as hb_auto_result or hb_auto_list_chains. Usage is only implied by the name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_conn_listB

List configured connectors

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and no output schema, the description carries the full disclosure burden but says nothing about return shape, pagination, filtering scope, or whether listing triggers any side effects. A read-only enumeration is implied by 'List' but never confirmed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single three-word phrase with no waste is well front-loaded, though it is arguably under-specified rather than concise. It earns its place but adds little beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool with no output schema the description is minimally viable, but an agent still lacks any sense of what a connector record contains or how results are scoped. It does not misuse complexity, but it leaves obvious questions open.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no semantics to explain beyond what the schema provides. Baseline of 4 applies for a parameterless definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (List) and resource (configured connectors), so an agent can tell it apart from the action-oriented siblings hb_conn_send/hb_conn_receive and even the state-oriented hb_conn_status. It stops short of clarifying what a 'connector' is or what fields are returned, so it is clear but not fully disambiguating.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this versus hb_conn_status or the other conn_* siblings. The agent must infer that this is the enumeration entry point for connectors purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_conn_receiveB

Get recent local inbox messages for a connector

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of results to return.
connectorYesConnector name.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says 'Get' and 'recent local inbox messages' but does not clarify whether receiving consumes or deletes messages, whether it marks them as read, whether the connector must exist, or any auth or rate-limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It states the action, scope, and target directly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and the schema covers both parameters, but with no annotations and no output schema, the description should say more about what 'receive' means behaviorally. It is adequate for basic selection but leaves gaps around return format and side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both the connector and limit parameters. The description adds the scoping phrase 'for a connector' and 'recent' implies limiting results, but it does not add syntax, format, or default details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (recent local inbox messages) with scope (for a connector). It is clearly different from sending or listing connectors, but it does not name or distinguish itself from siblings like hb_conn_list or hb_conn_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies retrieval usage but gives no explicit when-to-use guidance, no when-not-to-use conditions, and no alternatives. It does not mention hb_conn_send or other connector tools that an agent might confuse it with.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_conn_sendB

Queue a message for a connector without network delivery

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoConnector target, channel, or chat ID.
messageYesMessage text.
connectorYesConnector name.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the core side effect (queued, not sent over the network), but says nothing about persistence, auth requirements, whether queued messages are later delivered, or what happens on duplicates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler. The key behavioral constraint (queue, no network) appears immediately and every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description should carry more behavioral weight for a mutation-style tool. It omits what the queued message becomes, the distinction between connector and target, and any return/confirmation semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents connector, message, and target. The description adds no syntax or format detail beyond what the schema supplies, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb (queue) and resource (a message for a connector) plus a distinguishing scope constraint: no network delivery. That separates it from siblings like hb_conn_receive or hb_conn_status, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without network delivery' hints at the use case, but there is no explicit when-to-use, when-not-to-use, or named alternative (e.g., vs hb_conn_receive). The agent must infer the routing decision on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_conn_statusB

Check local connector health and queue counts

ParametersJSON Schema
NameRequiredDescriptionDefault
connectorYesConnector name.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It reveals that the tool reports health and queue counts, implying a read-only operation, but does not state whether it requires authentication, has side effects, rate limits, or how errors are handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no filler and the core purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter status tool, the description is minimally adequate. However, it does not explain what 'health' entails or what the queue counts represent, and it lacks usage guidance, leaving some contextual gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single required 'connector' parameter, so the schema already provides adequate parameter semantics. The description adds no additional meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Check') and two specific resources ('local connector health' and 'queue counts'), making the tool's purpose clear. It does not explicitly differentiate from siblings like hb_conn_list or hb_conn_receive, but the status/health focus is distinctive.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. The implied usage is a status check, but no conditions, prerequisites, or alternative tools are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_garden_findC

Search the garden

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of results to return.
queryYesSearch query.

TDQS

C2.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the entire behavioral burden, and it discloses nothing: not whether this is a read-only operation, whether results are ranked, how pagination or the "limit" default behaves, or what permissions are needed. For a search tool with zero annotation coverage this is a complete gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words with zero padding or repetition, so there is no bloat to penalize. However, it is a bare fragment with no structure, and the single phrase does not earn its place because it conveys nothing the tool name did not already convey.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description should at least characterize the search domain and result shape, and it does neither. The only fully specified surface is the two-parameter input schema, leaving the agent without enough context to invoke this confidently in a 50-tool environment.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — both "query" and "limit" (including its default of 10) are documented in the schema itself. The description adds no syntax, matching, or format guidance beyond that, so the baseline 3 applies where the schema does all the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Search the garden" does state a verb and a nominal resource, but "garden" is never defined and the phrase effectively restates the tool name hb_garden_find. With no differentiation from siblings like hb_kb_search, hb_garden_get, or hb_garden_run, an agent cannot tell what domain this searches or how it differs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative-tool guidance anywhere in the description. The verb "Search" faintly implies a retrieval use case, but nothing tells the agent when to prefer this over hb_kb_search or hb_garden_get.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_garden_getC

Get a specific entry by key

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesEntry key.

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not say whether the read is safe, what happens when the key is absent (error vs null), or whether the lookup is case-sensitive or namespaced — all meaningful for a retrieval tool with a free-form string key.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with zero filler — verb and lookup mechanism come first. It is efficient but so terse that there was no opportunity to front-load anything more useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, nothing in the structured data tells the agent what an entry contains or how failures surface. For a tool whose only behavior is returning a keyed object, the description should at least sketch the return shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter is documented as "Entry key." The description adds no format, prefix, or case rules beyond that, so it sits at the baseline appropriate when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and the keyed lookup, which does contrast with hb_garden_find's implied search. However, the resource is only described as "a specific entry" with no indication of what a garden entry is or which namespace it lives in, so the agent gets the mechanic but not the domain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus hb_garden_find (search) or hb_garden_put (write). The description implies read-by-key but never states the condition that selects this tool over its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_garden_putC

Store or update an entry

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesEntry key.
valueYesEntry value.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does imply upsert semantics ('store or update'), which is a genuine behavioral trait, but says nothing about overwrite behavior on existing keys, permissions, durability, or what is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence that is front-loaded and waste-free, but it is thin rather than tight — the brevity comes from under-specification rather than disciplined editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, the description should at minimum clarify overwrite behavior, scope of the store, and error cases. None of that is present, leaving real gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the two parameters (key, value) are already documented in the schema. The description adds no format, length, or naming-convention details beyond that, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb phrase 'store or update' plus 'an entry' does communicate a key-value upsert, which implicitly distinguishes it from read-only siblings like hb_garden_get and hb_garden_find. However, 'entry' and the 'garden' namespace are never grounded in a domain, so the purpose is only vaguely conveyed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many overlapping write/state siblings (hb_mem_store, hb_state_mem_set, hb_kb_ingest). No prerequisites, no exclusions, no mention of when an update is preferable to a create.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_garden_runC

Run a stored command if execution is explicitly enabled

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesEntry key.

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It hints that execution may be gated by a flag, but does not clarify what 'explicitly enabled' means, who can enable it, what happens when disabled (error vs no-op), or any side effects. For an execution tool this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence that front-loads the action and appends the condition. It is efficient, though the vagueness of 'stored command' stems from under-specification rather than verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given it is an execution tool with no annotations, no output schema, and only a one-word parameter description, the definition is under-specified. An agent lacks enough to know when to invoke it, what authority is required, or what a disabled state yields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% but the only parameter description is 'Entry key.', which is terse. The description adds no meaning about what the key references or how it maps to a stored command. Baseline 3 applies since schema covers the structure, but no extra value is added.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb ('Run') and a resource ('stored command'), which is clearer than the bare name. But 'stored command' is vague and the description does not distinguish this tool from siblings like hb_plug_run, hb_test_run, or hb_auto_run. An agent cannot confidently tell when to select this over those.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is a conditional clause ('if execution is explicitly enabled'), but no guidance on which sibling to use, no prerequisites explained, and no explicit when-to-use versus when-not. The condition is stated without clarifying how or where execution is enabled.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_kb_getB

Get a single knowledge entry with metadata

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesNumeric entry ID.
agent_idNoOptional agent filter

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. 'Get' clearly implies a non-mutating read, and 'with metadata' hints at the return shape, but there is no disclosure of auth requirements, behavior for missing IDs, or what 'metadata' includes. For a simple read the exposure is modest, but the gap is real.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words; the action and resource come first. It is efficient rather than padded, though it errs slightly toward under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with no output schema and no annotations, the description is minimally sufficient but thin: the 'with metadata' phrase is the only clue about the response, and nothing covers the optional agent_id's effect or failure modes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with both 'id' and the optional 'agent_id' documented in the schema, so the baseline of 3 applies. The description adds no meaning about the ID format or what 'agent filter' actually scopes, so it neither helps nor hurts.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (a single knowledge entry), and the word 'single' implicitly distinguishes it from the sibling list/search tools hb_kb_list and hb_kb_search. It stops short of naming those siblings explicitly, so an agent must infer the retrieval-by-ID vs. query distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: fetching one entry by ID. There is no statement of when to prefer this over hb_kb_list or hb_kb_search, nor any prerequisite or error condition (e.g. non-existent ID). Adequate but leaves routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_kb_ingestC

Add new knowledge (text, URL, or file content)

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNoList of tags.
sourceNoURL or file path
contentYesContent text.
agent_idNoAgent identifier for shared-knowledge provenance

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full disclosure burden. It implies a persistent mutation but says nothing about indexing, deduplication, chunking, provenance handling, permissions, or whether the operation is reversible — all material for an ingest tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the core verb front-loaded and no wasted words. It is efficient, though the brevity edges toward under-specification rather than optimal density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With four parameters, no annotations, and no output schema, the description should at minimum indicate what ingestion produces (e.g., a knowledge ID) and where the content lands. It instead covers only the action and input formats, leaving the agent short of what it needs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters; baseline is 3. The description's mention of 'URL or file content' loosely maps to the source/content fields but adds no format, size, or tagging syntax beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb and resource ('Add new knowledge') and enumerates the accepted content forms (text, URL, file content), so the action is unambiguous. It does not distinguish itself from adjacent write-oriented siblings such as hb_mem_store or hb_garden_put, which prevents a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to ingest into the knowledge base versus storing to memory (hb_mem_store) or the garden (hb_garden_put), nor any prerequisite or ordering advice. The agent is left to infer the appropriate context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_kb_listC

List tags in the knowledge base

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idNoOptional agent filter

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no return format, no pagination, no statement that it is a safe read, and no behavior when agent_id is omitted or matches nothing. Only a bare "list" implies read-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no waste. It is efficient, though its brevity edges toward under-specification rather than genuine economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description leaves critical questions unanswered: what entity is actually listed, what the response contains, and whether results are paginated. An agent could invoke it but cannot predict its output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with a single optional agent_id documented as "Optional agent filter", so the schema does the work; the description adds no clarification of what the filter means (filter tags by agent?) or what happens without it. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb+resource ("List tags"), but the resource is ambiguous against the tool name hb_kb_list and the sibling set (hb_kb_search, hb_kb_get, hb_kb_ingest) – an agent cannot tell whether this lists tags, KB entries, or tag-filtered entries. It also doesn't distinguish itself from hb_kb_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no exclusions, and no mention of alternatives such as hb_kb_search for content lookup or hb_kb_get for a specific entry. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_lock_checkA

Check whether a project directory currently has an active lock -- call this BEFORE any write action there (read-only).

ParametersJSON Schema
NameRequiredDescriptionDefault
project_dirYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that the operation is read-only and should precede writes, but omits authentication needs, return semantics, and failure behavior, leaving clear gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that states the action, scope, timing prerequisite, and read-only nature without any wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, one-parameter read tool with no output schema and no annotations, the description covers purpose, timing, and safety sufficiently. It could still say more about the return value or lock semantics, but the essential information is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter project_dir is only named in the schema. The description adds that it refers to a project directory, which clarifies meaning, but does not specify path format, relative vs. absolute, or any other constraint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: checking whether a project directory has an active lock. It does not explicitly distinguish itself from the sibling tool hb_lock_list, so it falls short of a 5 despite being clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit timing guidance: call this BEFORE any write action there. However, it does not mention alternatives or when-not to use it, such as hb_lock_list, so it is not fully complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_lock_listA

List all currently active locks across configured roots (read-only).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden: it usefully discloses that the operation is read-only and scoped to configured roots, which prevents an agent from expecting arbitrary paths. It says nothing about return shape, ordering, or whether stale locks are included.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence covering verb, resource, scope, and safety trait with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only listing tool with no output schema, the description covers what is listed, where it is scoped, and that it is safe. Only minor gaps remain (e.g., what fields each lock entry contains), which is acceptable at this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so no parameter-level explanation is required; baseline is 4. The description's mention of 'configured roots' correctly signals that scope is implicit rather than caller-supplied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (currently active locks) with scope (across configured roots). It does not explicitly differentiate from the sibling hb_lock_check, but the 'all currently active' phrasing largely conveys the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Use is implied by the listing semantics, but there is no explicit when-to-use guidance or routing hint against hb_lock_check (which presumably inspects a single lock). No exclusions or prerequisites are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_mem_consolidateA

Decay memory confidence and prune low-confidence entries (dry_run previews; dry_run=false applies)

ParametersJSON Schema
NameRequiredDescriptionDefault
decayNoAbsolute confidence reduction per entry.
dry_runNoPreview without applying. Set false to apply.
agent_idNoOptional agent filter
min_confidenceNoEntries below this (after decay) are pruned.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It does disclose the key behavioral trait: a mutation happens only when dry_run=false, otherwise it is a preview, and that entries get pruned. It stops short of stating irreversibility, whether pruning is permanent, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words; the dry_run safety note is appended efficiently. The nested parenthetical is slightly dense but earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-param mutation tool with no output schema and no annotations, the description covers purpose and the dry-run safety valve, and the schema fully documents parameters. Missing only irreversibility/scope caveats about what pruning destroys.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents decay, dry_run, agent_id, and min_confidence with bounds and defaults. The description only restates the dry_run semantics, adding little beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Decay memory confidence and prune low-confidence entries.' An agent can tell this is a memory-maintenance operation distinct from hb_mem_store or hb_mem_merge. However, it does not explicitly differentiate itself from the closest sibling (hb_mem_merge), which also mutates memory state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(dry_run previews; dry_run=false applies)' tells the agent how to invoke it safely, implying usage context. But it never states when this tool should be chosen over hb_mem_merge or other memory siblings, nor any prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_mem_contextC

Generate compact context string for prompt injection

ParametersJSON Schema
NameRequiredDescriptionDefault
focusNoOptional topic focus
agent_idNoOptional agent filter
max_tokensNoApproximate token budget

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not state what data source the context is drawn from, whether it is read-only, whether repeated calls mutate or consume memory, or what happens when the token budget is exceeded.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with no filler, and the artifact and its purpose are front-loaded. It is appropriately sized but arguably under-specified rather than over-worded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter tool with no output schema, the description should explain the returned string's shape and the source of context. It conveys the output type but leaves provenance and budget behavior unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so focus, agent_id, and max_tokens are already documented in the schema. The description adds nothing about semantics such as how the token budget is honored or what focus filtering actually selects.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb (generate) plus specific artifact (compact context string) and intended use (prompt injection). It is distinguishable from hb_mem_store/hb_mem_query by producing a consumable string rather than storing or filtering records, though it never names those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to call this versus hb_mem_query or hb_state_mem_get, nor any prerequisite (e.g., existing stored memories) or ordering advice. The agent must infer the trigger condition from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_mem_mergeB

Confidence-based merge of duplicate memories (dry_run previews; dry_run=false applies the merge)

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNoPreview merge without applying. Set false to apply.
agent_idNoOptional agent filter

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that dry_run previews and dry_run=false applies the merge, which is meaningful mutation context. However, it does not state irreversibility, what happens to merged memories, permission requirements, or conflict resolution beyond 'confidence-based'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with a parenthetical that covers both purpose and dry_run behavior. Every word earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter memory-mutation tool with no annotations and no output schema, the description is minimally viable. It covers the core action and dry_run semantics but omits sibling routing guidance and behavioral caveats such as irreversibility.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters fully. The description adds the 'confidence-based' framing but does not mention agent_id or add syntax/format detail beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action and resource: 'merge of duplicate memories', and adds 'confidence-based' to clarify the mechanism. It does not distinguish this tool from the sibling hb_mem_consolidate, which also deals with memory consolidation, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains dry_run behavior but gives no guidance on when to choose this merge tool over alternatives like hb_mem_consolidate or hb_mem_query. The agent must infer the use case from the phrase 'duplicate memories' alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_mem_queryC

Search memory by keyword

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of results to return.
queryYesSearch query
agent_idNoOptional agent filter
categoryNoCategory filter or category to store.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it discloses almost nothing. "Search" implies a read-only operation, but there is no statement about result ordering, pagination via limit, category semantics, or what happens when no match exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single phrase is front-loaded and waste-free, which is good, but the brevity here reads as under-specification rather than discipline. A slightly longer statement could route usage without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter tool with no annotations and no output schema, one line is insufficient. An agent lacks guidance on filtering by category or agent_id, how limit interacts with result ordering, and how this search relates to the other mem and kb siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented in the schema, making 3 the baseline. The word "keyword" hints at query matching style but adds no format or syntax detail beyond the schema's "Search query".

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Search memory by keyword" names a verb (search) and resource (memory), so the core purpose is legible. But with siblings like hb_mem_context, hb_mem_merge, hb_mem_consolidate, and hb_kb_search, a one-phrase description does nothing to distinguish which memory surface this targets or how it differs from context retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative routing is given. An agent cannot tell from this description whether to reach for hb_mem_query, hb_mem_context, or hb_kb_search first, and no prerequisites or result expectations are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_mem_storeC

Store a fact, lesson, or working memory entry

ParametersJSON Schema
NameRequiredDescriptionDefault
contentYesThe content to store
agent_idNoAgent identifier for shared-memory provenance
categoryYesMemory category
confidenceNoConfidence score (0-1)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it says nothing about persistence semantics, deduplication, whether entries are append-only or overwritable, or what provenance/agent_id actually does. 'Store' implies a mutation but no safety, auth, or reversibility context is given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler or redundancy. It is efficient, though the terseness is partly what leaves the behavioral and usage gaps unfilled.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with four parameters, zero annotations, and no output schema, the description does not cover enough ground. It never addresses what happens on store, whether duplicates are handled, or what the caller gets back.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented in the schema, establishing the baseline of 3. The description only re-lists the category values that the schema's enum already covers, adding no new meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Store') and the resource is clearly scoped by the enumeration of what may be stored ('fact, lesson, or working memory entry'). The verb alone distinguishes it from read-oriented siblings like hb_mem_query and hb_mem_context, though it never names those siblings explicitly to confirm the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to store versus query, merge, consolidate, or route memory, and no preconditions. An agent must infer the write-vs-read split purely from the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_plug_discoverC

Scan a local directory for plugin metadata

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoLocal filesystem path.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. 'Scan' implies a read-only local operation, but the description does not state permissions, side effects, recursion behavior, error handling, or output shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It is appropriately concise, though extremely terse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple discovery tool with no output schema and no annotations, the description leaves key context unstated: default path behavior, what 'plugin metadata' includes, and whether the operation is read-only. These gaps matter for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single path parameter, which is fully documented as a local filesystem path. The description adds no further parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (Scan), resource (local directory), and target (plugin metadata). It is clear what the tool does, though it does not explicitly differentiate itself from siblings like hb_plug_list or hb_plug_info.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives, nor on prerequisites or exclusions. The description only states what it does, leaving usage entirely implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_plug_infoC

Get local plugin metadata

ParametersJSON Schema
NameRequiredDescriptionDefault
pluginYesPlugin name.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read-only operation but does not disclose return format, permissions required, caching behavior, or error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no wasted words. It is appropriately concise for a simple retrieve operation, though its terseness contributes to gaps elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description should explain what metadata is returned or at least characterize the return value. It also lacks basic usage context and behavioral detail, leaving the agent with minimal information beyond the tool name and schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single 'plugin' parameter is already well documented in the schema. The description adds no additional meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('local plugin metadata'), distinguishing it from siblings like hb_plug_list and hb_plug_run by its metadata focus. However, it does not explicitly name or contrast with any sibling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no guidance on when to use this tool versus hb_plug_list, hb_plug_discover, or hb_plug_run. Usage is only implied by the phrase 'Get local plugin metadata', with no exclusions or alternatives mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_plug_listB

List discovered local plugins

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It does not state whether the operation is read-only, whether it triggers or refreshes plugin discovery, whether results are cached, or what permissions are required. The minimal phrasing leaves key behavioral traits undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or repetition. Every word contributes to stating the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple listing tool, the description is minimally adequate but leaves important contextual gaps. Most notably, it does not clarify how it relates to hb_plug_discover or whether listing implies prior discovery, and absent an output schema it gives no hint of what is returned.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so there are no parameter semantics to document. The baseline score of 4 applies for a zero-parameter tool whose schema also has nothing to clarify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'List' plus 'local plugins'. It also scopes the resource to 'discovered' plugins, which helps distinguish it from raw discovery. However, it does not explicitly differentiate itself from siblings like hb_plug_discover or hb_plug_info.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus hb_plug_discover, hb_plug_info, or hb_plug_run. The intended usage is only implied by the word 'List' and the 'discovered' qualifier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_plug_runB

Record a plugin dry-run without executing plugin code

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNoOptional argument object.
pluginYesPlugin name.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that plugin code is not executed, which is important for safety and intent. It does not disclose persistence side effects, permissions, idempotency, or what 'record' entails beyond the non-execution guarantee.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It states the action and the key constraint immediately. It is appropriately sized for a two-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description must carry more context than usual. It covers the critical non-execution behavior but does not explain what the recorded dry-run returns, whether it persists state, or how failures are reported. It is adequate but leaves meaningful gaps for an agent invoking a write-like operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the input schema. The description adds no additional meaning about the plugin name or args object beyond what the schema provides. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb plus resource: 'Record a plugin dry-run.' It also clarifies the critical boundary that plugin code is not executed. It does not explicitly distinguish itself from siblings like hb_plug_list, hb_plug_info, or hb_garden_run, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without executing plugin code' implies this is for dry-run recording rather than real plugin execution. However, there is no explicit when-to-use guidance, no when-not-to-use guidance, and no alternative tool is named. The agent must infer the usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_policy_listC

List policy-registry entries (read-only, canonical-only), optionally filtered.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
queryNoSearch query.
scopeNo
consumerNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses two behavioral traits—read-only and canonical-only (excluding non-canonical/overridden entries)—but says nothing about permissions, pagination, result count limits, or ordering for what is presumably a large registry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the core action front-loaded and the key constraint parenthesized. Nothing is wasted, though it is arguably too terse for a four-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With four parameters, 25% schema coverage, no annotations, and no output schema, the definition omits far too much: the meaning of kind/scope/consumer, the shape of returned entries, and any pagination behavior. An agent can invoke it only by guessing at filter semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only 'query' is documented), so three of four parameters (kind, scope, consumer) are opaque. The description says entries can be filtered but never explains what kinds, scopes, or consumers are valid or how they combine, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List policy-registry entries'), so an agent knows this is a read/listing operation over the policy registry. It does not name or contrast with the sibling hb_policy_resolve, so sibling differentiation is missing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'optionally filtered' hints that filters are supported but gives no when-to-use guidance, no conditions, and no mention of the alternative hb_policy_resolve. The agent must infer that resolution (not listing) is the alternative path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_policy_resolveC

Resolve the authoritative policy/rule/decision for a scope via the real policy-registry engine (read-only, canonical-only).

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNoSearch query.
scopeYes
consumerNo
required_kindNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It discloses that the operation is 'read-only' and 'canonical-only', which are useful safety and filtering traits, but it omits auth requirements, error behavior, and what 'resolve' actually returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded with the core action and resource, and the parenthetical qualifiers are compact. It contains no filler, though its extreme brevity leaves much unsaid.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given four parameters, one required field, no annotations, no output schema, and low schema coverage, the description is far too sparse. It does not explain how to use scope, consumer, required_kind, or query, nor what a resolved result looks like.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (just the 'query' parameter). The description mentions 'for a scope' but adds no real meaning beyond the parameter name; 'consumer' and 'required_kind' are left completely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Resolve') and resource ('authoritative policy/rule/decision for a scope'), making the tool's purpose clear. It implicitly distinguishes itself from sibling 'hb_policy_list' by indicating resolution rather than listing, but does not explicitly name that alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as hb_policy_list. The description only states what it does, leaving usage context entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_route_evaluateC

Rate a response for routing feedback loop (epsilon-greedy learning)

ParametersJSON Schema
NameRequiredDescriptionDefault
qualityYesQuality rating from 0 to 1.
route_idYesRouting decision ID.

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose a meaningful behavioral trait: this feeds a routing feedback loop / learning process, so an agent can infer it mutates routing state and affects future selections. It does not say whether the effect is immediate, persistent, or reversible, which leaves a gap for a write-like operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single, front-loaded sentence with no filler or repetition. It is efficient, though the brevity is partly why other dimensions are thin rather than a sign of strong structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-mutating learning tool with no annotations and no output schema, the description omits what the rating does, what is returned, and its prerequisite relationship to hb_route_select. The schema covers the inputs, but the behavioral contract an agent needs before invoking it is largely missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already defines both parameters (route_id, quality with a 0-1 range). The description adds no syntax, format, or constraint details beyond the schema. This meets the baseline 3 when structured fields do the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb ("Rate") and a target ("a response for routing feedback loop"), so the general purpose is inferable. However, "a response" is ambiguous and it does not explicitly distinguish itself from close siblings like hb_route_select or hb_route_stats. The purpose is vague-but-directional rather than specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit statement of when to call this tool, what must precede it (e.g., a prior hb_route_select), or which sibling to use instead. The "epsilon-greedy learning" parenthetical is context, not guidance. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_route_selectC

Analyze a prompt and recommend the best model/provider

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesThe prompt to route
constraintsNoOptional constraints (max_tokens, speed, cost)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not disclose whether the call is a read-only analysis, how constraints influence the recommendation, whether there are latency or rate-limit implications, or what triggers a reroute.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler or repetition. Appropriately sized for the tool's simplicity, though minimal enough that no structure beyond one clause exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter routing tool with no output schema and no annotations, the description should at least hint at the shape of the recommendation or how constraints affect it. The schema covers inputs, but the behavioral and output picture remains thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented, including the nested constraints object shape (max_tokens, speed, cost). The description adds nothing beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (analyze/recommend) and resource (model/provider routing) for an input prompt. It is clear what the tool does, but it does not distinguish itself from the adjacent hb_route_evaluate or hb_route_stats siblings in the same family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus hb_route_evaluate, hb_route_stats, or any other sibling. No prerequisites, no exclusions, and nothing about when routing is or is not appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_route_statsC

Get routing statistics and learning progress

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no read-only assurance, no auth requirements, no indication of whether stats are cumulative or per-session, and no return shape. For a no-annotation tool this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence with no padding, which is good, but the brevity comes at the cost of under-specification rather than genuine economy. Nothing is front-loaded because there is only one clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description is the only source of information about what this tool returns, and it does not describe the statistics, their format, or their freshness. An agent could call it, but could not anticipate the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to add beyond the empty schema. Baseline of 4 applies; nothing is missing or misleading on this dimension.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

It names a verb (Get) and a resource area (routing statistics and learning progress), which loosely separates it from hb_route_select and hb_route_evaluate. But 'routing statistics' and 'learning progress' are vague — it never says what is measured, over what window, or in what scope, so an agent cannot tell exactly what it retrieves.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus hb_route_select, hb_route_evaluate, or any sibling. The agent must infer that a stats getter is for observation rather than route selection, with no explicit trigger condition or exclusion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_state_dispatchC

Return connector-dispatch status

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoChannel/chat ID
messageYesMessage text.
connectorYesConnector name

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Return' implies a read, but it does not disclose whether this requires auth, whether it mutates dispatch state, why a 'message' is required to query status, or what the response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single five-word sentence, fully front-loaded with no filler. It is efficient, though its brevity borders on under-specification rather than tight conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and a confusing required 'message' parameter for what is called a status query, the description leaves too much unexplained. An agent cannot confidently determine what this returns or why it must supply message text.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (target, message, connector) are already documented in the schema. The description adds no meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb ('Return') and a resource ('connector-dispatch status'), so the purpose is minimally identifiable. However, 'connector-dispatch' is undefined jargon and it does not distinguish itself from close siblings like hb_conn_status, which also appears to report connector state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given. There is no mention of prerequisites, when to prefer this over hb_conn_status or the other hb_state_* tools, or what a 'dispatch' refers to.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_state_mem_getC

Get state memory entries

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNoMemory type.
queryNoSearch query.
agent_idNoOptional agent filter

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies a read operation via 'Get' but says nothing about permissions, side effects, pagination, or return behavior, leaving disclosure minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded phrase with no wasted words. However, it is extremely terse for a tool with three optional filters and no annotations, so it is concise without being adequately structured for selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description is too thin. It does not explain what state memory entries are, how type/query/agent_id interact, or what is returned, leaving significant ambiguity against sibling memory tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters are documented there, including the 'type' enum. The description adds no syntax or filter semantics beyond what the schema already provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb ('Get') and resource ('state memory entries'), enough to distinguish from hb_state_mem_set, but it does not differentiate from other memory query siblings such as hb_mem_query or hb_mem_context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no alternatives named, and no prerequisites are stated. It is not misleading, but an agent gets no help choosing this tool over hb_mem_query, hb_mem_context, or related state tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_state_mem_setC

Store a state memory entry

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesEntry key.
typeNoMemory type.
valueYesEntry value.
agent_idNoAgent identifier for state-memory provenance

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, yet it only says 'Store' – nothing about whether an existing key is overwritten, whether this requires provenance (the agent_id param hints at it), idempotency, or return behavior. A mutation tool with zero annotation coverage needs more.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence is front-loaded and waste-free, but it is under-specified rather than concise – the terseness leaves obvious gaps for a mutating, four-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutation tool with four parameters, no annotations, and no output schema demands explanation of overwrite behavior, provenance, and the fact/lesson distinction, none of which is present. The definition is incomplete for its complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are documented in the schema (key, type enum, value, agent_id). The description adds no syntax, format, or provenance semantics beyond what the schema already provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb ('Store') and resource ('state memory entry'), but the phrasing is nearly a restatement of the tool name and does not distinguish it from siblings like hb_mem_store or hb_state_mem_get. It conveys the basic action but nothing about scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus hb_mem_store (general memory) or hb_state_mem_get (retrieval). The sibling namespace makes routing ambiguous, and the description offers no help resolving it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_state_task_createC

Create a new task

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYesTask title.
agent_idNoAgent identifier for task provenance
priorityNoTask priority.
descriptionNoOptional description.

TDQS

C2.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden, and it discloses nothing: not whether the task is persisted, whether it is dispatchable, whether agent_id is required for provenance, or what side effects creation has. For a mutation tool with zero annotation coverage, this is a critical gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded and wastes no words, but it is under-specification rather than genuine conciseness — the brevity comes at the cost of all useful content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter mutation tool with no annotations and no output schema, the description is far too thin: it omits provenance semantics for agent_id, priority defaults, and any indication of the creation result. An agent could not call this confidently without reading the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (title, agent_id, priority, description) are already documented in the schema. The description adds no format, constraint, or relationship detail beyond that, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Create a new task" essentially restates the tool name hb_state_task_create with no added specificity. It gives a verb and resource but does nothing to distinguish the task-creation semantics from siblings like hb_state_task_list or hb_state_task_update, and it never explains what a "task" is in this system.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all: no prerequisites, no mention of the sibling hb_state_task_update or hb_state_dispatch as alternatives, and no note on when creating a task is appropriate versus dispatching one. The agent gets no routing help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_state_task_listC

List tasks with optional status filter

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNoStatus filter.
agent_idNoOptional agent filter

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the entire behavioral burden. It does not disclose default behavior when status is omitted (does it default to 'all'?), pagination, ordering, or return shape. For a tool with zero annotation coverage this is a substantial gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single terse phrase, front-loaded with the verb and resource and nothing wasted. It is efficient, though the fragmentary style leaves room for a useful clause or two about defaults.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only list tool with a fully documented schema and no output schema, the minimum viable information is present. It still lacks the default-status behavior and any hint of the result set, which an agent calling it blind would want to know.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the schema and the baseline is 3. The description restates the status filter (already in the schema with a full enum) and omits mention of the agent_id filter entirely, adding no meaning beyond the structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List tasks') plus a scope qualifier, so an agent immediately knows this is a read-only enumeration of tasks. It does not, however, differentiate itself from the sibling task tools (hb_state_task_create, hb_state_task_update, hb_state_dispatch), which the description never mentions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'with optional status filter' implies the filters exist but gives no when-to-use guidance, no prerequisites, and no alternative tools to consider. Nothing tells the agent when this tool is the right choice over hb_state_task_create or hb_state_task_update.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_state_task_updateC

Update a task (status, description, priority)

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNoStatus filter.
task_idYesTask ID.
agent_idNoOptional agent filter
priorityNoTask priority.
descriptionNoOptional description.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. The key question for an update tool — whether omitted fields are left unchanged or cleared — is not answered, nor is there any mention of permissions, reversibility, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the action and affected fields front-loaded; nothing is wasted. It is arguably under-specified rather than padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter mutation with no annotations and no output schema, the description is thin. Partial-update semantics, the effect of each enum transition (open/in_progress/done), and the role of the filter-like fields are all unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description lists three of the five fields but says nothing that adds meaning beyond the schema, and it ignores agent_id, which the schema oddly labels 'Optional agent filter' for an update operation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Update a task') and names the three mutable fields, so the agent knows it is a mutation of an existing task. It does not differentiate itself from siblings like hb_state_task_create or hb_state_task_list, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus hb_state_task_create or hb_state_task_list, nor any stated prerequisites. Usage is only implied by the verb 'Update'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_swarm_consensusC

Get majority vote from multiple independent agents

ParametersJSON Schema
NameRequiredDescriptionDefault
votersNoNumber of independent voters to plan.
questionYesQuestion to evaluate by consensus.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not disclose cost, latency, determinism, tie-breaking behavior, or whether the operation is read-only or spawns real external agents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It is appropriately sized for the amount of information it chooses to convey.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-agent consensus tool with no output schema and no annotations, the one-line description is too sparse. It omits result format, tie behavior, and how the independent voters are realized.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented. The description adds only the concept of majority vote and does not extend parameter meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Get) and resource (majority vote from multiple independent agents), so the basic action is clear. However, it does not differentiate this tool from sibling swarm tools like hb_swarm_parallel, hb_swarm_hierarchy, or hb_swarm_stigmergy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the other swarm coordination tools. The description gives no context, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_swarm_hierarchyD

Boss-worker delegation pattern

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYesTask description.
subtasksNoOptional subtasks for hierarchy planning.

TDQS

D1.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations and no output schema, so the description carries the full burden and delivers nothing. It does not disclose whether tasks are actually executed, whether worker agents are spawned, what the hierarchy depth means, or whether the call is synchronous — all critical for an orchestration tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four words is not conciseness but under-specification; there is no front-loaded statement of what the tool does or returns, so nothing earns its place because nothing is there.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a multi-agent orchestration tool with zero annotations, no output schema, and only a fragment of description text. Everything an agent needs to invoke it correctly (execution semantics, agent spawning, result shape, relationship to the other swarm tools) is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with two well-labelled parameters (`task`, `subtasks`), so the schema does the documentation work. The description adds no meaning beyond it, which is the baseline-3 case for a fully covered schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase "Boss-worker delegation pattern" essentially restates the tool name (swarm_hierarchy) in different words rather than stating a verb and resource. It hints at an orchestration strategy but never says whether the tool plans, executes, or just configures the hierarchy, and it does not distinguish itself from siblings like hb_swarm_parallel or hb_swarm_consensus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no routing to the sibling swarm patterns. An agent cannot tell from this text why it would pick hierarchy over parallel or consensus.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_swarm_parallelC

Split task into chunks and process in parallel

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYesTask description.
chunksYesTask chunks to assign across workers.
workersNoNumber of parallel workers to plan.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it delivers only a one-line mechanism. It says nothing about side effects, whether workers are real concurrency or just a planning hint, failure/partial-completion behavior, or result aggregation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no wasted words and the core action front-loaded. It is concise, but the brevity comes at the cost of under-specification rather than disciplined economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A parallel fan-out tool with no annotations, no output schema, and only a one-line description leaves critical questions unanswered: what gets returned, whether the split is automatic or manual, and how workers interact with the chunks. The description is too thin for the tool's apparent complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so task, chunks, and workers are already documented in the schema. The description restates the chunking concept but adds no syntax, format, or constraint detail (e.g., chunk-to-worker mapping), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb-and-resource action (split a task and process chunks in parallel), which is clearer than a tautology. However, it gives no differentiation from the sibling swarm tools (hb_swarm_consensus, hb_swarm_hierarchy, hb_swarm_stigmergy), so an agent cannot tell when this parallel primitive is the right choice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative-selection guidance is present. The agent must infer the context entirely from the name and the sibling list, which is exactly the situation the guidelines dimension is meant to penalize.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_swarm_stigmergyC

Indirect coordination via shared state (blackboard pattern)

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYesTask description.
iterationsNoNumber of coordination iterations to plan.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the underlying coordination pattern (shared state/blackboard), which is one behavioral trait, but says nothing about side effects, state mutation, permissions, return behavior, or reversibility. For a coordination tool with zero annotation coverage, this is a substantial gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise phrase with no filler, but it is under-specified for the tool's complexity. It is front-loaded only in the sense that everything is one fragment; there is no structure to guide the agent through purpose, usage, or behavior. Concise but not appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and a coordination task that likely returns a plan or state, the description should explain what the tool returns and how it operates. It only provides a conceptual label. It is not completely empty, but it is far from complete enough to invoke the tool confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description adds no additional parameter meaning, syntax, or constraints. Baseline 3 applies when the schema does the heavy lifting, and the description does not enhance it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the coordination mechanism ('indirect coordination via shared state') and differentiates it from other swarm strategies, but it never states the actual operation the tool performs. It is a noun phrase, not a verb+resource, so an agent must infer whether the tool plans, runs, or configures stigmergy coordination. Vague but not tautological.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no conditions, and no alternatives named. The description implies a context (indirect coordination), but it does not help the agent choose between this tool and siblings like hb_swarm_consensus or hb_swarm_parallel. Matches the 'no guidance' level.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_test_listB

List available test batteries

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing beyond the name. It does not say whether results are paginated, scoped, filtered, or whether the list is static or environment-dependent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded phrase with zero filler. Every word earns its place and the purpose is conveyed immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only list tool with no output schema and no annotations, the description is barely adequate. It leaves unanswered what a 'test battery' is in this system and how its output connects to hb_test_run, which is the key thing an agent needs for multi-step test workflows.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to convey; the baseline for a parameterless tool is 4. The description correctly implies a no-argument enumeration call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List ... test batteries'), so the agent knows this enumerates test batteries. It does not distinguish itself from siblings hb_test_run and hb_test_results, which also concern tests, so the agent must infer that this is the discovery/enumeration step rather than execution or result retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance: nothing says to call this before hb_test_run to discover battery identifiers, and no alternatives or exclusions are named. The agent must guess the ordering relative to the closely related test siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_test_resultsC

Get results of last test run

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoOutput format.
batteryNoTest battery name.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not say whether results are read-only, how long results are retained, what happens if no run has occurred, or whether a run is required beforehand — significant gaps for a results-retrieval tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no waste. It is efficient, though the brevity edges toward under-specification rather than true conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description must explain the return shape and the 'last run' scoping semantics, but it does neither. An agent cannot tell what fields come back or what determines the 'last' run.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (format enum, battery) are already documented in the schema. The description adds no syntax, defaults, or format meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ('Get') and resource ('results of last test run'), so an agent knows the operation. However, it does not differentiate itself from siblings hb_test_list or hb_test_run, and 'last test run' is ambiguous given an optional battery parameter.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus hb_test_run (which produces results) or hb_test_list. The implicit ordering dependency — that a run must exist first — is left for the agent to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_test_runC

Run a test battery or single test

ParametersJSON Schema
NameRequiredDescriptionDefault
testNoOptional single test name.
batteryYesTest battery name.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it says nothing about side effects, whether execution is blocking, timeouts, permission needs, or how failures surface. For an execution tool this is a substantial gap; 'Run' alone does not tell an agent what will happen.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It is efficient, though the brevity contributes to the missing behavioral and usage context rather than compensating for it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations mean the description must explain what execution produces and where results are retrieved (likely hb_test_results), but it does neither. For a test-execution tool with side effects and no structured coverage, the description is materially incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both 'battery' and 'test' are already documented in the schema. The description only hints that a single test may be run instead of a battery, which is already implied by the optional 'test' parameter, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Run') and resource ('test battery or single test'), which distinguishes it from the read-only siblings hb_test_list and hb_test_results. It does not explicitly name those siblings or explain the boundary, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus hb_test_list or hb_test_results, no prerequisites (must the battery already exist?), and no indication of when a single test is preferable to a battery. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_ticket_listA

List tickets in one lifecycle category (INBOX/ACTIONABLE/QUEUED/BLOCKED/WAITING/USER/PARKED/SOLVED) -- header fields only, read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of results to return.
categoryYesCategory filter or category to store.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It usefully discloses that only header fields are returned and that the operation is read-only, but says nothing about pagination behavior (limit defaults to 50, what happens beyond it), ordering, or whether empty categories error out.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the read-only and header-only caveats front-loaded after the core action. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description does well to state that only header fields come back, but an agent still lacks ordering, pagination, and error behavior for the required category parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already documented in the schema, giving a baseline of 3. The description adds the notion of 'lifecycle category' but does not clarify limit semantics or what 'header fields' means for the returned objects.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (tickets) scoped to one lifecycle category, and enumerates the valid categories inline. It implicitly distinguishes itself from hb_ticket_show by declaring 'header fields only', though it never names that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the category enumeration and the 'header fields only' qualifier, but the description never states when to prefer this over hb_ticket_show or what categories mean operationally. No explicit when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hb_ticket_showB

Show one ticket's header fields by ID (e.g. T-20260825-196589547), read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
ticket_idYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It declares read-only, which is a useful behavioral trait, but says nothing about permissions, error behavior for missing tickets, or what 'header fields' omits (e.g. comments, attachments). For a retrieval tool with zero annotations this is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the verb and resource, no waste. The example ID and read-only hint are packed efficiently rather than padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with no annotations and no output schema, the description should explain what 'header fields' includes, what a not-found response looks like, and how it differs from hb_ticket_list. It leaves these gaps unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It does provide a concrete ID format example (T-20260825-196589547), which adds meaning beyond the bare 'ticket_id' string type. Still, only one param is documented this way and no format constraints beyond the example are given.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Show) and resource (one ticket's header fields) with an ID example. It's clear and distinguishable from sibling hb_ticket_list, though it doesn't explicitly name that sibling to differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'by ID' and the example format, and the read-only note signals a safe inspection call. But there's no explicit when-to-use vs hb_ticket_list, no prerequisites or error conditions for invalid IDs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 51 tool updatesv0.1.0-alpha.29
    • First observedhb_api_discover
    • First observedhb_api_export
    • First observedhb_api_history
    • First observedhb_api_probe
    • First observedhb_auto_list_chains
    • First observedhb_auto_result
    • First observedhb_auto_run
    • First observedhb_auto_status
    • First observedhb_conn_list
    • First observedhb_conn_receive
    • First observedhb_conn_send
    • First observedhb_conn_status
    • First observedhb_garden_find
    • First observedhb_garden_get
    • First observedhb_garden_put
    • First observedhb_garden_run
    • First observedhb_kb_get
    • First observedhb_kb_ingest
    • First observedhb_kb_list
    • First observedhb_kb_search
    • First observedhb_lock_check
    • First observedhb_lock_list
    • First observedhb_mem_consolidate
    • First observedhb_mem_context
    • First observedhb_mem_merge
    • First observedhb_mem_query
    • First observedhb_mem_store
    • First observedhb_plug_discover
    • First observedhb_plug_info
    • First observedhb_plug_list
    • First observedhb_plug_run
    • First observedhb_policy_list
    • First observedhb_policy_resolve
    • First observedhb_route_evaluate
    • First observedhb_route_select
    • First observedhb_route_stats
    • First observedhb_state_dispatch
    • First observedhb_state_mem_get
    • First observedhb_state_mem_set
    • First observedhb_state_task_create
    • First observedhb_state_task_list
    • First observedhb_state_task_update
    • First observedhb_swarm_consensus
    • First observedhb_swarm_hierarchy
    • First observedhb_swarm_parallel
    • First observedhb_swarm_stigmergy
    • First observedhb_test_list
    • First observedhb_test_results
    • First observedhb_test_run
    • First observedhb_ticket_list
    • First observedhb_ticket_show

TDQS

C2.7/5.0

Scored across 51 tools

Disambiguation4/5

Most tools are cleanly separated by domain prefix and verb. Minor overlaps exist: hb_api_probe vs hb_api_discover (probe-all-strategies vs schema auto-detect), and hb_mem_* vs hb_state_mem_* memory families could be confused by an agent. The many *_run/*_list/*_get tools are differentiated by their domain prefix, so ambiguity is limited.

Naming Consistency4/5

Strong, predictable hb_<domain>_<verb> pattern across nearly all tools (hb_kb_search, hb_mem_store, hb_lock_check). A few deviations: hb_auto_list_chains inverts the order used by hb_auto_run/status/result, and accessors are inconsistently named hb_garden_get / hb_kb_get vs hb_ticket_show.

Tool Count2/5

51 tools is well above the heavy threshold and far beyond what an agent can reliably select from in one session. The surface bundles at least a dozen separate domains (plugins, memory, routing, KB, swarm, state, garden, API, tests, connectors, automation, policy, tickets, locks), which should be split into focused servers.

Completeness3/5

Coverage is broad per domain but shallow: plugins have no install/enable, KB and garden lack update/delete, tasks have create/update but no delete, and policy/ticket tools are read-only by design. Read-only design is defensible, but the missing mutating operations leave lifecycle dead ends in several domains.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    D
    maintenance
    A durable multi-agent orchestrator for software development with explicit run graphs, checkpoint/resume capabilities, and project memory exposed through MCP resources and tools. It enables coordinated agent workflows for coding, review, repair, CI, and approval with SQLite-backed memory retrieval and pluggable research backends.
    10
    -
  • A
    license
    Not graded
    quality
    F
    maintenance
    Local-first, auditable memory for AI agents. Provides durable context for MCP hosts with SQLite storage, CLI, and MCP tools for memory management.
    2
    Apache 2.0
  • A
    license
    A
    quality
    F
    maintenance
    Local-first MCP memory server that gives AI coding agents long-term memory via SQLite and sqlite-vec, with optional LLM-powered layering. No gateway or API key required.
    7
    131 npm
    4
    MIT