Skip to main content
Glama

Run a crew of Claude Code and Codex workers from one lead session: visible tmux panes, exact state from hooks, one local SQLite store.

  • Visible workers: each one is a real tmux pane you can read and type into.

  • Exact state: workers report it through their own CLI's hooks, so nothing polls.

  • One local SQLite store: no daemon, and nothing leaves your machine.

hive: a lead plans, spawns workers, and is woken when they finish

Install

Requirements: macOS, Node ^22.14.0 || >=23.6.0, Claude Code, and tmux for the agent tools. codex is optional, only needed if a project opts a worker into it; see docs/install.md.

npm install -g @cmgmyr/hive
hive setup           # pins the hive command to one interpreter, prints the MCP line
brew install tmux
claude mcp add --scope user hive -- "$(command -v node)" "$(npm root -g)/@cmgmyr/hive/dist/index.js"
ln -s "$(npm root -g)/@cmgmyr/hive/claude-plugin" ~/.claude/skills/hive   # optional: session-start kickoff
hive doctor          # verify: node, ABI, tmux, claude, database, hooks all green

Working from a clone instead? See docs/install.md.

Put ~/.local/bin on your PATH below your version manager's block. See Node version and the interpreter pin for why the order matters.

Related MCP server: claude-mux.mcp

First run

Run cd ~/Code/your-project && hive. It is shorthand for hive lead, and it opens a lead window running Claude in this project's tmux session, with the lead session named after the project so your other Claude Code sessions can address it by that name; ask it to triage, and it reads the standing process and proposes work. Spawn workers with agent_spawn, and watch or take over any of them with tmux -CC attach -t hive-main (or plain tmux attach).

How it works

Each Claude Code session runs its own hive MCP server over stdio, and every instance reads and writes one SQLite database (WAL mode) at ~/.hive/hive.db, so every session sees the same state. There is no daemon and nothing leaves your machine. State is scoped to a project (a directory), resolved from the working directory; a lead spawns workers into tmux panes locked to that project.

Why not subagents?

Subagents

hive workers

Visibility

report at the end

live terminal you read and type into

Persistence

vanish with the conversation

pads and todos outlive every session

Lifetime

die with the parent

keep running when the lead detaches

Scope

one session

several sessions, terminals, humans

Hive workers can still use subagents. See Why not subagents? for the longer answer.

Running several projects

If you keep more than one project registered, hive queen starts a single lead that reads all of them. hive portfolio sorts them into waiting on you, stuck, moving and quiet, and hive next attaches you to the lead that needs you most. The queen writes into another project only through that project's lead, and hive records each write. The queen guide covers setup and limits.

Status

hive is a personal daily-driver tool, released low-key. It is single-user by design and dogfooded daily by its author on macOS. The test suite also runs on Linux in CI, but nobody drives hive there yet. Issues are welcome; for bigger changes, open a discussion first. See CONTRIBUTING.md. MIT licensed.

Docs

Page

What's there

Architecture

Seven diagrams: process topology, module layering, spawn sequence, wake lifecycle, worker state, project scoping, store and server identity

Patterns

Standing trades, refused approaches, evidence standards, and guard shapes distilled from the project's own decisions and dead-ends

Concepts

Vocabulary, why not subagents, identity, the workflow, project scope, the shared store

Daily driver

A day with hive, starting a session, watching workers, wake-ups

The queen

One lead across every project: hive queen, the portfolio, hive next, its reach and audit trail

Commands

Every hive subcommand and what it does

Configuration

The HIVE_* environment variables

Profiles

Standing instructions across projects, the session-start plugin

Projects

hive init, hive.yml, automatic backups, pads and todos from the shell

Dashboard

The generated dashboard: enabling it, where it lives, what it shows

Install details

The interpreter pin, iTerm settings, the status line, MCP scope, codex workers, updating, uninstalling

Troubleshooting

Common errors and their fixes

Tools

The MCP tools: what each does and when to use it

tmux settings

Attach modes, pane options, and what to put in ~/.tmux.conf

Development

Building and testing hive itself

Updating

Run hive upgrade for a global npm install, or hive upgrade --check to preview the commands without changing your install. In a git checkout, the default prints the recipe; hive upgrade --run executes it. Upgrade reports Claude Code and Codex registration repairs without editing their configs. Restart every session with hive loaded after the upgrade. See Updating and recovery.

Available Tools

46 tools
actor_pruneA
DestructiveIdempotent

Delete every actor that owns no rows anywhere in the store and has not been active in the last minute: agents.actor_id, agents.parent_actor_id, todos.locked_by, todo_comments.author, kv.updated_by, leases.owner, pads.updated_by, wakes.owner, wakes.deliver_actor, agent_state_log.actor_id. The scan is global across every project, never scoped to the caller's: actors carry no project_id, so an actor can own rows in a project the caller cannot see, and a project-scoped scan would misread that actor as inert and delete it. Never prunes the caller's own actor. Refuses under HIVE_PROJECT_LOCK=1: this is a whole-store sweep. Run this after project_prune when sweeping the store: an empty project owns no agents rows either, so today the order cannot orphan an actor, but that stops being true the day project deletion ever covers a non-empty project, and this ordering is the one that stays safe if it does. Immediate, permanent: no dry-run mode.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorsNo
deletedYes
held_backYes
hive_noticeNo
held_back_liveYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, but the description adds material context beyond them: the scan is global across every project and never scoped to the caller, the caller's own actor is never pruned, it refuses under a project lock, and it is immediate and permanent with no dry-run. These are exactly the facts an agent needs before triggering an irreversible store-wide sweep.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the deletion criteria, then the safety constraints, then ordering guidance — a sensible structure. It is dense and slightly discursive in the project_prune ordering rationale ('this is the one that stays safe if it does'), which could be trimmed, but nearly every sentence carries operational weight for a destructive tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is complex and destructive, yet the description covers the selection predicate, scope, self-protection, lock behavior, ordering, and irreversibility. An output schema exists, so return values need not be explained. Nothing an agent needs in order to invoke it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and schema coverage is 100%, so there is nothing for the description to disambiguate. Baseline 4 applies for a parameterless tool; no param semantics are needed or missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Delete) plus resource (actors) and the exact predicate that selects them: owns no rows anywhere in the store and inactive in the last minute. It enumerates the actual ownership columns across 11 tables, so an agent understands precisely what 'owns nothing' means. Clearly separable from project_prune and other sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to run it (after project_prune when sweeping the store) and when it will not run (refuses under HIVE_PROJECT_LOCK=1). It also names the related tool (project_prune) and the ordering rationale, giving the agent a concrete operational condition rather than inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_closeA
DestructiveIdempotent

Kill an agent's tmux window and mark it closed, addressed by name (or agent_id). Capture handoffs (todo comments, pads) BEFORE closing; terminal output is not retained. Closing yourself requires confirm_self=true. Refuses a lead target whose pane is live; retires one whose pane is confirmed dead. A worker may never close a lead, live or dead.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoThe worker's name, e.g. "impl" or "DEVX-123". Prefer passing this rather than agent_id, but pass only one: if both arrive, agent_id wins and this is ignored. A partial name works when it matches exactly one running worker, so "123" finds DEVX-123.
agent_idNoNumeric agent id. Use name instead unless you have the id to hand.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
confirm_selfNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYes
noteNo
closedYes
parkedNo
agent_idYes
hive_noticeNo
park_releasedNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and idempotentHint=true, and the description adds real context beyond them: terminal output is not retained, handoffs must be captured first, and the lead/worker authorization rules. This is unusually rich behavioral disclosure for a destructive tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the primary action, then conditional rules in descending importance, each clause carrying a real constraint. No filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values needn't be explained, and annotations cover the safety profile. The description closes the remaining gaps: data loss (terminal output), prerequisites (handoff capture), and authorization boundaries.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, and the description reinforces the key semantics: name-vs-agent_id precedence, the confirm_self gate, and the partial-name matching rule. It doesn't add new parameter detail for project_id beyond the schema, but it meaningfully supplements the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource pair ('Kill an agent's tmux window and mark it closed') and immediately names the addressing scheme (name or agent_id). An agent can distinguish this from siblings like agent_park or agent_prune without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use rules: capture handoffs BEFORE closing, confirm_self=true needed to close yourself, refuse a live lead / retire a dead one, workers never close leads. It names the alternatives implicitly (park vs close) and states the conditions that gate the call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_listA
Idempotent

List this project's agents with live status. Without include_closed, this is every running agent, in full. With include_closed, it is every agent (running and closed/parked), newest first, bounded by limit (default 20, max 100); when the receipt carries next_before_id, page through the rest by passing it back as before_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax rows to return when include_closed is true, newest first. Default 20, max 100. Ignored otherwise.
before_idNoPage cursor for include_closed: only rows with id below this value. Pass the previous receipt's next_before_id to get the next page.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
include_closedNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare idempotentHint=true and destructiveHint=false, so safety is partly covered; the description adds real behavioral detail the annotations don't — result ordering (newest first), the limit bound (default 20, max 100) and the cursor-based paging via next_before_id/before_id. It does not explain why readOnlyHint=false on what reads as a list operation, which is the one unexplained behavioral gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then layers the include_closed variants and paging in one dense but well-ordered passage. Every clause earns its place; it is slightly compressed but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by describing the receipts and the next_before_id cursor, so an agent knows how to page and what ordering to expect. Combined with tested conditional behavior for include_closed and limit, it is nearly sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the description reinforces the include_closed conditional, the limit default/max, and — beyond the schema — the pagination contract of feeding the receipt's next_before_id back as before_id. The only parameter not touched in prose is project_id, which the schema documents fully.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("List this project's agents") plus scope ("with live status" and "this project's"), which cleanly separates it from single-agent tools like agent_status. It never names a sibling explicitly, so differentiation must be inferred from the scope phrasing rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use guidance for include_closed ("Without include_closed, this is every running agent... With include_closed, it is every agent") and an explicit when-not for project_id ("Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope"). It stops short of routing the agent to alternatives such as agent_status or wake_list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_message_getA
Idempotent

Read one agent-to-lead message in full, by the id in a "[hive:worker NAME] [message #N ...]" pointer line. hive stores a message here only when it shortens one: text over 300 characters sent to a lead by someone who is not that lead. Every other send is typed with its sender tag and stores nothing, so there is no id to read. Messages are pruned after 7 days, and a lookup for a pruned id says so rather than reporting it missing.

ParametersJSON Schema
NameRequiredDescriptionDefault
message_idYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses behavior no structured field carries: the storage trigger (>300 chars, non-lead sender), the 7-day pruning window, and that a pruned id yields a distinct "pruned" response rather than a missing-record error. Note readOnlyHint=false while the description says "Read" – the annotation is more conservative, not contradicted, but the mismatch is worth knowing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action, then the id-source, then the existence condition and retention caveat. Dense and largely waste-free, though the retention/clarification sentence could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and minimal annotation coverage, the description carries the burden well: it covers invocation, precondition, retention, and the pruned-id response. Only the behavior for a never-existed or malformed id is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%; project_id is documented in the schema, and the description compensates for the undocumented message_id by explaining it comes from the "[hive:worker NAME] [message #N ...]" pointer line. That is meaningfully more than the bare integer type the schema gives.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Read one agent-to-lead message in full") and pins the identifier source to the exact pointer-line format. An agent can distinguish this from agent_send/agent_output without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the precondition precisely: a message is stored only when a >300-char message is sent to a lead by a non-lead, and otherwise "there is no id to read." That effectively tells the agent when this tool is and isn't applicable, though it does not name a sibling alternative for the other cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_outputA
Idempotent

Read the rendered terminal of an agent (default 50 lines, max 200), addressed by name or agent_id. Read REAL output before declaring a worker done.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoThe worker's name, e.g. "impl" or "DEVX-123". Prefer passing this rather than agent_id, but pass only one: if both arrive, agent_id wins and this is ignored. A partial name works when it matches exactly one running worker, so "123" finds DEVX-123.
linesNo
agent_idNoNumeric agent id. Use name instead unless you have the id to hand.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and destructiveHint=false, consistent with the 'read' framing. The description adds the line-window defaults and the 'verify real output before declaring done' workflow cue, but says nothing about what the rendered terminal contains or how it behaves when the agent has exited. readOnlyHint=false is conservative and sits in mild tension with 'Read', though destructiveHint=false and idempotentHint=true keep them reconcilable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, with the operation and its size bounds front-loaded before the usage cue. Every clause carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with no output schema and well-annotated parameters, the description covers operation, addressing, and window size adequately. It would be complete at 5 if it hinted at the returned content (text lines) or the empty/partial-name edge case beyond what the schema says.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so the schema already explains name/agent_id precedence and the project_id scoping rule. The description contributes the one thing the schema omits: the default line count of 50 against the schema's max of 200, plus the name-or-id addressing summary.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read the rendered terminal of an agent') with concrete bounds (default 50 lines, max 200), which is far more specific than the bare name. It implicitly separates itself from agent_list/agent_status by being about output content, but it never names a sibling to differentiate against, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Read REAL output before declaring a worker done' gives a concrete trigger for use rather than relying on agent status alone. There is no explicit when-not-to-use or named alternative tool, so it is clear context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_parkA
Destructive

Park a claude or codex worker for the night: kill its pane, mark the row PARKED rather than plain closed, record the branch, and hand back a board line plus the one call that brings it back. Use this instead of agent_close when the lane is paused, not finished - closed alone cannot tell a next-morning lead which is which. Resume it with agent_resume.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoThe worker's name, e.g. "impl" or "DEVX-123". Prefer passing this rather than agent_id, but pass only one: if both arrive, agent_id wins and this is ignored. A partial name works when it matches exactly one running worker, so "123" finds DEVX-123.
agent_idNoNumeric agent id. Use name instead unless you have the id to hand.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
confirm_selfNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
cwdYes
nameYes
noteNo
parkedYes
agent_idYes
todo_idsYes
parked_atYes
board_lineYes
session_idYes
hive_noticeNo
parked_branchYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and readOnlyHint=false, and the description adds real behavioral detail beyond them: it kills the pane (concrete destructive act), records a distinct PARKED state rather than closed, and captures the branch. It does not discuss permissions or idempotency/reversibility beyond the resume pointer, so a 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the verb and its effects, then the routing guidance, then the resume pointer. Dense but every clause carries information; minor redundancy in restating PARKED vs closed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the destructive effect, the state distinction, and the resume path. For a destructive mutation tool this is nearly complete, with only permissions/irreversibility left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the schema descriptions are rich (name partial-match rules, agent_id precedence, project_id scoping). The description adds no parameter-level syntax or meaning beyond what the schema already provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (park) and resource (claude/codex worker) and enumerates the concrete effects: kill the pane, mark the row PARKED, record the branch, return a board line plus the resume call. It explicitly distinguishes itself from the sibling agent_close.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use rule ('when the lane is paused, not finished') and names the alternative it replaces (agent_close), plus the follow-up path (agent_resume). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_renameA
Destructive

Change a worker's display name. Its actor_id (agent:N) does not change, so every pad write, todo comment and lease it has already made stays attributable. A live claude worker is also told to retitle its own session, which shows up in its pane; that arrives as a user turn, so rename between assignments rather than mid-task. Refuses a lead target outright.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoThe worker's name, e.g. "impl" or "DEVX-123". Prefer passing this rather than agent_id, but pass only one: if both arrive, agent_id wins and this is ignored. A partial name works when it matches exactly one running worker, so "123" finds DEVX-123.
agent_idNoNumeric agent id. Use name instead unless you have the id to hand.
new_nameYesThe new display name. No other running worker may have it, case aside.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYes
noteNo
tailNo
actor_idYes
agent_idYes
retitledYes
hive_noticeNo
previous_nameYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations: attribution of existing pad writes/todos/leases is preserved, a live claude worker receives an out-of-band retitle delivered as a user turn visible in its pane, and lead targets are rejected. These are consequences an agent could not infer from destructiveHint/readOnlyHint alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action, then packs the identity invariant, worker-side effect, timing advice and refusal rule into three dense sentences with no filler. Slightly long, but every clause carries operational information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a full input schema and an output schema present, return values need no explanation; the description instead covers the risky edges (lead refusal, mid-task timing, session retitle side effect). Nothing essential to calling it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents name-vs-agent_id precedence, partial-name matching, uniqueness of new_name, and the project_id scope override. The description adds the actor_id format note but no further parameter detail, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource combination (change a worker's display name) and immediately clarifies the identity invariant (actor_id unchanged). An agent can distinguish this from siblings like agent_spawn or agent_park without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete timing guidance ("rename between assignments rather than mid-task") and an explicit refusal condition ("Refuses a lead target outright"). It does not name an alternative sibling for renaming-adjacent needs, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_resumeA

Resume a CLOSED claude or codex worker from its recorded session id (claude --resume / codex resume): a fresh pane, the same actor_id, and the worker's full prior context. Addressed by name or agent_id among closed agents (agent_list(include_closed: true)). Send it its next instruction with agent_send once resumed - this tool does not.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoThe worker's name, e.g. "impl" or "DEVX-123". Prefer passing this rather than agent_id, but pass only one: if both arrive, agent_id wins and this is ignored. A partial name works when it matches exactly one running worker, so "123" finds DEVX-123.
agent_idNoNumeric agent id. Use name instead unless you have the id to hand.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYes
actor_idYes
agent_idYes
hive_noticeNo
tmux_targetYes
branch_driftNo
was_parked_atNo
landed_in_projectNo
resumed_session_idYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already declaring readOnlyHint=false, idempotentHint=false, and destructiveHint=false, the description adds useful behavior beyond the structured data: a fresh pane, the same actor_id, full prior context, and the fact that it does not send a follow-up instruction. It stops short of describing error or permission behavior, so it is not fully exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and is only three compact sentences. Each sentence carries necessary information: what is resumed, how it is addressed, and what to do after resuming.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich input schema, output schema, and annotations, the description supplies the remaining context an agent needs: the worker must be CLOSED, targeting uses name or agent_id among closed agents, and sending the next instruction is a separate tool call. Nothing essential for correct invocation appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the parameters thoroughly. The description still adds meaning by specifying that the target must be addressed by name or agent_id among closed agents, and by pointing to agent_list(include_closed: true) for finding them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Resume a CLOSED claude or codex worker from its recorded session id'. It also distinguishes the result from creating a new worker by noting 'a fresh pane, the same actor_id, and the worker's full prior context', so an agent can tell it apart from siblings like agent_spawn.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit conditions: resume a CLOSED worker, addressed by name or agent_id among closed agents, findable via agent_list(include_closed: true). It also names the alternative follow-up action: 'Send it its next instruction with agent_send once resumed - this tool does not.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_sendA
Destructive

Type into an agent's terminal, addressed by name (or agent_id). text of any shape is prefixed with the sender tag, delivered as one bracketed paste and submitted with Enter unless submit=false. ONE EXCEPTION: text over 300 characters sent to a LEAD by anyone who is not that lead is stored and delivered as a one-line pointer instead, because a lead's pane is a human's own window; the receipt says so and names agent_message_get for the full text. Worker-bound text is never shortened at any length. Alternatively pass keys (tmux key names like Escape, C-c, Enter). wait_ms (250-10000) returns the terminal tail after sending. A claude worker is already briefed by agent_spawn. A worker whose screen hive cannot classify is REFUSED on the text path entirely (its brief is at the spawn receipt's brief_path); keys still reaches it. A pane in tmux copy mode is REFUSED too, and retriably: tmux clears its bracketed-paste flag there, so the paste would lose its markers and the Enter would be eaten - leave copy mode (or agent_send(keys: ["-X", "cancel"]) to cancel it deliberately) and send again.

ParametersJSON Schema
NameRequiredDescriptionDefault
keysNotmux key names, e.g. ["Escape"] or ["C-c"].
nameNoThe worker's name, e.g. "impl" or "DEVX-123". Prefer passing this rather than agent_id, but pass only one: if both arrive, agent_id wins and this is ignored. A partial name works when it matches exactly one running worker, so "123" finds DEVX-123.
textNo
submitNoAppend Enter after text. Defaults to true.
wait_msNo
agent_idNoNumeric agent id. Use name instead unless you have the id to hand.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare it is destructive and non-idempotent; the description adds substantial context beyond them — sender-tag prefixing, single bracketed paste plus Enter, the 300-character lead-truncation rule, two distinct refusal conditions, and the copy-mode bracketed-paste-flag explanation. This is exactly the behavioral disclosure annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and defers edge cases, and every sentence carries information. However it is delivered as one dense block with the refusal/exception rules run together, which costs some scanability on re-read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still explains what is returned (terminal tail via wait_ms, the receipt naming agent_message_get), plus the refusal and truncation behavior. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 71%, and the description compensates well: submit defaults to true, keys take tmux key names, wait_ms (250-10000) returns the terminal tail, and text is prefixed and pasted as one bracket. It adds real meaning over the schema for text, which has no inline description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource: 'Type into an agent's terminal, addressed by name (or agent_id).' It immediately contrasts with the read path by naming agent_message_get for full text, so the agent can distinguish it from siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly covers when to use the text path vs the keys path, when it is refused (unclassifiable hive screen, copy mode), the alternative for full text (agent_message_get), and even how to recover from copy mode. Conditions and alternatives are stated rather than implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_spawnA

Spawn a worker agent (default: claude, or the project's hive.yml agents: default). A claude worker is briefed automatically: the full brief is appended to its system prompt, so send it its assignment directly. A command or harness that resolves to a known harness (claude, codex) not listed in the project's hive.yml agents: is refused; absent agents: means claude only. A command hive cannot classify the screen of (claude and codex both do; a harness with no entry does not) can be spawned but NOT typed into: the receipt carries brief_path and says so, and agent_send's text path and wakes both refuse that pane. Set read_only: true to block local file writes and mutating shell while keeping Hive MCP available; agent_resume preserves this mode and agent_status reports it. Only bare claude/codex executables are supported in this mode, with only effort-setting extra_args. The worker is locked to this project. Humans can watch with: tmux attach -t hive-main.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoWorking directory, e.g. a git worktree path. Defaults to the project root.
nameNoDisplay name; defaults to worker-N. This is how you address the worker later.
modelNoPassed as --model to the agent command.
layoutNoHow to arrange the lead's window when placement is split. main-vertical gives the lead the left half with workers stacked on the right; tiled (default) splits evenly. Projects can set a default in hive.yml.
commandNoRaw agent command to run. Overrides harness when both are given. Defaults to the project's hive.yml agents: default, or claude. Refused if it resolves to a known harness the project's agents: list does not allow.
harnessNoSpawn a known harness by name (e.g. "codex") instead of a raw command. Ignored when command is also given. Must be in the project's hive.yml agents: list (default: claude only).
placementNosplit (default): the worker appears as a pane in the lead's window, auto-tiled, so the whole crew shares one screen. window: its own tmux window (an iTerm tab under control mode).
read_onlyNoPrevent local file writes and mutating shell commands while allowing Hive MCP tools. Only bare claude/codex executables. Extra args allow codex -c model_reasoning_effort=low|medium|high|xhigh, or claude --effort low|medium|high|xhigh|max; all other arguments are refused.
extra_argsNoExtra CLI arguments.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYes
noteNo
tailNo
readyNo
exitedNo
layoutNo
actor_idYes
agent_idYes
read_onlyNo
brief_pathNo
codex_homeNo
hive_noticeNo
tmux_targetYes
instructionsNo
config_warningsNo
worktree_installNo
landed_in_projectNo
codex_instructionsNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (which only say non-read-only, non-destructive), the description discloses substantial behavior: harness refusal rules, that unclassifiable panes can be spawned but not typed into, that the receipt carries brief_path, that read_only blocks writes while keeping MCP, that agent_resume preserves the mode, and that the worker is project-locked. This is exactly the extra context the annotations do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first clause, and the remaining sentences each cover a distinct rule (briefing, refusal, typing restriction, read_only, project lock, monitoring). It is dense and long, but no sentence is pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter spawn tool with annotations and an output schema (the receipt), the description covers the critical edge cases an agent must know: default resolution, refusals, read_only constraints, project scoping, and a human-monitoring command. Nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: the semantics of read_only mode (only bare claude/codex, only effort extra_args), the interaction between command and harness, and the project_id scoping constraint. This goes beyond restating field docs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb (spawn) and resource (worker agent) and immediately names the default harness resolution rules. An agent can distinguish this from agent_resume, agent_send, and agent_list without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operational context: the claude worker is auto-briefed so send the assignment directly, refusal conditions for unlisted harnesses, and the fallback when agents: is absent. It names related tools (agent_send, agent_resume, agent_status) but does not fully frame them as alternatives to choose between, so it stops short of a full when/when-not rubric.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_statusA
Idempotent

Detailed status for one agent, addressed by name (or agent_id), including a short tail of its terminal. include_brief=true returns the exact brief this worker was given; hive keeps that copy because an appended system prompt appears in no transcript.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoThe worker's name, e.g. "impl" or "DEVX-123". Prefer passing this rather than agent_id, but pass only one: if both arrive, agent_id wins and this is ignored. A partial name works when it matches exactly one running worker, so "123" finds DEVX-123.
agent_idNoNumeric agent id. Use name instead unless you have the id to hand.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
include_briefNoReturn the full injected brief, not just its path. Defaults to false.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true, destructiveHint=false, openWorldHint=false. The description adds real value by disclosing what the response contains (terminal tail, optional brief) and the non-obvious rationale that hive stores the brief because an appended system prompt appears in no transcript. However, readOnlyHint=false sits in mild tension with a description that portrays a pure inspection call, and pagination/return shape are not covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with purpose and the addressing mechanism, zero filler. The second sentence drifts slightly into hive implementation trivia that an agent doesn't strictly need in order to call the tool, which keeps it out of the top band.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the return-shape burden; it does so adequately by naming the terminal tail and the optional full brief. Safety semantics are covered by annotations and the parameter surface is fully covered by the schema, so the remaining gaps (full return structure, empty-result behavior) are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so name/agent_id/project_id/include_brief are already fully documented in the schema. The description reinforces the name-or-agent_id selector and the meaning of include_brief=true, but adds no syntax, format, or priority detail beyond what the schema states (the 'agent_id wins' precedence lives in the schema, not here). Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource pairing ('Detailed status for one agent') plus the selector (name/agent_id) and the distinctive return ('a short tail of its terminal'), which distinguishes it from list-style siblings like agent_list. It stops short of explicitly naming which sibling to use instead, so it lands just below the top band.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: 'one agent, addressed by name' tells the agent this is the single-agent counterpart to agent_list, and include_brief is explained. There is no explicit when-to-use/when-not or mention of agent_output/agent_message_get as alternatives, so an agent must infer the routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

helpA
Read-onlyIdempotent

Hive usage guidance. Omit topic for an overview, or pass one of: workflow, profiles, agents, wakes, projects, identity, actors, pads, todos, kv, leases.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicNo

TDQS

A4.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, covering the safety profile. The description adds little behavioral context beyond restating that it supplies guidance; with annotations carrying that burden, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero waste. The primary behavior (omit for overview) is front-loaded, and the valid topics follow compactly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple help tool with one optional parameter, no output schema, and annotations covering the safety profile, the description is complete. It explains both invocation modes and all valid topic values, leaving no gaps an agent would need to fill.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: it fully specifies that omitting the topic yields an overview and enumerates every valid topic string. This effectively documents the single optional parameter despite the schema having no enum or description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool provides Hive usage guidance and enumerates the exact topics available, making its purpose and scope unmistakable. It is clearly distinct from all sibling tools, which are operational actions on agents, projects, pads, todos, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit usage instructions: omit the topic for an overview, or pass one of the listed topics for specific guidance. This directly tells the agent when to use each mode without requiring inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kv_deleteB
DestructiveIdempotent

Delete a shared value by key.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
keyYes
deletedYes
project_idYes
hive_noticeNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered by structured data. The description adds only the word "shared," hinting the value is cross-agent/cross-project scoped, but says nothing about behavior on a missing key or whether deletion is recoverable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded with the action. Nothing is wasted, though it is arguably too terse to be fully informative for a destructive operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema and complete destructive/idempotent annotations reduce the burden, and for a two-parameter delete the description is barely sufficient. It still omits error behavior for a nonexistent key and the project-scoping default, leaving real gaps for a destructive tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: project_id carries rich inline guidance while key has none. The description's "by key" merely restates the required parameter name and adds no format, uniqueness, or lookup semantics, and it ignores project_id entirely — so it does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Delete a shared value") and the selector ("by key"), which cleanly distinguishes it from siblings kv_get, kv_set, and kv_list. It stops short of explicitly naming those alternatives, so it is clear but not maximally differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or prerequisite guidance is given. It never mentions kv_set/kv_get as the paired operations or warns that deletion is permanent. The agent must infer everything from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kv_getC
DestructiveIdempotent

Get a shared JSON value by key.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

TDQS

C2.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=true, yet the description frames this as a plain read ('Get a shared JSON value'). Nothing in the text explains the destructive/non-read-only behavior, so the description actively misleads relative to the structured metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no waste. It is arguably too terse for the behavioral complexity implied by the annotations, but as prose it is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 50% parameter coverage, the description needed to carry more weight, and it leaves the destructive/non-read-only nature of the operation completely unexplained. An agent could call this believing it is a safe read.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: project_id carries a strong, self-contained usage constraint, but the required 'key' parameter is undocumented. The description only adds 'by key,' which is already obvious from the schema, so it does not compensate for the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('shared JSON value') and identifies the lookup key. It does not explicitly distinguish itself from siblings like kv_list or kv_delete, but the intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no mention of alternatives such as kv_list or kv_delete, and no stated prerequisites. The agent has to infer that this is the single-key read path from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kv_listC
DestructiveIdempotent

List shared values, optionally filtered by key prefix.

ParametersJSON Schema
NameRequiredDescriptionDefault
prefixNo
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

TDQS

C2.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=true, yet the description presents this purely as a read/enumeration operation ('List shared values'). The description offers no acknowledgment of the surprising destructive/non-read-only behavior, leaving the contradiction unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with the primary action first and the modifier second. Nothing is wasted and it is appropriately sized for a two-parameter list tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with no output schema, the definition is minimally adequate, but the unexplained destructive annotation and the undocumented prefix param leave real gaps. An agent cannot fully predict the tool's side effects or filtering behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: project_id carries a strong inline description, while prefix is undocumented in the schema and only implied by the description's 'filtered by key prefix'. The description compensates marginally for the prefix gap but adds no format or matching semantics (e.g. exact vs prefix match behavior).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('shared values') with the optional prefix filter, so the agent knows it enumerates key-value entries. It does not explicitly distinguish itself from the sibling kv_get (single-value fetch), though 'List' vs 'Get' is fairly self-evident.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance is 'optionally filtered by key prefix' — there is no statement of when to use this versus kv_get or the other kv_* siblings, and no mention of scope or prerequisites. The agent gets context only by inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kv_setB
Destructive

Set a small shared JSON value other sessions can discover. Optional TTL in seconds.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
valueYesAny JSON value.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
ttl_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
keyYes
project_idYes
hive_noticeNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-read-only, destructive write (overwrites), so the safety picture is partially covered. The description usefully adds cross-session visibility and TTL semantics, but does not clarify overwrite behavior, the meaning of 'small', or any key/value limits despite destructiveHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and followed by the optional modifier. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists and annotations cover the destructive/overwrite profile, so the description need not explain returns. Still, for a cross-session shared-value write tool it omits overwrite semantics and 'small' size expectations, leaving notable gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%; project_id is well documented in the schema while key and ttl_seconds are not. The description adds the unit for ttl_seconds ('in seconds'), a real gain, but says nothing about key naming rules or project scoping beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('Set a small shared JSON value') with scope info ('other sessions can discover'), clearly distinct from kv_get/kv_delete by name and semantics. No explicit sibling differentiation, but the resource is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'other sessions can discover' hints at the cross-session sharing use case, but there is no explicit when-to-use, when-not, or reference to alternative tools. The agent must infer usage entirely from the name and sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lease_acquireA

Try to take a named lease on state another session could also change, such as a shared checkout, a dev database or a port (non-blocking). Re-taking your own lease extends it. Leases expire on their own.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesStable and specific, like "db:dev".
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
ttl_secondsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
keyYes
held_byNo
acquiredYes
extendedNo
expires_atNo
project_idYes
hive_noticeNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds genuinely useful traits beyond the annotations: non-blocking acquisition, re-acquiring your own lease extends it, and leases self-expire. The annotations only cover the safety profile, so the contention and expiry semantics are meaningful added context. What happens on a failed acquire (error vs. false return) is not stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the purpose, and the parenthetical placement of '(non-blocking)' is efficient. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and annotations carry the safety profile. The description covers the acquisition/extension/expiry semantics an agent needs; only the failure mode and ttl meaning are left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description never mentions key, ttl_seconds, or project_id. With schema coverage at only 67%, ttl_seconds in particular is undocumented in both the schema and the prose, so the description does nothing to compensate for the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('take a named lease') plus resource and the kind of state it guards, with concrete examples (shared checkout, dev database, port). It is clearly the acquisition half of the lease pair, though it never names the sibling lease_release to sharpen the contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the situation that calls for it (state another session could also change) and flags that it is non-blocking, so the agent knows it fails rather than queues. No explicit when-not-to-use or pointer to the alternative tool (lease_release, or a lock-free path).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lease_releaseC
DestructiveIdempotent

Release a lease you own.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
keyYes
releasedYes
project_idYes
hive_noticeNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true and readOnlyHint=false, so the safety profile is covered structurally. The description adds one genuinely non-structured fact — that the caller must own the lease — but says nothing about failure behavior for unowned keys, side effects on the leased resource, or error semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the action front-loaded and no wasted words. It is efficient, though the brevity borders on under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, for a destructive, ownership-gated mutation, the description omits what happens on success, whether other holders are affected, and what error to expect when the lease is not owned — leaving real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: project_id carries a thorough description, but the required 'key' parameter has no schema description and the tool description does not explain it either. The description therefore does nothing to compensate for the undocumented required parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource pair ('Release a lease'), which an agent can distinguish from the sibling lease_acquire without opening either schema. It stops short of explicitly naming the counterpart tool or the scope of the release, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'you own' implies a precondition but gives no explicit when/when-not guidance, no reference to lease_acquire or lease lifetimes, and no indication of what to do if ownership is unclear. Usage must be inferred entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pad_appendB

Append content to the end of a pad. Optional expected_revision guards against concurrent writes.

ParametersJSON Schema
NameRequiredDescriptionDefault
pad_idYes
contentYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
expected_revisionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
pad_idYes
revisionYes
hive_noticeNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-read-only, non-idempotent, non-destructive behavior, and the description usefully confirms the additive append semantics plus the concurrency angle. It stops short of saying what happens on an expected_revision mismatch (error vs. overwrite) or whether appends are atomic.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the operation and then the concurrency guard. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations and an output schema present, the description only needs to carry the mutation semantics, which it does. Adding the failure mode for a stale expected_revision would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%: only project_id is documented in the schema, and the description explains expected_revision's purpose (guarding concurrent writes), which the schema does not. pad_id and content remain undocumented, though their names are largely self-evident.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Append content to the end of a pad') plus the additive scope, which distinguishes it from siblings like pad_write and pad_edit. It does not name those siblings explicitly, so differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no alternatives named, even though pad_write, pad_edit, and pad_read sit in the same family. 'Append' implies the intent, but an agent must infer when to append versus rewrite or edit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pad_archiveA
Destructive

Archive a pad (or unarchive with archived=false). Archiving frees the name for a new active pad; the old content stays readable by pad_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
pad_idYes
archivedNoDefault true. Pass false to unarchive.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
pad_idYes
archivedYes
revisionYes
hive_noticeNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and readOnlyHint=false, so the description's job is to add context — and it does: the operation is reversible (unarchive) and content is preserved and still readable by pad_id. That materially softens the 'destructive' reading and tells the agent the effect is renameable rather than erased.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero padding, with the primary action and its reversal front-loaded before the side-effect explanation. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers the required parameter's effect plus reversibility. The only thin spot is the project_id scope override, which is left entirely to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%, with archived and project_id already documented in the schema; pad_id is bare but self-evident. The description's 'unarchive with archived=false' merely restates the schema's default note, adding little syntax or format detail beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Archive a pad') and immediately covers the inverse operation ('unarchive with archived=false'). It also implicitly distinguishes itself from the sibling pad_delete by noting the old content stays readable, so an agent can tell the two apart without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for use ('Archiving frees the name for a new active pad') and the condition that flips behavior (archived=false). It does not explicitly name when to prefer it over pad_delete or note any prerequisites, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pad_deleteA
DestructiveIdempotent

Permanently delete a pad. Irreversible; prefer pad_archive. Optional expected_revision guards against deleting a pad someone just updated.

ParametersJSON Schema
NameRequiredDescriptionDefault
pad_idYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
expected_revisionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
pad_idYes
deletedYes
hive_noticeNo

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, and the description reinforces irreversibility rather than contradicting. It adds real value beyond the annotations by disclosing the concurrency-guard purpose of expected_revision, though it doesn't say what happens on a revision mismatch (fail vs. ignore).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the destructive action and the safer alternative; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the description covers irreversibility, the archive alternative, and the revision guard. The only omission is the failure behavior when expected_revision doesn't match.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%, so the description must compensate; it explains expected_revision ('guards against deleting a pad someone just updated') which is the non-obvious parameter. pad_id is self-evident and project_id is documented in the schema, leaving only minor gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Permanently delete a pad') and immediately contrasts it with the safer sibling 'pad_archive', so an agent can distinguish it from the other pad_* tools without reading schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the preferred alternative ('prefer pad_archive') and gives the reason ('Irreversible'), which is exactly the when-to-use/when-not guidance needed for a destructive tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pad_editB
Destructive

Replace one literal occurrence of old_text with new_text in a pad. old_text must match exactly once; include surrounding context to disambiguate.

ParametersJSON Schema
NameRequiredDescriptionDefault
pad_idYes
new_textYes
old_textYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
expected_revisionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
pad_idYes
revisionYes
hive_noticeNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false, so the agent knows this is a non-idempotent mutation. The description adds useful behavioral context: only one literal occurrence is replaced, and old_text must match exactly once. It does not explain concurrency behavior, error handling on multiple matches, or the role of expected_revision, so it adds only moderate value beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with no wasted words. The core operation and the critical old_text matching constraint are front-loaded and immediately actionable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations cover the destructive safety profile. The description adequately covers the core edit behavior, but it omits expected_revision semantics entirely and does not mention pad_id or project_id behavior, leaving gaps for a mutation tool with five parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%, so the description must compensate, but it covers only old_text and new_text semantics. It does not explain pad_id or expected_revision, and expected_revision is a critical optional parameter for a destructive edit tool. Partial compensation leaves significant parameter gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: replace one literal occurrence of old_text with new_text in a pad. It is clear and unambiguous, but it does not distinguish itself from sibling tools like pad_write or pad_append, so it falls short of the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an important usage constraint — old_text must match exactly once, and surrounding context should be included to disambiguate — which implies when this tool is appropriate. However, it does not state when to use pad_edit instead of alternatives such as pad_write or pad_append, leaving alternative selection unaddressed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pad_listB
Idempotent

List pads without full content. query matches names and content (returns a snippet); tags matches any listed tag.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNo
limitNo
queryNo
offsetNo
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
include_archivedNo

TDQS

B3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description describes a pure listing operation ('List pads'), but the annotations declare readOnlyHint=false, i.e., the tool is not read-only. For a tool whose stated purpose is listing, the annotation asserting non-read-only behavior is directly at odds with the description's implication, and the description does nothing to reconcile it (e.g., explain side effects such as access tracking or logging).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: purpose first, then filter semantics. No wasted words and the front-loading is correct, though the terseness borders on under-specification rather than optimal concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 6 parameters, no output schema, and only safety-oriented annotations, the description should carry more load. It covers purpose and two filter parameters but omits pagination behavior (limit/offset defaults), archived-inclusion semantics, and the non-read-only behavior, leaving an agent with gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% (just project_id), so the description must compensate. It does add real semantics for query (matches names AND content, returns a snippet) and tags (matches any listed tag), but says nothing about limit, offset, or include_archived, leaving a substantial documentation gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List pads') with a meaningful scope qualifier ('without full content') that implicitly distinguishes it from the full-content sibling pad_read. The purpose is clear, though it never names the sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without full content' implies when to prefer this over pad_read, and the query/tags clauses hint at filtering use cases. However, there is no explicit when-to-use/when-not-to-use statement or named alternative, so guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pad_readC
Idempotent

Read a pad's content, revision, and metadata by pad_id or name.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
pad_idNo
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

TDQS

C2.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description asserts a read-only operation, while the annotations declare readOnlyHint=false, i.e. the tool may mutate state. That is a direct contradiction of the stated behavior, and the description adds no compensating context (no permissions, no scope behavior, no return detail).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; nothing redundant is present, though it is arguably too terse given the gaps it leaves unaddressed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no output schema and annotations that cannot be trusted to convey the safety profile, the description is too thin: it never clarifies what 'revision' and 'metadata' contain, how name/pad_id resolution works, or the project-scope rule that the schema only partly covers.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%: project_id is documented in the schema (including the 'explicit override only' guidance), but name and pad_id are bare. The description repeats the two lookup keys without explaining precedence if both are supplied, whether the name is scoped to the current project, or the pad_id format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read) plus the resource (a pad) and enumerates what is returned: content, revision, metadata. Among siblings (pad_write, pad_append, pad_edit, pad_archive, pad_delete, pad_list) the read semantics are unambiguous, though it does not explicitly distinguish itself from pad_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the verb 'Read' and the lookup keys ('by pad_id or name'). There is no guidance on when to use this instead of pad_list, nor any exclusions or prerequisites, leaving the agent to infer the selection condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pad_writeA
Destructive

Create a pad, or fully overwrite one by passing pad_id plus expected_revision. Pad names are unique per project. Prefer pad_append/pad_edit for targeted changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
tagsNo
pad_idNoPass with expected_revision to overwrite.
contentYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
expected_revisionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
pad_idYes
revisionYes
hive_noticeNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false; the description reinforces this by stating the operation 'fully overwrites' and that overwrite requires passing expected_revision (optimistic concurrency). It also discloses that pad names are unique per project, implying a duplicate-name failure. It does not say what happens if expected_revision is stale, which would round this out.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the create/overwrite distinction, followed by the uniqueness constraint and the alternative-tool routing. No filler; each sentence changes how the agent would call it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and annotations cover the safety profile, so return values and destructiveness need not be restated. The description covers both modes, the concurrency contract, uniqueness, and siblings. The only material gap is behavior on revision mismatch and tag merge semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (pad_id and project_id documented in-schema), so the description must compensate for name, tags, content, expected_revision. It adds meaning for pad_id+'expected_revision' and for name uniqueness, but says nothing about tags semantics (replace vs merge) or content format. Baseline 3 given partial coverage and partial compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource and distinguishes two modes: create a new pad, or 'fully overwrite one by passing pad_id plus expected_revision'. It also names the sibling pad_append/pad_edit as the lighter-weight alternatives, so an agent can separate this from neighboring pad tools without reading schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Prefer pad_append/pad_edit for targeted changes', giving the condition that selects the alternative. It lacks an explicit when-not for this tool (e.g. don't use to add a single line), but the guidance is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_addA
Idempotent

Register a directory as a project. Defaults to the current working directory. Returns the existing project if the path is already registered.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYes
nameYes
pathYes
created_atYes
hive_noticeNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and destructiveHint=false. The description reinforces this with the concrete behavior 'returns the existing project if the path is already registered,' which is useful added context. It omits permissions/auth requirements, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: what it does, the path default, and duplicate behavior. Front-loaded and waste-free.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and annotations cover the safety profile. The one gap is the undocumented 'name' parameter, but otherwise the description is complete for a simple registration tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for 2 parameters, so the description must compensate. It explains the 'path' parameter's default (cwd) but says nothing about 'name', leaving half the parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Register a directory as a project') and adds scope details (default path, duplicate handling). It does not differentiate from siblings like project_list or project_select, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The default-to-cwd behavior and the 'returns existing project if already registered' note imply usage context, but there is no explicit when-to-use-this-vs-alternatives guidance and no mention of project_select or project_list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_listA
Idempotent

List registered projects and the currently selected one.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, which is in tension with a verb ("List") that reads as a pure read. The description does nothing to resolve that: no mention of side effects, permissions, or why a listing operation is not flagged read-only. With annotations present the bar is lower, but it adds zero context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. The most important information — that this lists projects — is the first word.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read tool with no output schema, the description does state what is returned (the project list plus the current selection). It could say more about the return shape (names vs. IDs, how the selection is marked), but the core need is covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters and an empty properties object, so there is nothing for the description to disambiguate; the baseline for a no-param tool applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (registered projects) plus an extra detail — the currently selected one — so the agent knows exactly what comes back. It sits clearly alongside project_add/project_select/project_prune in the sibling family, though it never names those siblings or distinguishes itself from them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when/when-not guidance and no alternatives named. Usage is only implied: an agent can infer that this is the discovery call before project_select or project_prune, but the description never says so.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_pruneA
DestructiveIdempotent

Delete every registered project that owns no rows anywhere in the store (pads, todos, kv, leases, agents, wakes, command_trust), verified individually before each delete. Never prunes the caller's own project. With project_id, removes exactly that one project instead of sweeping, and only if it owns no rows; otherwise refuses and names what it owns. Refuses under HIVE_PROJECT_LOCK=1: this is a whole-store sweep, and a project-locked session may only touch its own project. Immediate, permanent: no dry-run mode.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorsNo
deletedYes
held_backYes
hive_noticeNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive/idempotent, but the description adds substantial value beyond them: per-project verification before each delete, refusal semantics naming what a project owns, the HIVE_PROJECT_LOCK=1 guardrail, and 'immediate, permanent: no dry-run mode.'

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, all load-bearing, with the scope stated first and the destructive/permanent caveat last. Slightly dense but nothing is wasted or out of place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value details are not needed here. For a destructive, gated, parameterizable prune, the description covers scope, preconditions, guardrails, and irreversibility – everything an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single parameter has no schema description, so the description carries the load: it explains that project_id switches from sweep to targeted deletion, that it still requires owning no rows, and that failure names the owned rows. It could arguably note the positive-integer constraint, but the behavioral meaning is well conveyed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (delete/prune) and resource (registered projects owning no rows), scoping it across named row types (pads, todos, kv, leases, agents, wakes, command_trust). It is clearly distinguishable from siblings like project_add, project_list, and project_select.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly covers when to sweep vs. when to pass project_id, the condition that makes deletion legal (owns no rows), the refusal under HIVE_PROJECT_LOCK=1, and the hard exclusion of the caller's own project. No inference is left to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_selectA
Idempotent

Set which project later tools act on in this session.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare this is a non-read-only, idempotent, non-destructive, closed-world operation. The description adds genuinely useful context beyond that: the effect is session-scoped state that mutates the behavior of subsequent tools. It does not cover what happens on an invalid or nonexistent project_id, nor whether the selection persists across sessions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that carries the effect and the scope with zero filler. Nothing to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tiny one-parameter session-state tool with no output schema and annotations covering the safety profile, the description conveys the essential behavior. Missing only edge-case behavior on invalid ids and persistence semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter with 0% schema description coverage, so the schema gives only type and bounds (positive integer). The description's 'which project' implies the parameter is a project identifier but adds no format, source, or validation meaning beyond that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Set') and the resource it affects ('which project later tools act on'), plus the scope ('in this session'). This clearly distinguishes it from project_add, project_list, and project_prune, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'later tools act on' implies this should be called before project-scoped operations, giving indirect usage context. However, it does not state when-not to use it, what happens if a project is already selected, or point to any alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

queen_audit_listA
Read-onlyIdempotent

List confirmed queen writes into other projects, newest first. The queen may read all targets; other callers see only their own project. Default 20, maximum 100 entries. History is retained for 30 days with a 20,000 id-range backstop. A crash between a completed terminal send and its audit insert can leave that send unrecorded.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
project_idNoFilter by the target project id. Only the queen may name another project.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the safety profile (readOnly, idempotent, non-destructive). The description goes well beyond: scoping of visibility by caller role, a 30-day retention window with a 20,000 id-range backstop, and the caveat that a crash between a completed send and its audit insert can leave a send unrecorded. That last point is exactly the kind of reliability disclosure annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with what the tool does, then visibility, then limits, then the caveat. Each sentence carries distinct information; the crash caveat is arguably appendix material but is still worth its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, read-only list tool with no output schema, the description covers scope, access, bounds and retention. It does not describe what fields an audit entry contains, which an agent might want, but the absence of an output schema makes that a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% – project_id is documented in the schema, but limit is not. The description compensates by stating 'Default 20, maximum 100 entries', which is materially different from the schema's nominal maximum and adds real meaning. It does not restate the queen-only rule for project_id, but the schema already covers that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List confirmed queen writes into other projects') plus ordering ('newest first'). No sibling tool covers this audit view, so the agent can immediately place it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the access model that governs when the tool returns anything: the queen sees all targets, other callers see only their own project. It does not name an alternative tool or state a when-not condition, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

todo_archiveA
Destructive

Archive a todo (or unarchive with archived=false), mirroring pad_archive. Archived todos are excluded from todo_list by default; todo_get always reaches them by id. Refuses when this todo still blocks another todo that is not completed.

ParametersJSON Schema
NameRequiredDescriptionDefault
todo_idYes
archivedNoDefault true. Pass false to unarchive.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
todo_idYes
archivedYes
project_idYes
hive_noticeNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint, idempotentHint, readOnlyHint) by disclosing the visibility side effect: archived todos drop out of todo_list by default while todo_get still reaches them by id. It also states a hard precondition that causes refusal, which isn't derivable from the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with zero filler, front-loaded with the primary action and then the effects and refusal rule. Every clause carries information the agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers the mutation's effect, visibility semantics, and failure precondition. Nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% and both non-trivial parameters (archived, project_id) already carry their own descriptions. The description only restates the archived=false unarchive behavior and adds no syntax or constraint detail for project_id or todo_id, so it does little beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Archive a todo') and immediately covers the inverse operation ('unarchive with archived=false'), which an agent can act on without opening the schema. It also distinguishes itself from siblings by explicitly referencing pad_archive, todo_list, and todo_get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for both modes of use and names the refusal precondition (todo still blocks an incomplete todo), which tells the agent when the call will fail. It stops short of naming an alternative for the delete-vs-archive decision, so it is not fully explicit about alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

todo_blockA

Add a blocker: todo_id cannot start until blocker_id completes. Cycles are rejected.

ParametersJSON Schema
NameRequiredDescriptionDefault
todo_idYes
blocker_idYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
todo_idYes
blocker_idYes
project_idYes
hive_noticeNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose the safety profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false), so the description's job is to add what they cannot. It does exactly that by disclosing that cycles are rejected — a concrete failure mode an agent must anticipate before calling. It stops short of saying what happens on a duplicate blocker, whether cross-project blockers are permitted, or whether any permission is required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, and the operative definition of the argument roles is front-loaded ahead of the constraint. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the annotations carry the mutation/idempotency profile. The description covers purpose, edge direction, and the cycle-rejection failure mode, which is enough to call it correctly; it is only slightly thin on the duplicate-edge and unblock-counterpart cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%, with todo_id and blocker_id documented as bare positive integers, so the description must carry their meaning — and it does, by defining the direction of the edge (todo_id waits on blocker_id), which is the single most important semantic here and is not derivable from the schema. project_id is richly documented in the schema itself, so it needs no help from the description; the remaining gap is minor.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb-plus-resource ('add a blocker') and defines the resulting semantic relation — todo_id cannot start until blocker_id completes — which is far more informative than the tool name alone. It does not, however, name its obvious counterpart sibling todo_unblock, so an agent must infer the reverse operation from the tool list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: an agent can infer this is the tool for expressing a dependency between two todos, and the sibling todo_unblock clearly reverses it. There is no explicit 'use when / do not use when' framing, no mention of todo_unblock, and no guidance on when a dependency should be modeled as a blocker versus some other mechanism.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

todo_commentA

Add a comment to a todo. Use for handoffs: changed files, tests run, decisions, remaining risk.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYes
todo_idYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
todo_idYes
comment_idYes
project_idYes
hive_noticeNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false and openWorldHint=false, so the safety profile is covered. The description adds the handoff framing but says nothing about whether comments are append-only, editable, or ordered, so it adds modest value beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, action first, elaboration second, with zero filler. The purpose is front-loaded before the handoff guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the tool is a simple 3-param write. However, the undocumented body/todo_id semantics leave an agent guessing about comment content constraints and todo_id validity given the 33% schema coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 33%: only project_id carries a description, while body and todo_id are bare in the schema and wholly unexplained in the description. With low coverage the description is expected to compensate, and it does not clarify comment length/format or what todo_id must reference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Add a comment to a todo'), which is clearly distinct from the todo_get/todo_update/todo_complete siblings by action and artifact. It never explicitly names or contrasts a sibling, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete usage context ('Use for handoffs: changed files, tests run, decisions, remaining risk') so an agent knows what content belongs here. It offers no when-not-to-use guidance or named alternative (e.g. todo_update for field changes), keeping it below a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

todo_completeA
Destructive

Mark a todo complete (or reopen with completed=false). Returns todo ids that this completion newly unblocked.

ParametersJSON Schema
NameRequiredDescriptionDefault
todo_idYes
completedNoDefaults to true.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
todo_idYes
completedYes
project_idYes
hive_noticeNo
newly_unblockedYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose destructiveHint=true, idempotentHint=false and readOnlyHint=false, so the safety profile is largely covered. The description adds genuinely non-obvious side-effect information — that completion newly unblocks other todos and returns their ids — which an agent could not infer from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the primary action and the flag behavior, with the return side effect last. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be spelled out, and annotations carry the destructive/idempotency profile. The description covers the mutation, its inverse, and the unblocking side effect; only permission or scope prerequisites are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, and the description adds real meaning for the 'completed' flag beyond the schema's 'Defaults to true' by explaining that false reopens the todo. todo_id's semantics are left to inference, but the dual-mode explanation pushes past the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Mark a todo complete') plus the inverse operation ('reopen with completed=false'), which is more precise than the bare name. It does not explicitly differentiate itself from siblings like todo_update or todo_archive, so it falls short of the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent how to reopen (completed=false), which is a real usage condition, and the schema's project_id guidance covers the scope question. There is still no when-not guidance or named alternative for general todo edits, so it is clear context rather than full routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

todo_createB

Create a project-scoped todo. Pass a short slug to have it render beside the id wherever this todo is listed; left unset, those surfaces fall back to the title instead (truncated in the dashboard). Optionally pass blocked_by todo ids to encode ordering. Returns a slim receipt.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNoObjective, owned files, acceptance criteria.
slugNoShort label, ~3-5 words (40 chars max). When set, it renders beside the id: bracketed in `hive todos`, `hive todo <id>`, and the session-start digest, and as the row's whole label in the dashboard. Left unset, those surfaces fall back to the title - shown in full in `hive todos`/`hive todo <id>`/the digest, truncated to this same length in the dashboard. Free text, not a pad-style slug. Pass "" to clear a previously-set slug back to unset.
tagsNo
titleYes
priorityNo
blocked_byNo
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
todo_idYes
project_idYes
hive_noticeNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (non-readOnly, non-destructive, non-idempotent), so the description carries a lower burden. It adds downstream behavioral context — where the slug renders and that truncation happens in the dashboard — plus 'Returns a slim receipt'. It does not state authorization requirements or any side effects beyond rendering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the core action, then parameter behavior, then the return. No filler. Slightly over-weighed toward the slug's rendering surface relative to the remaining six parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the return value needn't be explained ('slim receipt' suffices), and annotations cover safety. However, with 7 parameters and sub-50% schema coverage, the description leaves tags, priority, and title semantics to inference, which is thin for a creation tool with an ordering feature.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 43%, so the description should compensate. It explains the slug rendering contract and blocked_by ordering semantics, but body, tags, title, and priority receive no description-level explanation, and the slug/blocked_by detail largely duplicates what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource+scope: 'Create a project-scoped todo.' This clearly distinguishes creation from the sibling mutation tools (todo_update, todo_archive, todo_complete). It stops short of explicitly naming which sibling to prefer, but the operation is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives conditional usage for two parameters: when to pass a slug (optional, rendering effect) and when to pass project_id ('Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope'). This is real guidance, but it is parameter-scoped rather than tool-scoped — there is no statement of when to create a todo vs. use another tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

todo_getC
Idempotent

Read one todo in full: body, blockers, what it blocks, and optionally comments.

ParametersJSON Schema
NameRequiredDescriptionDefault
todo_idYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
include_commentsNo

TDQS

C2.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description says the tool 'Read[s]' a todo, but the annotations declare readOnlyHint=false, meaning the operation is not read-only. Nothing in the description explains this discrepancy or any side effect. This is an annotation contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the returned fields are enumerated compactly right after the verb and resource.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the return payload, which partially compensates. However, it leaves the annotation contradiction unexplained and does not document two of three parameters, so an agent lacks full context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%: todo_id and include_comments have no schema descriptions. The description's 'optionally comments' loosely maps to include_comments, but it says nothing about todo_id or the sensitive project_id override parameter, so it does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) and resource (one todo) and enumerates the returned content: body, blockers, blocked-by, comments. It is clearly distinguishable from the sibling todo_list by scoping to a single item, though it doesn't name the sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no prerequisites, and no mention of alternatives such as todo_list. The single-item scope hints at usage, but nothing tells the agent when this is preferable to listing or updating.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

todo_listA
Idempotent

List todo summaries. is_blocked=false finds dispatchable work. query matches title, body, and slug. Archived todos are excluded by default; include_archived=true retrieves them too.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNo
limitNo
queryNo
offsetNo
statusNo
priorityNo
is_blockedNo
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
include_archivedNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real behavior beyond the schema: archived todos are excluded by default (a default the schema's optional include_archived boolean cannot express), and query matches title, body, and slug rather than just title. Worth noting: the annotations set readOnlyHint=false while the description presents a purely read/retrieve operation, but the description never claims a mutation, so this reads as an imprecise annotation rather than a misleading description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short clauses, front-loaded with the core action and then the highest-value filters, with no padding or repetition. The clause-per-parameter style is slightly telegraphic but every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, nine parameters at 11% coverage, and no mention of pagination behavior (limit/offset defaults) or tags filtering, which an agent needs to page or narrow results correctly. The defaults it does cover (archived exclusion, dispatchable work) are the most important ones, so the definition is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 11% (only project_id is documented), so the description carries most of the burden and it explains three of nine parameters: query (which fields it matches), is_blocked, and include_archived. tags, limit, offset, status, and priority are left to the schema's bare enum/type declarations, including the pagination defaults that a list tool should state.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("List todo summaries") and immediately qualifies the scope with filter semantics, which is more than a restatement of the name. It does not, however, name or contrast itself with the obvious sibling todo_get or the other todo_* tools, so the agent must infer the list-vs-single distinction from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete selection cues: is_blocked=false "finds dispatchable work" and include_archived=true is the only way to see archived items. Both are when-to-use conditions an agent can act on. It stops short of naming an alternative tool or stating when not to use this one (e.g., use todo_get for a single todo).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

todo_unblockB
Destructive

Remove one blocker relationship from a todo.

ParametersJSON Schema
NameRequiredDescriptionDefault
todo_idYes
blocker_idYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
removedYes
todo_idYes
blocker_idYes
project_idYes
hive_noticeNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, so the agent knows this mutates state and repeated calls differ. The description adds essentially nothing beyond that—it doesn't say whether the blocker record is deleted or just the link, nor whether the blocker_id must pre-exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence that is front-loaded with the core operation. Efficient, though it could have spent a few more words on parameter/behavior clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema and annotations already covering safety, the description needn't explain returns. However, for a destructive two-required-param mutation it leaves the identity of todo_id vs blocker_id and the required relationship state unclear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%: todo_id and blocker_id are undocumented in both schema and description. The description never clarifies which todo is the blocked one versus the blocker, or the parameter meanings, leaving the two required params ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Remove') and resource ('blocker relationship from a todo'), which clearly distinguishes it from the inverse sibling todo_block. It does not, however, explicitly name or contrast itself with the sibling, keeping it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance, no prerequisites (e.g. that the relationship must exist), and does not reference the related todo_block or alternative tools. An agent must infer all usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

todo_updateB
Destructive

Update todo fields. Omitted fields are preserved. Returns a slim receipt.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNo
slugNoShort label, ~3-5 words (40 chars max). When set, it renders beside the id: bracketed in `hive todos`, `hive todo <id>`, and the session-start digest, and as the row's whole label in the dashboard. Left unset, those surfaces fall back to the title - shown in full in `hive todos`/`hive todo <id>`/the digest, truncated to this same length in the dashboard. Free text, not a pad-style slug. Pass "" to clear a previously-set slug back to unset.
tagsNo
titleNo
statusNo
todo_idYes
priorityNo
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
todo_idYes
project_idYes
hive_noticeNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false, so the safety profile is covered. The description usefully adds that omitted fields are preserved (partial update, not full replace) and that it returns a slim receipt, but does not explain what 'destructive' entails here — e.g., whether passing empty values clears fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short front-loaded sentences with no filler. "Returns a slim receipt" is slightly redundant since an output schema exists, but it is cheap and does no harm.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with annotations and an output schema, the safety and return concerns are largely handled elsewhere. Still, the undocumented parameters and the unstated meaning of its destructive hint leave the definition thinner than the tool's complexity warrants.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% — six of eight parameters (body, tags, title, status, priority, todo_id) carry no schema description, and the tool description adds no per-parameter meaning at all. It fails to compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (update todo fields), which is clear on its own. However, it offers no differentiation from sibling mutators like todo_complete, todo_block, or todo_archive, so an agent can't tell from the description why it would pick this over those.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Omitted fields are preserved" conveys partial-update semantics, which implies when the tool is appropriate (patch-style edits). But it never states when to prefer it over todo_complete/todo_archive for terminal states, nor any prerequisites, leaving routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wake_cancelA
DestructiveIdempotent

Cancel a pending wake-up you own, or - if you are a running lead - any pending wake-up in this project. Cancelling any wake also cancels the hold notices already filed about IT (modal-hold, unsubmitted-input, and one-shot block), since a notice about a wake that no longer exists has nothing left to say; unlike a finish notice these never expire on their own, since the thing they report may well still be true an hour later. Cancelling a standing watch also cancels the FINISH notices it has already filed but not yet delivered. It does NOT cancel a standing watch's own per-worker block notice (a crew member stopped on a dialog): that carries no parent link, so one already filed still delivers, and it may still be true - the worker is probably still on that dialog - but it no longer claims anything about the watch's own liveness, deliberately.

ParametersJSON Schema
NameRequiredDescriptionDefault
wake_idYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

Output Schema

ParametersJSON Schema
NameRequiredDescription
wake_idYes
cancelledYes
hive_noticeNo
cancelled_noticesYes

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (destructive=true, idempotent=true), the description discloses rich side-effect semantics: cancelling a wake also cancels its hold notices, cancelling a standing watch cancels undelivered FINISH notices, but deliberately does NOT cancel a watch's per-worker block notice. This is exactly the kind of non-obvious behavior an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose well, but the back half is a dense, comma-heavy exposition of notice semantics that is far longer than needed to convey the key point (which notices are and aren't cancelled). Several clauses could be tightened without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the description thoroughly covers ownership, scope, and side effects. The main gap is the undocumented wake_id parameter, but overall an agent has enough to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%; project_id is documented in the schema, but wake_id has no description. The description touches on project scope ('any pending wake-up in this project') yet does not clarify what wake_id refers to or its format beyond the schema's integer constraint. Baseline 3 is appropriate given partial coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('cancel') and resource ('pending wake-up') with an explicit scope rule: own wakes, or any project wake if you are a running lead. This clearly distinguishes it from siblings like wake_set, wake_get, and wake_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The ownership/lead-permission condition is clear implicit guidance for who may call it, and the project_id description in the schema reinforces staying in current scope. It does not, however, explicitly contrast with alternatives such as wake_update or state when cancellation is preferable to updating.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wake_getA
Idempotent

Read one wake-up by id, in this project, with its UNTRUNCATED body. wake_list truncates body at 120 chars; use this to see exactly what a wake will say, or to confirm what wake_update just changed.

ParametersJSON Schema
NameRequiredDescriptionDefault
wake_idYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations present, the bar is lower, but there is an unresolved tension: the description frames this as a pure read while readOnlyHint=false leaves open whether invoking it mutates state (e.g. marks the wake delivered/consumed). idempotentHint=true and destructiveHint=false are consistent with a read, but the description never addresses the readOnlyHint=false signal, nor does it describe the return shape. It does add real value by contrasting truncation behavior with wake_list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler. The core capability (untruncated read) is front-loaded and the sibling comparison follows immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must carry the return expectations; 'with its UNTRUNCATED body' and 'see exactly what a wake will say' do that adequately. The only gap is what happens to the wake's state after the read, which is a minor omission for an otherwise complete definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: project_id is well documented in-schema, wake_id is not. The phrase 'in this project' hints at the scoping semantics but adds no syntax, format, or id-provenance detail beyond what the schema already carries, so the baseline 3 fits.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (read one wake-up by id), a scope (in this project), and a distinguishing property (UNTRUNCATED body vs wake_list's 120-char cut). It names the sibling it differs from, so an agent can route without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names two use cases: seeing exactly what a wake will say, and confirming what wake_update just changed — both of which are conditions the alternative (wake_list) cannot satisfy. The project_id schema description reinforces the stay-in-scope default.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wake_listA
Idempotent

List pending wake-ups in this project, plus recently_delivered: the last 10 one-shot wakes that have already fired, with their delivery state (typed_at, held_at/held_reason, confirmation). A one-shot wake leaves the pending list the moment it fires; recently_delivered is where to check whether it was actually typed and, if its target has a confirmation channel, acknowledged.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

It discloses real behavior beyond the annotations: one-shot wakes leave the pending list the instant they fire, recently_delivered is capped at the last 10, and delivery state is exposed as typed_at / held_at+held_reason / confirmation. That is valuable operational context. It does not, however, explain why readOnlyHint is false for what reads as a listing operation, leaving a mild unexplained gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The primary result set is front-loaded in the first clause, and the second sentence explains the recently_delivered lifecycle efficiently. It is somewhat dense and packs several field names into one long sentence, but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining returns and does so well, enumerating both result sets and the delivery-state fields. The only shortfall is not addressing the readOnlyHint=false annotation for a list operation, which an agent might want clarified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single project_id parameter is fully documented there, including the 'use ONLY when explicitly asked' restriction. The description only restates the current-project scoping ('in this project') and adds no syntax or override nuance beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List pending wake-ups in this project') and goes further by naming the secondary result set (recently_delivered) and its contents. An agent can distinguish it from wake_get, wake_set, wake_cancel and wake_update without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly tells the agent when to consult recently_delivered ('where to check whether it was actually typed'), which is useful routing guidance. However, it never contrasts this tool with sibling retrieval tools like wake_get or states prerequisites for listing, so usage is only implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wake_setA

Schedule a wake-up: after delay_seconds the body is typed into the target session's terminal as a fresh user turn (prefixed [hive wake #N]). Defaults to delivering to THIS session. Use instead of polling. Write the body self-contained: ids, context, next action - it may arrive in a session that has none of this conversation. Delivering to your OWN lead pane, where the context is already there, prefer the action, the ids, and a pointer to where the detail lives.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYes
deliver_toNoDeliver to a spawned agent instead of this session.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
delay_secondsYes
repeat_every_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
due_atYes
wake_idYes
repeatingYes
deliver_toYes
hive_noticeNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover safety (readOnlyHint=false, destructiveHint=false, idempotentHint=false), and the description adds real mechanism the annotations cannot: the body is injected as a fresh user turn into a terminal, tagged with a wake counter, and defaults to THIS session. It omits failure/expiry behavior and how the wake can be cancelled (wake_cancel exists), which keeps it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core operation, then defaults, then the body-authoring guidance. Every sentence carries information, though the body-writing advice in the final sentence is somewhat verbose for the delivery it provides.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need not be described, and the definition covers the injection mechanism, the default target, and how to author the body. The remaining gap is repeat_every_seconds semantics and the wake lifecycle (cancellation/update) relative to its wake_* siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 40%; deliver_to and project_id carry schema descriptions, and the description usefully explains delay_seconds' effect and prescribes what body should contain. However repeat_every_seconds is left entirely undocumented in both schema and description, leaving a real gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (schedule a wake-up) and goes further by explaining the delivery mechanism: after delay_seconds the body is typed into the target session's terminal as a fresh user turn, prefixed [hive wake #N]. This is enough to distinguish it from wake_when_idle, wake_get, and wake_cancel without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use instead of polling" is an explicit when-to-use directive, and the closing sentence gives conditional guidance for the own-lead-pane case. It stops short of naming the sibling alternatives (wake_when_idle, agent_send) or stating when not to use it, so it is clear context rather than a full routing rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wake_updateA
Destructive

Edit a pending wake-up you own, in place, without minting a new id. Provide any subset of delay_seconds, body, repeat_every_seconds. delay_seconds is RELATIVE TO NOW, exactly as in wake_set: it moves the next fire time to now + delay_seconds. repeat_every_seconds only changes the interval used for firings AFTER this one; on its own it does not move the next fire time. Only a still-pending wake can be edited; use wake_get to read the result back. delay_seconds and repeat_every_seconds only apply to a delay wake (from wake_set) - an idle wake (from wake_when_idle) fires on watched-agent state and max_wait_seconds instead, so only body can be edited on one.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNo
wake_idYes
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
delay_secondsNo
repeat_every_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
due_atYes
updatedYes
wake_idYes
hive_noticeNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint=true, idempotentHint=false, readOnlyHint=false) by disclosing subtle runtime semantics: delay_seconds is relative to NOW, repeat_every_seconds only affects firings after the current one and does not move the next fire time by itself, and idle wakes accept only body. These are non-obvious mutation effects an agent could otherwise get wrong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then the caveats, with no filler sentences. It is information-dense but runs long, and the repeat/delay interplay takes two passes to parse fully.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-param mutation tool with an output schema (no need to describe returns), the description covers the pending-only constraint, field applicability per wake type, and the relative timing semantics. Nothing material is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%, so the description must compensate, and it does for delay_seconds, repeat_every_seconds and body with precise relative-vs-absolute semantics. wake_id and the project_id override are left to the schema (project_id has its own description), a minor remaining gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (edit a pending wake-up) and the key differentiator from wake_set: in-place edit without minting a new id. An agent can separate it from the create/cancel siblings immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly scopes usage: 'Only a still-pending wake can be edited', and spells out which fields apply to a delay wake vs an idle wake from wake_when_idle. It even names wake_get as the read-back alternative, so alternatives are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wake_when_idleA

Wake up when watched agents go idle (exact state from Claude Code hooks) or max_wait_seconds passes - except delivery HOLDS past that bound instead, for as long as the target pane is on a dialog or has unsubmitted human text in it, rather than pasting the wake body into either (.claude/rules/tmux-and-panes.md). Two shapes for workers, plus one for the queen, and you pass EXACTLY ONE of them. agents=[...] is a ONE-SHOT over a named list: mode=any fires on the first fresh idle transition, mode=all fires when every watched agent is idle (returns already_satisfied without scheduling anything if they all are now), and either way it stops watching once it fires. scope="project" is a STANDING WATCH over the crew you spawn in this project, including workers spawned later: it never stops watching, and on each finish it delivers a roster naming who finished and who is still going, until max_wait_seconds runs out or you wake_cancel it. You may hold ONE standing watch per project: a second call is refused and names the one already running, since two would report every finish twice. Use the standing watch when you are running more than one worker - a one-shot leaves every other worker unwatched from the moment it fires. Use either instead of polling. agents=[...] refuses a lead: it watches worker state, which a lead does not write. lead_project_id is the QUEEN's alone: a one-shot that fires when another registered project's running lead ENDS A TURN (never 'finished its work'), stored in and delivered to the queen's own project, and ended with a named reason if that lead's pane dies, is reissued, or restarts.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYes
modeNoDefaults to any. Only meaningful with agents.
scopeNoWatch the crew you spawn in this project as a STANDING watch that keeps watching after each finish, including workers spawned later. Mutually exclusive with agents.
agentsNoAgents to watch, as a ONE-SHOT. Mutually exclusive with scope.
deliver_toNoDeliver to a spawned agent instead of this session.
project_idNoDifferent project override. Use ONLY when the user explicitly asks for another project by name; otherwise stay in the current scope, even when results are empty.
lead_project_idNoQUEEN ONLY: wake when the running lead of this OTHER registered project ends a turn. The wake is stored in the queen's own project and delivered only to the queen; the watched lead is read, never typed into. Mutually exclusive with agents and scope; refuses project_id and deliver_to.
max_wait_secondsNoFor agents=[...]: how long to wait for idle before firing anyway, default 900. For scope="project": THE WATCH'S LIFETIME, default 14400 (4 hours), after which it delivers one last wake saying it has expired and stops watching. Not a hard deadline either way: delivery holds past it while the target pane is on a dialog or has unsubmitted text, until the pane clears.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeNo
noteNo
scopeNo
statusNo
wake_idNo
standingNo
watchingNo
deliver_toNo
expires_atNo
lead_watchNo
hive_noticeNo
watching_nowNo
max_wait_secondsNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far beyond the annotations, it discloses that delivery HOLDS past max_wait_seconds while the target pane is on a dialog or holds unsubmitted text, that only ONE standing watch per project is allowed and a second call is refused, that a one-shot stops watching after firing, that mode=all may return already_satisfied, and that lead_project_id ends with a named reason if the pane dies/is reissued/restarts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core behavior is front-loaded in the first sentence, but the remainder is one extremely dense run-on covering three call shapes with heavy parentheticals, which forces re-reading. Most clauses carry information, yet the packing hurts scannability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and the description still covers every edge case an agent needs: the three shapes, their mutual exclusions, default lifetimes, the dialog/typing delivery hold, the single-standing-watch limit, and lead end-of-turn semantics vs 'finished its work'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 88%, so the baseline is 3, but the description adds real meaning: mode is only meaningful with agents and defines any/all firing semantics; scope='project' never stops and delivers a roster each finish; max_wait_seconds carries different defaults (900 vs 14400) and is explicitly 'not a hard deadline' for either shape.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a concrete verb and resource (wake when watched agents go idle / max_wait_seconds passes) and enumerates the three call shapes (agents, scope, lead_project_id). It distinguishes those shapes internally, but never names or contrasts the generic wake_set/wake_cancel siblings, so an agent must infer the boundary from elsewhere.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing is provided: 'Use the standing watch when you are running more than one worker - a one-shot leaves every other worker unwatched from the moment it fires' and 'Use either instead of polling.' Mutual exclusions (scope vs agents, lead_project_id refusing project_id/deliver_to) are stated outright.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

whoamiA

Show this session's actor identity and effective project scope. Call this first in a new session.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description says 'Show', implying a read-only operation, but annotations declare readOnlyHint=false. This directly contradicts the annotations. Additionally, idempotentHint=false for an identity query is misleading, and no extra behavioral context is provided.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose followed by usage guidance. No wasted words, and the structure is optimal for a simple tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (0 params, no output schema) and the clear purpose, the description is complete enough. It tells the agent what it does and when to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline score is 4. The description does not need to add parameter semantics, and the schema coverage is 100% (though empty).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Show' and a specific resource 'this session's actor identity and effective project scope'. It clearly distinguishes from siblings like agent_status or agent_list, which concern other agents, not the current session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states 'Call this first in a new session', giving clear context for when to use it. However, it does not mention when not to use it or name any alternative tools, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv1.4.0
    • Changedagent_spawn2 fields changed
      • addedInput schema / properties / read_only
        Added value: +{
        +  "description": "Prevent local file writes and mutating shell commands while allowing Hive MCP tools. Only bare claude/codex executables. Extra args allow codex -c model_reasoning_effort=low|medium|high|xhigh, or claude --effort low|medium|high|xhigh|max; all other arguments are refused.",
        +  "type": "boolean"
        +}
      • addedOutput schema / properties / read_only
        Added value: +{
        +  "type": "boolean"
        +}
    • Changedlease_acquire1 field changed
      • changedInput schema / properties / key / description
        Previous value: -"Stable and specific, like \"file:src/api/routes.ts\"."New value: +"Stable and specific, like \"db:dev\"."
    • Changedproject_prune1 field changed
      • addedInput schema / properties / project_id
        Added value: +{
        +  "exclusiveMinimum": 0,
        +  "maximum": 9007199254740991,
        +  "type": "integer"
        +}
    • Addedqueen_audit_list
    • Changedwake_when_idle2 fields changed
      • addedInput schema / properties / lead_project_id
        Added value: +{
        +  "description": "QUEEN ONLY: wake when the running lead of this OTHER registered project ends a turn. The wake is stored in the queen's own project and delivered only to the queen; the watched lead is read, never typed into. Mutually exclusive with agents and scope; refuses project_id and deliver_to.",
        +  "exclusiveMinimum": 0,
        +  "maximum": 9007199254740991,
        +  "type": "integer"
        +}
      • addedOutput schema / properties / lead_watch
        Added value: +{
        +  "additionalProperties": {},
        +  "propertyNames": {
        +    "type": "string"
        +  },
        +  "type": "object"
        +}
  2. 28 tool updatesv1.2.2
    • Changedactor_prune1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedagent_close1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedagent_park1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedagent_rename1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedagent_resume1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedagent_spawn1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedkv_delete1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedkv_set1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedlease_acquire1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedlease_release1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedpad_append1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedpad_archive1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedpad_delete1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedpad_edit1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedpad_write1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedproject_add1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedproject_prune1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedtodo_archive1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedtodo_block1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedtodo_comment1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedtodo_complete1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedtodo_create1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedtodo_unblock1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedtodo_update1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedwake_cancel1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedwake_set1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedwake_update1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
    • Changedwake_when_idle1 field changed
      • addedOutput schema / properties / hive_notice
        Added value: +{
        +  "type": "string"
        +}
  3. 1 tool updatev1.2.0
    • Changedagent_list2 fields changed
      • addedInput schema / properties / before_id
        Added value: +{
        +  "description": "Page cursor for include_closed: only rows with id below this value. Pass the previous receipt's next_before_id to get the next page.",
        +  "exclusiveMinimum": 0,
        +  "maximum": 9007199254740991,
        +  "type": "integer"
        +}
      • addedInput schema / properties / limit
        Added value: +{
        +  "description": "Max rows to return when include_closed is true, newest first. Default 20, max 100. Ignored otherwise.",
        +  "maximum": 100,
        +  "minimum": 1,
        +  "type": "integer"
        +}
  4. 45 tool updatesv1.0.0
    • First observedactor_prune
    • First observedagent_close
    • First observedagent_list
    • First observedagent_message_get
    • First observedagent_output
    • First observedagent_park
    • First observedagent_rename
    • First observedagent_resume
    • First observedagent_send
    • First observedagent_spawn
    • First observedagent_status
    • First observedhelp
    • First observedkv_delete
    • First observedkv_get
    • First observedkv_list
    • First observedkv_set
    • First observedlease_acquire
    • First observedlease_release
    • First observedpad_append
    • First observedpad_archive
    • First observedpad_delete
    • First observedpad_edit
    • First observedpad_list
    • First observedpad_read
    • First observedpad_write
    • First observedproject_add
    • First observedproject_list
    • First observedproject_prune
    • First observedproject_select
    • First observedtodo_archive
    • First observedtodo_block
    • First observedtodo_comment
    • First observedtodo_complete
    • First observedtodo_create
    • First observedtodo_get
    • First observedtodo_list
    • First observedtodo_unblock
    • First observedtodo_update
    • First observedwake_cancel
    • First observedwake_get
    • First observedwake_list
    • First observedwake_set
    • First observedwake_update
    • First observedwake_when_idle
    • First observedwhoami

TDQS

A3.5/5.0

Scored across 46 tools

Disambiguation4/5

Tools are grouped by clear domain prefixes and most have distinct resource+action boundaries. Some pairs overlap slightly, such as agent_output vs agent_status and wake_set vs wake_when_idle, but the detailed descriptions make the intended use cases mostly separable.

Naming Consistency5/5

Nearly all names follow a predictable snake_case domain_action pattern: agent_spawn, pad_append, todo_complete, wake_cancel, kv_set, lease_acquire. A few standalone names like whoami and help are conventional exceptions rather than inconsistent naming.

Tool Count2/5

With 46 tools, the surface is well beyond the 3-15 well-scoped range and lands in the too-many category under the rubric. While the domain is broad, the set is heavy enough that discovery and selection will be burdensome for an agent.

Completeness4/5

The server covers most lifecycle operations for projects, agents, pads, todos, KV, leases, and wakes, including archive, blocking, comments, and audit surfaces. Minor gaps exist, such as no lease listing/status tool and no direct todo delete, but agents can work around these.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables inter-session communication and coordination for multiple Claude Code instances through a shared SQLite database. Supports real-time messaging, shared state management, and resource locking to facilitate parallel development workflows between AI agents.
    11 npm
    8
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Collects session JSONL from Codex and Claude Code into local SQLite memory, exposed through MCP with tools for memory context, search, get, put, forget, sleep, and status.
    -
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables coordinating Claude Code and Codex across separate Git worktrees with shared issue ownership, file reservations, messages, and explicit handoffs.
    265 PyPI
    MIT