Skip to main content
Glama

Brainy

CI License: AGPL-3.0 Dependencies: none

A shared brain for your AI agents — and for the humans who work with them.

Brainy is a self-hosted knowledge and task backbone that Claude, ChatGPT and any other MCP-capable AI connect to through one controlled endpoint. Every agent reads the same knowledge, works on the same task list and leaves an audit trail — so your AIs can hand work to each other instead of living in separate chat silos.

   ChatGPT ─┐                                   ┌─ Git-versioned knowledge (Markdown)
   Claude  ─┼──►  MCP endpoint  ──►  Brainy  ───┼─ Tasks with atomic claim/lease
   Your bot ┘     (OAuth / tokens)   ACL+Audit  └─ Spaces, roles, append-only audit log
                                       ▲
                          Humans: web admin UI (+ optional Telegram approvals)

Why Brainy?

  • AIs that collaborate. ChatGPT creates a task, Claude claims it, does the work and writes the result back; a human reviews. Coordination happens through tasks — no hidden agent-to-agent magic.

  • One source of truth. Knowledge lives as plain Markdown in a Git repository. Every write is a commit, so you get history, diffs and rollback for free.

  • Safe by default. No shell, no filesystem, no eval over MCP. Path allowlist, per-space ACLs, secret detection on writes, rate limits, hashed tokens and an append-only audit log.

  • No double work. Atomic task claims with leases and claim tokens guarantee that two agents never process the same task at the same time.

  • Human in the loop. Agents can propose knowledge changes; a human approves or rejects them (web UI or Telegram). Tasks can require review/approval before they count as done.

  • Zero dependencies. Pure Python standard library + SQLite. No pip install, tiny attack surface.

Related MCP server: auxly-memory-cli

Why not just a markdown file?

A shared notes file in a Git repo is a great knowledge store — and Brainy keeps that part on purpose: knowledge is plain Markdown in Git, with history, diff and rollback. The problem is everything a bare file doesn't do once more than one agent touches it:

  • Two agents overwrite each other. Nothing serializes concurrent writes to the same file. Brainy writes go through optimistic concurrency (expected_git_commit) plus a server-side write lock: exactly one writer commits, the others get a Conflict and re-read — no lost edits.

  • No task hand-off. A file can list TODOs, but it can't stop two agents from grabbing the same one. Brainy tasks use an atomic claim (UPDATE ... WHERE status='READY', one row): exactly one agent wins, the rest get a conflict and move on.

  • No per-area permissions. A file is all-or-nothing. Brainy has Spaces with per-space ACLs, so an agent only reads/writes the areas you grant it.

  • No approval step. Agents can propose changes for a human to approve (web UI or Telegram) instead of writing straight to the source of truth.

  • "Who changed what?" beyond git blame. Every action — reads of task state, claims, writes, rejections — lands in an append-only audit log enforced by a DB trigger.

Both guarantees above are covered by a concurrency test (tests/test_concurrency.py): 20 threads racing on one task yield exactly 1 winner and 19 conflicts, and 20 threads writing the same file yield exactly 1 commit and 19 conflicts, with the repo left clean.

How is Brainy different?

Honest comparison with the tools people mention most. Beads is closest to Brainy's task side, and Central Brain to its knowledge side; neither is a one-to-one match. Facts are from each project's own docs (links below); where something isn't documented, it says so rather than guessing.

Brainy

Beads beads

Central Brain cb

Shared notes file

What it is

Knowledge base + task queue

AI-agent issue/task tracker

AI memory layer

Markdown in a repo

Knowledge store

Markdown in Git

— (issue tracker)

Plain files + local vector index

Markdown in Git

Task queue with atomic claims

✅

✅

not documented

❌

Per-area permissions (ACL)

✅ Spaces/roles

❌

❌

❌

Append-only audit log

✅ (DB-enforced)

partial (Git/DB history)

❌

git blame only

Human approval for writes

✅ (web/Telegram)

❌

❌

❌

MCP server

✅

✅

✅

❌

Multi-agent concurrency control

✅ claims + write lock

✅ claims + concurrent writers

not documented

❌

Dependencies

none (stdlib + SQLite)

bundles Dolt (Go)

bundles embedding model

n/a

Self-hosted

✅

✅

✅ (cloud optional)

✅

License

AGPL-3.0 (+ commercial)

MIT

proprietary

n/a

version-controlled SQL database; earlier versions used SQLite + JSONL).
Apple-silicon macOS, MCP integration; pricing per the site at time of writing: $12/mo solo,
$35/mo for 4 licenses).

Features

Area

What you get

Knowledge

list_documents, get_document, search_knowledge, write_document (optimistic concurrency via Git commit), append_document, propose_write

Tasks

create_task, claim_task, renew_claim, complete_task, fail_task, release_task, dependencies, priorities, review/approve/reject

Resource claims

claim_resource, renew_resource, release_resource, list_resource_claims — reserve any resource (a free-form key per space, e.g. a file path or area) with an atomic lease, same model as task claims

Access

Spaces (tenants/areas), roles ADMIN / EDITOR / AGENT / READER, per-space ACL, service tokens, OAuth 2.1 (for Claude/ChatGPT remote connectors)

Operations

Web admin UI, audit log, agent registry + dispatcher framework, backup & verified restore scripts

Quickstart (Docker)

Try it in one command — no clone, no setup. Runs a throwaway demo (temporary database + example knowledge, deleted on exit) as an MCP stdio server:

docker run --rm -i ghcr.io/memokar/brainy:latest --demo

For a real install, run the server (prebuilt image, published on every release):

docker run -d --name brainy -p 127.0.0.1:8765:8765 -v brainy-data:/data ghcr.io/memokar/brainy:latest
docker logs brainy   # prints your one-time ADMIN token on first start

Or build it yourself with Compose:

git clone https://github.com/memokar/brainy.git
cd brainy
docker compose up -d
docker compose logs brainy   # prints your one-time ADMIN token on first start

Brainy now listens on http://127.0.0.1:8765 (MCP endpoint: /mcp, admin UI: /admin — log in with the token). For remote AI connectors put it behind HTTPS (see deploy/nginx-brainy.conf.example) and set BRAINY_PUBLIC_BASE_URL.

Quickstart (bare metal, Linux, Python ≥ 3.10)

export BRAINY_DB_PATH=$PWD/data/brainy.db
export BRAINY_KNOWLEDGE_ROOT=$PWD/data/knowledge
export BRAINY_WEB_SESSION_KEY=$PWD/data/web_session.key

cp -r examples/knowledge "$BRAINY_KNOWLEDGE_ROOT"
git -C "$BRAINY_KNOWLEDGE_ROOT" init -q && git -C "$BRAINY_KNOWLEDGE_ROOT" add -A \
  && git -C "$BRAINY_KNOWLEDGE_ROOT" commit -qm "initial knowledge"

python3 scripts/bootstrap.py "$BRAINY_DB_PATH" --with-token   # prints ADMIN token once
python3 scripts/init_prod_db.py "$BRAINY_DB_PATH"             # seeds default spaces
python3 scripts/serve.py

Connecting an AI

  • Claude Code:

    claude mcp add --transport http brainy http://127.0.0.1:8765/mcp --header "Authorization: Bearer <token>"
  • Local stdio clients (e.g. Claude Desktop): run Brainy as a subprocess:

    {"mcpServers": {"brainy": {"command": "python3", "args": ["/opt/brainy/scripts/stdio.py"],
      "env": {"BRAINY_DB_PATH": "/var/lib/brainy/brainy.db",
              "BRAINY_KNOWLEDGE_ROOT": "/opt/brainy-knowledge", "BRAINY_TOKEN": "<token>"}}}}

    Try it without any setup: python3 scripts/stdio.py --demo (temporary data, deleted on exit).

  • Any other MCP client with custom headers: endpoint https://<your-host>/mcp, header Authorization: Bearer <service token>.

  • Claude.ai / ChatGPT remote connectors: use the OAuth 2.1 flow (discovery at /.well-known/oauth-authorization-server). Set BRAINY_PUBLIC_BASE_URL to your HTTPS URL.

Brainy is listed in the official MCP Registry as io.github.memokar/brainy.

Give each AI its own principal (e.g. claude, chatgpt) with role AGENT and only the spaces it needs. Every action then shows up in the audit log under that name.

Example: two coding agents in one repo

Two agents working the same repository can step on each other's files. With resource claims each one reserves the area it's about to touch — a free-form key per space — and the other backs off until it's released or the lease expires:

// Agent A, before editing the auth code:
claim_resource   { "space": "shared", "resource_key": "repo:app/src/auth/**", "lease_seconds": 1800 }
// -> { "claim_token": "…", "lease_until": "…" }   A now owns that area

// Agent B, about to touch the same area:
claim_resource   { "space": "shared", "resource_key": "repo:app/src/auth/**" }
// -> Conflict: already claimed  → B picks a different area (e.g. "repo:app/src/api/**")

renew_resource   { "space": "shared", "resource_key": "repo:app/src/auth/**", "claim_token": "…" }   // A keeps working
release_resource { "space": "shared", "resource_key": "repo:app/src/auth/**", "claim_token": "…" }   // A done → free again
list_resource_claims { "space": "shared" }   // who holds what right now

Same guarantee as task claims — exactly one holder at a time, expired leases free up automatically (see tests/test_resources.py). Resource claims reuse the task permissions/scopes (can_claim_tasks / brainy:tasks:write), so any agent that can claim tasks can claim resources.

Configuration

All configuration comes from environment variables — see .env.example. Secrets (tokens, keys) are never stored in the repository or the knowledge base.

Extensions

The core ships a worker plugin interface and a deterministic MockWorker. Real workers that let agents execute tasks autonomously (e.g. Claude Code, Codex) are separate extensions loaded via BRAINY_WORKER_PLUGINS. See docs/extensions.md.

Running the tests

for t in tests/test_*.py; do python3 "$t" || exit 1; done

Status & roadmap

Brainy runs in production for its author. Current limitations:

  • Code comments are still partly German; all user-facing text is English.

  • Single-node design (SQLite). PostgreSQL only if real multi-writer load appears.

License

Copyright (C) 2026 Mehmet Karakolcu

Brainy is dual-licensed:

  • Open source: GNU AGPL-3.0. Free for everyone, including companies — but if you modify Brainy and offer it to others (also as a network service), you must publish your changes under the AGPL.

  • Commercial license: for companies that want to use or embed Brainy without AGPL obligations. See COMMERCIAL.md.

Contributions require agreeing to the Contributor License Agreement.

  1. Beads — https://github.com/steveyegge/beads (MIT, Go; stores issues in Dolt, a

    ↩
  2. Central Brain by NeuroAIgent — https://neuroaigent.ai (local-first AI memory, Windows +

    ↩

Available Tools

27 tools
append_documentC
Destructive

Append to an allowed target (e.g. shared/todo.md).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
textYes
commit_messageYes
expected_git_commitNo

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false, so the safety profile is covered structurally. The description adds only the vague 'allowed target' allowlist notion and no detail on commit behavior, the effect of expected_git_commit conflicts, or what mutation actually occurs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence, which is front-loaded, but the brevity is under-specification rather than economy. The one sentence does not earn its place by covering enough ground for a 4-parameter mutation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent write tool with four undocumented parameters, no output schema, and no sibling differentiation, the description is far too thin. An agent has no basis for choosing this over write_document or for supplying commit_message/expected_git_commit correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the burden for four parameters. It only implies a path-like target via the 'shared/todo.md' example and says nothing about text, commit_message, or expected_git_commit semantics (e.g. optimistic-concurrency checking).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Append') and gives a concrete example target ('shared/todo.md'), so the operation is recognizable. However, the resource is described only as an 'allowed target', and it is not distinguished from the sibling write_document or propose_write, leaving ambiguity about when this is the right tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus write_document, propose_write, or get_document. The phrase 'allowed target' hints at an allowlist constraint but never explains what qualifies or what happens when the target is not allowed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

approve_taskA
Idempotent

Approve the approval gate (-> APPROVED). Requires approve permission; no self-approval.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
task_idYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it is a non-destructive, idempotent mutation (readOnlyHint=false, idempotentHint=true, destructiveHint=false). The description adds meaningful context beyond that: an auth requirement (approve permission) and a constraint (no self-approval), plus the concrete state transition to APPROVED.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the action and resulting state are front-loaded and the constraints follow. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with annotation safety coverage and no output schema, the state transition and auth/constraint context make it mostly usable. The unresolved gap is the completely undocumented parameters, especially the purpose of the optional note.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Two parameters with 0% schema description coverage, so the description carries the full burden and delivers nothing about either. The required task_id and the optional note are never explained, leaving the agent to guess what note is for and what task_id references.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Approve) and resource (the approval gate), plus the resulting state (-> APPROVED), so the operation is unambiguous. It does not explicitly name the contrasting sibling reject_task, but the verb itself is distinctive enough to separate it from the other task tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides prerequisites (requires approve permission; no self-approval) which is genuine usage guidance. However, it never states when to choose this over reject_task or review_task, so routing among gate-related siblings is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claim_taskC

Claim a task atomically (lease + claim_token).

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYes
lease_secondsNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare the mutation profile (readOnly=false, idempotent=false, destructive=false), so the safety bar is lower. The description usefully adds 'atomically' and mentions the lease/claim_token artifacts, but omits the most decision-relevant behavior: what happens if the task is already claimed, how the lease expires, and whether claim_token is returned or supplied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or repetition. Brevity is appropriate, though the parenthetical packing two concepts into one clause slightly clouds the meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent mutation tool with no output schema and 0% parameter coverage, the description is too thin: it does not explain failure modes, lease duration semantics, or what a successful claim returns (the claim_token is mentioned but not framed as an output). An agent could not reliably invoke or recover this tool from the definition alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden and mostly fails: task_id is unexplained (though self-evident) and lease_seconds is only obliquely referenced as 'lease' with no default, unit, or range. The mention of claim_token is ambiguous because it is not a parameter at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (claim) and resource (task), plus the concurrency semantics (atomically) that distinguish this from a plain update. It does not name the closely related siblings (renew_claim, release_task), so an agent must infer those boundaries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the obvious alternatives in the sibling list (renew_claim for extending a claim, release_task for giving one up). The parenthetical hints at the mechanics but gives no routing information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

complete_taskB
Idempotent

Complete a task — governance-gated (AUTO->COMPLETED, REVIEW->AWAITING_REVIEW, APPROVAL->AWAITING_APPROVAL).

ParametersJSON Schema
NameRequiredDescriptionDefault
resultNo
task_idYes
claim_tokenYes
result_refsNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, so safety is covered. The description adds genuinely useful behavior beyond that — the outcome depends on governance mode (AUTO->COMPLETED, etc.) — but it does not explain the required claim_token, whether completion can be undone, or what happens to a partial result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence, front-loaded with the verb and resource. The parenthetical state mapping is dense but every element carries meaning, and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with annotations covering safety and no output schema, the description covers the governance outcome but leaves the required claim_token's role and the meaning of result/result_refs unexplained. An agent can select the tool but may not invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are 4 parameters, yet the description names none of them. The required task_id and claim_token and the optional result/result_refs are entirely undocumented, so the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Complete) and resource (task) and even enumerates the resulting states per governance mode, so the operation is unambiguous. It does not, however, differentiate itself from siblings like approve_task, review_task, or fail_task, which are the natural alternatives for ending a task's lifecycle.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The governance-state transitions imply when completion is legitimate, giving some context. But there is no explicit when-to-use guidance and no mention of the sibling tools (approve_task, review_task, fail_task) an agent should consider instead, so routing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_taskD

Create a task.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNo
spaceYes
titleYes
statusNo
priorityNo
risk_levelNo
target_refNo
descriptionNo
action_classNo
dependenciesNo
execution_modeNo
preferred_agentNo
required_capabilitiesNo

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose this is a non-read-only, non-idempotent, non-destructive mutation. The description adds no behavioral context beyond that: nothing about required permissions, what a duplicate call does given idempotentHint=false, or side effects. It does not contradict the annotations, but contributes nothing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence is not conciseness here; it is under-specification. There is no front-loaded intent, scope, or return behavior to justify the brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter mutation tool with no output schema and no parameter documentation anywhere, the description is wholly inadequate. An agent has no basis to supply sensible values for the optional fields or predict the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 13 parameters and 0% schema description coverage, the description carries the full explanatory burden and says nothing about any parameter. Critical fields like space, title, status, priority, action_class, dependencies and required_capabilities are left entirely opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Create a task" names a verb and resource but is essentially a tautological restatement of the tool name with no scope, no distinction from siblings like claim_task or propose_write. It conveys nothing an agent couldn't infer from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many related siblings (approve_task, claim_task, propose_write, write_document). No prerequisites, no conditions, no alternatives are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

disable_agentC

Disable an agent (ADMIN).

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_nameYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false, so the description carries the behavioral burden. It adds the ADMIN auth requirement, which is genuine value, but says nothing about what disabling actually does (agent state, in-flight tasks, reversibility) for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence, front-loaded with the verb and resource, no filler. It is efficient, though the brevity edges toward under-specification rather than optimal conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, minimal annotations, and an undocumented parameter, the description is too thin. It should at least describe the effect of disabling and the agent_name argument.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single 'agent_name' parameter is undocumented. The description adds no detail on whether this is a name, ID, or path, or how it is resolved, leaving the agent to guess the accepted format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Disable an agent'), so the operation is unambiguous and pairs naturally against the sibling 'enable_agent'. It does not explicitly name that sibling or delimit scope, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The '(ADMIN)' tag signals a required privilege level, which is useful, but there is no guidance on when to disable vs. use a sibling like enable_agent, pause_dispatcher, or release_task. No preconditions or exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

enable_agentB

Enable an agent (ADMIN).

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_nameYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=false already tells the agent this is a mutation, so the description needn't restate that. It does add real context annotations do not carry: the operation requires ADMIN privileges. It says nothing about reversibility or what state change occurs on the agent, leaving that gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence with no filler, and the scoping/privilege note is front-loaded. Efficiency is high, though the terseness borders on under-specification rather than polished conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter toggle with no output schema, the definition is minimally adequate. It omits what enabling actually does, error behavior for unknown agents, and the relationship to disable_agent, so it is not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single agent_name parameter, but the name itself is self-describing. The description implies the parameter identifies the target agent but adds no format, ID-vs-name, or existence requirements beyond what the schema shows.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Enable an agent'. The sibling set contains disable_agent, so the operation's polarity is unambiguous. It stops short of explicitly differentiating from that sibling, which keeps it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The '(ADMIN)' marker signals a permission precondition, which is useful, but there is no guidance on when to use this versus disable_agent or how it relates to task/claim workflows. No exclusions or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fail_taskC
Idempotent

Mark a task as failed (idempotent).

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNo
task_idYes
claim_tokenYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=true, so the parenthetical '(idempotent)' merely repeats structured data and earns no credit. The description adds nothing about what failing a task does beyond the state change, whether the claim is consumed, or any required permission.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is efficient and front-loaded with no wasted words, but the terseness reflects under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-mutating tool with a required claim token, no output schema, and zero parameter documentation, the description is too thin. It omits the claim prerequisite and the effect of the operation, which an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across all three parameters, so the description carries the full burden of explaining them and does not. It says nothing about task_id, the required claim_token, or the optional reason, leaving the meaning of claim_token entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Mark a task as failed'. An agent can tell this transitions a task into a failed state. However, it offers no differentiation from closely-related siblings like reject_task or complete_task, which also terminate a task's lifecycle.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus reject_task, complete_task, or release_task, nor any stated prerequisites (such as needing to hold a claim first). The agent must infer positioning in the task lifecycle entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_statusC
Read-onlyIdempotent

Status of an agent.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_nameYes

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered structurally. The description adds nothing beyond that: no indication of what "status" contains (running/paused/failed), whether an unknown agent_name errors or returns empty, or any polling/rate-limit caveat.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is one short, front-loaded clause with no filler, which is appropriate in form. However, the brevity comes from under-specification rather than efficiency, so it only meets the minimum bar.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of explaining what the returned status looks like and how to interpret it, which it entirely omits. For a lookup tool with a required parameter, this leaves the agent guessing about the response shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter has 0% schema description coverage and the description does nothing to compensate. It doesn't say whether agent_name is a display name or an identifier, nor where to obtain a valid value (e.g., from list_agents).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Status of an agent" is essentially a restatement of the tool name (get_agent_status) with no verb and no scope. It doesn't differentiate from siblings like list_agents or get_task, which an agent could easily confuse with it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no mention of alternatives such as list_agents or get_execution_job. The agent must infer that this is for polling a single named agent's state.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_documentB
Read-onlyIdempotent

Read a knowledge document (path allowlist, ACL).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds genuinely new context beyond them: access is gated by a path allowlist and ACL, which tells the agent reads can fail for authorization reasons. It stops short of describing failure behavior or return content.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the core action front-loaded and no padding. It is efficient, though the brevity borders on under-specification rather than being tight and complete.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A simple one-parameter read tool with no output schema, so the description need not explain return values, but it also gives no indication of what a successful read returns or how the ACL/allowlist manifests as an error. Adequate for a low-complexity tool, with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One required parameter with 0% schema description coverage, so the description must carry the load. 'path allowlist' hints that the path is validated against permitted locations, but no format, root, or examples are given, leaving the single parameter only partially clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read) and resource (knowledge document), which separates it from task/agent siblings like get_task or get_agent_status. However it does not explicitly distinguish itself from the closest alternatives, list_documents and search_knowledge, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance and no named alternative. An agent must infer on its own that this fetches a single document by path rather than listing or searching, despite three relevant siblings existing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_execution_jobB
Read-onlyIdempotent

Read an execution job (without dispatch_token).

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral fact — that the returned job omits dispatch_token — but says nothing about error behavior or what a missing/invalid job_id yields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no padding, and the core action is front-loaded before the parenthetical caveat. It is efficient, though arguably under-specified rather than concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and only one required parameter, so the bar is low and the annotations carry the safety story. Even so, for a fetch tool the description omits how to obtain job_id and what the job record contains, and the dispatch_token remark is unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter job_id has no description in either place. The description never explains what a job_id is, its format, or where to obtain it (presumably from list_execution_jobs), leaving the only parameter under-specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Read an execution job"), which cleanly separates it from the plural sibling list_execution_jobs. However, it never names that sibling or otherwise differentiates explicitly, and the parenthetical about dispatch_token is cryptic rather than clarifying.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no routing to alternatives such as list_execution_jobs. The agent is left to infer that this fetches one job by id rather than enumerating jobs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_taskC
Read-onlyIdempotent

Read a task.

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that — no note on whether the task must exist, whether it returns stale data, or what happens for an unknown/invalid task_id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded sentence with zero padding, which is structurally clean. But the brevity comes from under-specification rather than efficiency — the sentence does not carry enough weight to justify its place as the tool's entire guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a bare input schema, the description should at minimum indicate what a returned task contains or how failures surface. It leaves the agent without any behavioral or return context beyond the read-only annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter task_id has 0% schema description coverage and is not explained in the description at all — no format, source, or example is given. The description therefore fails to compensate for the schema gap on the one parameter the agent must supply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Read a task" gives a clear verb (read) and resource (task), so the basic operation is identifiable. However, it offers no differentiation from siblings like list_tasks, review_task, or get_execution_job, which also surface task information, so an agent cannot tell precisely when this one is the right pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of alternatives, and no prerequisites or exclusions. With a crowded task-tool family (claim_task, approve_task, review_task, list_tasks), the description does nothing to route the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_agentsB
Read-onlyIdempotent

List the agent registry (no secrets).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral note — output contains no secrets — but says nothing about pagination, ordering, or result size limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler, but it is so terse that there is essentially no structure to evaluate beyond the leading verb.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description bears some responsibility for describing what is returned; 'the agent registry (no secrets)' gives only a partial picture (no field list, pagination, or registry contents). For a zero-parameter read this is adequate but leaves clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for 0 params is 4. No parameter-related gaps exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the agent registry') and adds a scope qualifier ('no secrets') that distinguishes it from a credential-dumping read. It does not explicitly name siblings like get_agent_status, but the plural listing verb makes the distinction reasonably clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance and does not point to alternatives such as get_agent_status for a single agent. The only usage signal is the implicit notion that 'list' means enumerating all agents.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_documentsC
Read-onlyIdempotent

List canonical Markdown documents (optional space).

ParametersJSON Schema
NameRequiredDescriptionDefault
spaceNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds the "canonical Markdown" qualifier, suggesting drafts or non-canonical variants are excluded, which is modest but real behavioral context. It says nothing about pagination, ordering, or result volume, which matters for a list operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the qualifier front-loaded and no filler. It is efficient, though the brevity comes at the cost of the detail noted in other dimensions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list tool with an undocumented parameter, no output schema, and no annotation-covered return behavior, the description is too thin. An agent still does not know what a returned document listing looks like or how the space filter shapes results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single "space" parameter, and the description only repeats its name and optionality, both of which the schema already conveys. It does not explain what a "space" is or what happens when the filter is omitted, so the compensation burden is not met.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ("List ... documents") plus a qualifier ("canonical Markdown") that separates it from write_document and append_document. It does not, however, distinguish itself from get_document or search_knowledge, which are the closest siblings an agent would weigh.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus get_document (single fetch) or search_knowledge (query-based). The parenthetical "(optional space)" describes the parameter, not the usage context, so the agent must infer the routing decision.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_execution_jobsC
Read-onlyIdempotent

List execution jobs (without dispatch_token).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
statusNo
task_idNo

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered structurally. The description does add one useful behavioral detail beyond the annotations: results omit the sensitive dispatch_token. It says nothing about default limits, pagination, or ordering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the operation front-loaded and no padding. It is efficient, though its brevity borders on under-specification rather than purposeful concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter list tool with no output schema and zero schema documentation, the description should explain the filters and return shape. As written, an agent knows only that jobs are listed without dispatch_token, which is insufficient to call the tool confidently with filters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all three parameters (limit, status, task_id), and the description adds no semantics for any of them. It mentions dispatch_token, which is an output field, not a documented input, so it does not compensate for the parameter gap at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource ("List execution jobs") and scopes it with "(without dispatch_token)", which tells the agent the token field is excluded from results. It is distinguishable from the singular sibling get_execution_job only by the list-vs-get verb; the description never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus get_execution_job, nor any mention of filtering by status or task_id, even though those parameters exist. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_runnable_tasksB
Read-onlyIdempotent

Tasks runnable by this principal (READY, deps satisfied, ACL, required_capabilities, preferred_agent).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description usefully adds the filtering semantics, which is the most important behavioral fact here. However it says nothing about ordering, pagination behavior, or what a returned task record contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded line with no filler; the scope leads and the filter criteria follow. It is arguably too terse, omitting entirely rather than padding, but every fragment earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries some burden for return shape, which it does not describe beyond the implicit 'task list' and the filter conditions. The enumerated gating criteria partly compensate for that gap, making it minimally adequate rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter (limit) at 0% schema description coverage, and the description never mentions it or its bounding/pagination behavior. With a documented param left entirely to the schema and no compensating detail, the description does not earn the 0-param baseline of 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (tasks) with a scoping qualifier, 'runnable by this principal', that separates it from the sibling list_tasks by implying only executable-now tasks are returned. The parenthetical enumerates the gating conditions (READY, deps satisfied, ACL, required_capabilities, preferred_agent). It does not explicitly name or contrast with the sibling, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no named alternative. The scope phrase 'runnable by this principal' hints at the use case (finding tasks I can actually claim/execute, versus the general list_tasks), but the agent must infer that routing decision itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_spacesB
Read-onlyIdempotent

List visible knowledge spaces.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is fully covered structurally. The description's only added behavioral signal is 'visible', implying results are filtered by the caller's access permissions rather than returning every space. It says nothing about pagination, ordering, or result volume.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no filler. It is appropriately sized for a parameterless list tool, though it is arguably too terse to convey context about what a space is.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only list, the description is close to sufficient, but in a 26-tool set with search_knowledge and list_documents it gives no indication of how spaces relate to documents or why an agent would enumerate them first. A sentence on the space/document hierarchy would close the gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies; there is nothing in the schema that the description could add meaning to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('knowledge spaces'), with a scope qualifier ('visible') that hints at access-scoped results. It does not explicitly contrast with siblings, but no sibling shares the 'space' resource, so the agent can distinguish it readily. The nature of a 'space' itself is left undefined.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as search_knowledge or list_documents that operate on content within spaces. The usage is only weakly implied by the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tasksC
Read-onlyIdempotent

List tasks of a readable space.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNo
limitNo
spaceNo
statusNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so the description only needs to add context beyond safety. It adds none: it never mentions pagination behavior, defaults, or result ordering despite a `limit` parameter implying paged output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words, but it is terse to the point of under-specification rather than efficiently concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter listing tool with no output schema and no parameter documentation anywhere, the description is too thin. An agent cannot determine valid values for type/status or how limit/space affect results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters (type, limit, space, status) with no enums. The description only alludes to "space" and explains none of the other three, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb+resource ("List tasks") is present, but the qualifier "of a readable space" uses undefined terminology and does not distinguish this tool from the sibling list_runnable_tasks, which also lists tasks. The core purpose is inferable but the scope is vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no alternatives named. Given the crowded sibling set (list_runnable_tasks, get_task), the description gives the agent nothing to route on.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pause_dispatcherA

Pause the dispatcher globally (ADMIN, kill switch).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint=false annotation confirms this is a mutation, and the description adds genuinely new context beyond it: the operation requires ADMIN privileges and acts as a global kill switch. It does not state reversibility or what happens to in-flight tasks, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no waste, front-loading the action, its scope, and the privilege requirement. Every element earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter mutation with no output schema, the essentials — action, scope, and admin requirement — are covered. Minor gaps remain around reversibility and the effect on running agents/tasks, but they are not critical for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to convey; the baseline for a parameterless tool is 4. Schema coverage is 100% trivially and nothing is missing on this dimension.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Pause the dispatcher') plus a scope qualifier ('globally'), making the action unambiguous. It does not explicitly name the complementary sibling resume_dispatcher, so it falls short of full sibling differentiation, but the verb makes the contrast obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Calling it a 'kill switch' implies emergency/urgent usage and 'globally' implies broad impact, so the when-to-use is implied. However, no alternative is named and no explicit condition for pausing versus resuming is stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

propose_writeA

PROPOSE a change to a knowledge document (no direct write). The owner approves via Telegram; only then is it committed. Use this when you do not have direct write permission.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
contentYes
commit_messageYes
expected_git_commitNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, which is a thin profile. The description adds real behavioral context: the write is deferred, an owner approves via Telegram, and only then is it committed. It does not say what happens on rejection or how long approval takes, but the deferred-commit model is the key trait.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place, with the no-direct-write constraint and approval flow front-loaded. No waste or repetition of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Behavioral flow is well covered, but with no output schema the agent has no idea what a successful propose returns (a task id? a pending-approval handle?), and the undocumented params leave gaps. Adequate but incomplete for a mutation-style tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters, and the description provides no per-parameter meaning at all. Names like 'content' and 'commit_message' are mostly inferable, but 'expected_git_commit' is an optimistic-concurrency token whose semantics are entirely unexplained in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (PROPOSE a change) and resource (knowledge document), and immediately clarifies scope with '(no direct write)', which separates it from the write_document sibling. An agent can tell what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives the selecting condition: 'Use this when you do not have direct write permission.' It does not name write_document as the alternative, but the condition is clear enough to route correctly between the two.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reject_taskA
Idempotent

Reject at the approval gate (-> REJECTED). Requires approve permission.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
task_idYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it is a non-read-only, idempotent, non-destructive mutation. The description adds two pieces of behavioral context beyond that: the resulting state (REJECTED) and the required authorization ("Requires approve permission"). It does not explain whether the note is persisted or what happens to a task already rejected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with the action and state transition front-loaded and no wasted words. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter mutation with annotations present and no output schema, the action, state transition and permission requirement are covered. The complete absence of parameter documentation is the remaining gap, so it is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both parameters, so the description carries the full burden, yet it says nothing about task_id (the required identifier) or note. An agent gets no guidance on what task_id accepts or what note is for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (reject) and the resulting state transition (-> REJECTED), which lets an agent distinguish it from approve_task and review_task. It does not name the sibling tools explicitly, so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"At the approval gate" implies the trigger condition for use, which is useful context. However, it never states when not to use it or points to the alternative (approve_task) for the opposite decision, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

release_taskC
Idempotent

Release your own claim.

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYes
claim_tokenYes

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true, destructiveHint=false, and readOnlyHint=false, covering the safety profile. The description adds the meaningful constraint that only your own claim can be released, but says nothing about what happens to the task afterward or why a claim_token is required for proof of ownership.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler, but it is terse to the point of under-specification rather than cleanly scoped. Conciseness here is achieved by omission rather than by efficient phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a state-mutating tool (readOnlyHint=false) with two required params at 0% schema coverage and no output schema. The description omits the ownership requirement mechanics, the effect on task state, and error conditions, so it is not complete enough for confident invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description offers no meaning for either required parameter: task_id and claim_token are undocumented in both places. The description does not compensate for the coverage gap, so the agent must infer what claim_token represents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (release) and resource (claim) with the scoping qualifier 'your own', so an agent can distinguish it from claim_task and renew_claim. It is clear but thin, giving no detail about what releasing actually does to the task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no mention of alternatives, and no prerequisites. An agent cannot tell from the description whether releasing is preferable to complete_task, fail_task, or simply letting the claim expire.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

renew_claimC
Idempotent

Extend the lease.

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYes
claim_tokenYes

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (not read-only, idempotent, not destructive), so the description's job is to add nuance — e.g., what happens if the lease has already lapsed, or whether renewal resets the expiry window. It adds none of that; "Extend the lease" merely restates the mutation the annotations already imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words are not concision here, they are under-specification. There is no wasted sentence, but there is also no front-loaded payload beyond the verb, leaving the definition effectively empty.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating claim-lifecycle tool with two undocumented required parameters, no output schema, and a crowded sibling set, the description is far too thin. The annotations supply only the basic safety hints; the description never explains lease semantics, failure modes, or how it pairs with claim_task/release_task.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and both required parameters (task_id, claim_token) are undocumented. The description says nothing about them — it doesn't explain that claim_token identifies the lease being extended or that task_id scopes it, so an agent gets no help beyond the bare names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Extend the lease" gives a specific verb and an object, but the object ("lease") is never tied to the tool's actual resource, a claim on a task. Nothing distinguishes it from siblings like claim_task, release_task, or complete_task, which all operate on the same claim lifecycle.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative is named. An agent must infer from the name alone that this renews an existing claim's lease rather than acquiring (claim_task) or ending (release_task) one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resume_dispatcherA

Resume the dispatcher globally (ADMIN).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false, so the description carries real weight by disclosing the ADMIN permission requirement and the global (not per-agent) scope. It stops short of describing side effects such as whether paused tasks immediately resume or whether the change is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler; the action is front-loaded and the two qualifiers that matter (global scope, ADMIN) are appended compactly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter state-change action with an annotation covering its safety profile and no output schema, the description supplies the two missing essentials: privilege level and scope. Only the immediate effect on in-flight paused work is unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema baseline of 4 applies. There is nothing parameter-related the description could add.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Resume the dispatcher') plus an explicit scope ('globally'), so the action is unambiguous. It does not name pause_dispatcher, but the inverse pairing is obvious from the sibling names and the tool name itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the name and the existence of pause_dispatcher, but the description never states when to resume versus when not to, nor any precondition (e.g., dispatcher must currently be paused). Adequate but leaves inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_taskB
Idempotent

Decide the review gate (decision=accept|reject). Requires review permission; no self-review.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
reworkNo
task_idYes
decisionYes

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-read-only, idempotent, non-destructive, so the safety profile is covered. The description adds value beyond that by disclosing the review-permission requirement and the no-self-review constraint, both of which are real behavioral limits an agent must respect. It stops short of saying whether the decision is terminal or what a reject with rework does.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and its allowed values, with no filler. It is arguably too terse for a four-parameter mutation tool, but every clause carries information and nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-read-only gate tool with no output schema and 0% parameter documentation, the description covers permissions and self-review but omits the semantics of 'note' and 'rework' and any routing relative to approve_task/reject_task. An agent still cannot confidently choose this tool or populate all its arguments from the description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters, so the description must carry the burden and it largely does not. It does supply the accepted values for 'decision' (accept|reject), which the schema lacks as an enum, but 'note' and especially the boolean 'rework' are left completely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb 'Decide' plus the parenthetical 'decision=accept|reject' makes the action understandable, but 'the review gate' is jargon and the description never clarifies that this is a task-review decision. It also fails to distinguish the tool from the direct siblings approve_task and reject_task, which appear to make the same kind of gate decision.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Requires review permission; no self-review' gives real preconditions for use, which is more than nothing. However, there is no guidance on when to pick this tool over approve_task/reject_task, nor on when a review gate should be decided at all, leaving the alternative selection entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_knowledgeC
Read-onlyIdempotent

Search knowledge case-insensitively (ACL, optional space).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
spaceNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds only "case-insensitively," and the cryptic "ACL" fragment raises more questions than it answers (is ACL enforced silently, or must the caller hold it?). Nothing on pagination or result ordering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short and front-loaded, with no filler sentences. However, the dangling parenthetical reads like an unfinished fragment rather than structured guidance, so brevity comes at the cost of clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A 3-parameter search tool with no output schema and 0% schema coverage needs the description to carry matching semantics, scope, and result shape. The single fragment leaves required-vs-optional, ACL behavior, and limit semantics unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all three parameters. The description mentions "space" as optional (consistent with the schema) but adds no format, matching semantics, or guidance for "limit"; "ACL" is not a parameter at all. It fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb+resource ("Search knowledge") but "knowledge" is an ambiguous resource — it could overlap with list_documents/get_document. No sibling differentiation despite a crowded document-related sibling set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical "(ACL, optional space)" hints at scoping but never says when to prefer this over list_documents or get_document, nor what ACL means in practice (auto-filtered vs. required permission). No when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

write_documentC
Destructive

Write a knowledge document (git commit; ACL/secret guard).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
contentYes
commit_messageYes
expected_git_commitNo

TDQS

C2.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=true, so the mutation profile is covered. The description adds genuinely non-obvious behavior beyond that: writes are persisted as a git commit and pass through an ACL/secret guard. It stops short of stating overwrite semantics for an existing path, which would be the natural next detail for a destructive write.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the verb and resource first and qualifiers in a compact parenthetical. Nothing is wasted, though the telegraphic parenthetical borders on under-specification rather than tightness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive 4-parameter write tool with no output schema and 0% parameter documentation, the description omits too much: it never explains the parameters, whether an existing document is overwritten, or how expected_git_commit interacts with the git commit behavior. Annotations cover the safety signal but not the operational details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and none of the four parameters (path, content, commit_message, expected_git_commit) are explained in the description. The mention of 'git commit' does not clarify the expected_git_commit optimistic-concurrency parameter, so the description fails to compensate for the total absence of schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ('Write a knowledge document') and the parenthetical identifies the mechanism (git commit) and the guardrails (ACL/secret guard). It is distinguishable from append_document by the write-vs-append framing, though it never names that sibling or propose_write explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus append_document or propose_write, nor any stated preconditions. The 'ACL/secret guard' hint implies writes may be rejected, but no condition for choosing this tool is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 27 tool updatesv0.1.0
    • First observedappend_document
    • First observedapprove_task
    • First observedclaim_task
    • First observedcomplete_task
    • First observedcreate_task
    • First observeddisable_agent
    • First observedenable_agent
    • First observedfail_task
    • First observedget_agent_status
    • First observedget_document
    • First observedget_execution_job
    • First observedget_task
    • First observedlist_agents
    • First observedlist_documents
    • First observedlist_execution_jobs
    • First observedlist_runnable_tasks
    • First observedlist_spaces
    • First observedlist_tasks
    • First observedpause_dispatcher
    • First observedpropose_write
    • First observedreject_task
    • First observedrelease_task
    • First observedrenew_claim
    • First observedresume_dispatcher
    • First observedreview_task
    • First observedsearch_knowledge
    • First observedwrite_document

TDQS

B3/5.0

Scored across 27 tools

Disambiguation4/5

Each tool targets a distinct resource+action: the task lifecycle verbs (claim/complete/fail/approve/reject/review/release/renew) are well differentiated by their governance roles, and document tools (write/append/propose/search) are clearly separated. Only mild potential confusion exists between the various task gate tools, but descriptions clarify intent.

Naming Consistency5/5

Every tool follows a strict verb_noun snake_case pattern (list_agents, create_task, claim_task, write_document, pause_dispatcher), with predictable verbs for read/list/create/mutate operations. No convention mixing is present.

Tool Count3/5

27 tools is on the heavy side and exceeds the comfortable 3-15 range. However, the surface spans five legitimate subdomains (tasks, documents, agents, execution jobs, dispatcher), so the count is defensible rather than bloated.

Completeness4/5

Task and document lifecycles are well covered (create/claim/complete/fail/approve/reject/review plus read/write/append/search/propose). Execution jobs are read-only (get/list with no create or claim path) and no delete operations exist, which are minor workable gaps.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    The infrastructure for AI teams: a self-hosted server that gives a fleet of agents shared semantic memory, tasks, direct messages, and session handoff. Any agent that speaks HTTP participates: Claude Code, AutoGen, raw API scripts, anything.
    47
    9
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Local-first, file-based memory layer for AI agents — one shared Markdown vault across Claude, Codex, Gemini, Cursor and any MCP client. Provides read/write memory tools with an audit trail, per-agent trust levels, and Git sync; no cloud and no lock-in.
    2
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Git-native long-term memory for AI agents: your markdown files are the source of truth, the search index is a disposable projection rebuilt from git, and every memory the agent writes is a reviewable git commit. Served over one OAuth-secured MCP endpoint with hybrid lexical+semantic recall and a gated, git-first commit_note write tool.
    7
    9
    AGPL 3.0