Skip to main content
Glama

TaskFlow MCP

npm version

A local Model Context Protocol (MCP) server that gives AI agents structured task planning, execution tracking, and guided research workflows.

Quick Start • Client Setup • Tools • Documentation

Table of Contents 📌

Related MCP server: codeweave-mcp

Overview ✨

TaskFlow MCP helps agents turn vague goals into concrete, trackable work. It provides a persistent task system plus research and reasoning tools so agents can plan, execute, and verify tasks without re‑sending long context every time.

Why Use It ✅

  • Lower token use: retrieve structured task summaries instead of restating context.

  • Smarter workflows: dependency‑aware planning reduces rework.

  • Better handoffs: tasks, notes, and research state persist across sessions.

  • More reliable execution: schemas validate tool inputs.

  • Auditability: clear task history, verification, and scores.

How It Augments Modern AI Tools 🧭

TaskFlow MCP complements modern AI tooling. Tools like GitHub CLI and Skills help with repo workflows and onboarding, while TaskFlow MCP focuses on durable task state, structured planning/execution, and repeatable workflows across sessions. Use it to add persistent task memory and structured agent prompts on top of your existing toolchain.

What Is MCP? 🤔

MCP is a standard way for AI tools to call external capabilities over JSON‑RPC (usually STDIO). This server exposes tools that an agent can invoke to plan work, track progress, and keep context consistent across long sessions.

How TaskFlow Works 🧭

TaskFlow MCP adds a structured workflow layer on top of normal LLM chat. The server validates tool inputs and returns deterministic, structured prompts for planning and research, while persisting task state on disk so agents can resume without re‑sending long context.

flowchart LR
  subgraph Host["MCP Host: VS Code"]
    subgraph Client["MCP Client"]
      Agent["Agent / Model"]
    end
  end

  Agent -- "JSON-RPC (STDIO)" --> Server["MCP Server (taskflow)"]
  Server -- "Structured prompts / results" --> Agent
  Server --> Store["Data Store (DATA_DIR/.mcp-tasks)"]

In practice:

  • The host runs the MCP client and the model.

  • The client calls MCP tools over JSON‑RPC via STDIO.

  • The server validates inputs, builds structured prompts, and returns them to the client.

  • The data store keeps task state across sessions so the agent can resume without context loss.

Quick Start 🚀

pnpm install
pnpm build
pnpm start

Installation 📦

# npm
npm install

# yarn
yarn install

# pnpm
pnpm install

Basic Usage ▶️

Start the server

pnpm start

Configure data directory (optional)

# PowerShell
$env:DATA_DIR="${PWD}\.mcp-tasks"

Client Setup 📎

Use npx to run the MCP server directly from GitHub. Replace <DATA_DIR> with your preferred data path.

Path examples:

  • Windows: <DATA_DIR> = C:\repos\mcp-taskflow\.mcp-tasks

  • macOS/Linux: <DATA_DIR> = /Users/you/repos/mcp-taskflow/.mcp-tasks

VS Code (.vscode/mcp.json)

{
  "servers": {
    "mcp-taskflow": {
      "type": "stdio",
      "command": "npx",
      "args": ["mcp-taskflow"],
      "env": {
        "DATA_DIR": "<DATA_DIR>"
      }
    }
  }
}

Claude Desktop (settings JSON)

{
  "mcpServers": {
    "mcp-taskflow": {
      "command": "npx",
      "args": ["mcp-taskflow"],
      "env": { "DATA_DIR": "<DATA_DIR>" }
    }
  }
}

Codex (config.toml)

[mcp_servers.mcp-taskflow]
type = "stdio"
command = "npx"
args = ["mcp-taskflow"]
env = { DATA_DIR="<DATA_DIR>" }
startup_timeout_sec = 120

Tools Overview 🧰

TaskFlow MCP exposes a focused toolset. Most clients surface these as callable actions for your agent.

Planning

  • plan_task: turn a goal into a structured plan

  • split_tasks: split a plan into discrete tasks with dependencies

  • analyze_task: capture analysis and rationale

  • reflect_task: record reflections and improvements

Task Management

  • list_tasks: list tasks by status

  • get_task_detail: show full details for a task

  • query_task: search tasks by keyword or ID

  • create_task: create a task directly

  • update_task: update status, notes, dependencies, or metadata

  • delete_task: remove a task by ID

  • clear_all_tasks: clear the task list

Workflow

  • execute_task: mark a task in progress and generate an execution prompt

  • verify_task: score and mark a task complete

Research & Project

  • research_mode: guided research with state tracking

  • process_thought: capture a structured reasoning step

  • init_project_rules: create or refresh project rules

  • get_server_info: get server status and task counts

Example: Agent-in-the-Loop (ReBAC Feature) 🧪

Below is a simple, human‑readable script that shows how a user might ask an agent to plan and execute a feature. The agent uses TaskFlow MCP tools behind the scenes, but you don’t need MCP details to follow the flow.

Plain Chat vs TaskFlow (ReBAC Example)

Without TaskFlow (plain chat) Prompt:

Create a ReBAC system from scratch.

Typical outcome:

  • The model returns a large, one‑shot answer.

  • No durable task list or dependencies.

  • Hard to resume later without re‑explaining context.

  • Team members have no shared, structured view of progress.

With TaskFlow (structured workflow) Prompt:

Create a ReBAC system from scratch. Plan the work, split tasks, then execute and verify.

Typical outcome:

  • The model generates a plan via plan_task.

  • Tasks are created and tracked via split_tasks (with dependencies).

  • Each task is executed and marked in progress via execute_task.

  • Results are verified and scored via verify_task, with adjustments logged.

  • State is persisted in the datastore, so anyone can list_tasks and get_task_detail to continue or review.

Why this matters for teams

  • The task list, notes, and verification results are stored on disk and can be shared in the repo or a shared data directory.

  • A teammate can open the same workspace and immediately see the current task state without reading long chat history.

Dependency management

  • Tasks can declare explicit prerequisites, so the agent knows what must happen first.

  • Dependencies prevent blocked work: a task can’t be executed until its upstream tasks are complete.

  • Dependencies are stored with tasks, so any teammate can see the critical path and pick up the next unblocked item.

  • For example: “Integrate with existing auth” cannot start until both “Define ReBAC model” and “Design storage layer” are completed.

User “I want to add a Relationship‑Based system. Create a task list and start working through it.”

Agent “I’ll plan the work, split it into tasks, and track progress as we go.”

Agent (planning summary) “Here’s a structured plan with dependencies and a clear execution order.”

  1. Define the ReBAC model

    • Entities, relations, and permission checks

    • Identify critical query patterns

  2. Design the storage layer

    • Graph storage schema and indexes

    • Query strategy for path checks

  3. Implement the policy engine

    • Relationship evaluation

    • Cached query results

  4. Integrate with existing auth

    • Request context mapping

    • Backward compatibility

  5. Add tests and benchmarks

    • Unit tests for rules

    • Integration tests for policy decisions

Agent (starts work) “I’ll mark the first task as in progress and add notes as I go.”

Progress updates

  • Task 1: In progress — “Drafted entity/relationship schema and example checks”

  • Task 1: Completed — “Added model doc and validation rules”

  • Task 2: In progress — “Evaluating graph storage options”

Task verification example (with scoring and challenges) Agent “I’ve verified Task 1 and logged a score.”

  • Score: 92/100

  • Checks passed: model completeness, schema validation, examples included

  • Challenges: ambiguous relationship naming in legacy data; resolved by adding a normalization step and a short mapping table

  • Next step: start Task 2 with the normalized model in place

Why this helps

  • The agent keeps a durable task list and status updates.

  • You can stop and resume without losing context.

  • Large features become manageable, with explicit dependencies.

Documentation 📚

Document

Purpose

docs/API.md

Tool overview and API surface

docs/ARCHITECTURE.md

High-level architecture

docs/PERFORMANCE.md

Benchmarks and performance targets

AI_AGENT_QUICK_REFERENCE.md

Agent workflow reference

SECURITY.md

Threat model and controls

CONTRIBUTING.md

Contribution workflow and changesets

CHANGELOG.md

Release notes

Development 🛠️

pnpm test
pnpm type-check
pnpm lint

Versioning 🏷️

This project uses Changesets for versioning and release notes. See CONTRIBUTING.md for guidance.

Release and Git-Based Usage 🚢

Git-based execution assumes the repository is buildable and includes a valid bin entry in package.json. For production or shared use, prefer a tagged release published via Changesets.

Typical flow:

  1. Add a changeset in your PR.

  2. CI creates a release PR with version bumps and changelog entries.

  3. Merging the release PR publishes to npm and creates a GitHub release.

Use git-based execution for fast testing; use npm releases for stable installs.

# pnpm
pnpm dlx git+https://github.com/CalebGerman/mcp-taskflow.git mcp-taskflow

# npx (fallback)
npx git+https://github.com/CalebGerman/mcp-taskflow.git mcp-taskflow

Prerequisites:

  • bin entry points to dist/index.js

  • pnpm build completes successfully

License 📄

MIT. See LICENSE.md.

Credit 🙏

Inspired by:

https://github.com/cjo4m06/mcp-shrimp-task-manager

Also informed by related MCP server patterns and workflows:

https://www.nuget.org/packages/Mcp.TaskAndResearch

Available Tools

16 tools
analyze_taskC

Deeply analyze task requirements and propose a high-level approach. Includes threat modeling, design decisions, and architectural considerations.

ParametersJSON Schema
NameRequiredDescriptionDefault
summaryYes
initialConceptYes
previousAnalysisNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it only lists output content (threat modeling, design decisions, architectural considerations). It never says whether the analysis is persisted, whether it mutates task state, what permissions are needed, or how it interacts with the task lifecycle in this server.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and followed by a scoping list of what the analysis covers. Little waste, though the enumerated 'includes' clause is somewhat generic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three undocumented parameters, no annotations, and no output schema, the description should explain inputs and the nature/return of the analysis. It omits both, leaving an agent unable to construct a correct call with confidence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description mentions none of the three parameters. The required 'summary' and 'initialConcept' and the optional 'previousAnalysis' are left entirely opaque, with no hint of expected content, format, or how previousAnalysis is used to continue an iteration.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('analyze') and resource ('task requirements') and states the deliverable ('propose a high-level approach'), which is more than a tautology. It does not, however, differentiate this from close siblings like plan_task or reflect_task, so an agent cannot tell from the text alone which analysis-flavored tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus the many overlapping siblings (plan_task, reflect_task, split_tasks, research_mode). No prerequisites, no exclusions, no lifecycle position (e.g., before planning) are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clear_all_tasksB

Clear all tasks after confirmation. Creates backup before deletion.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses destructive intent plus a mitigating behavior (a backup is created before deletion) and a confirmation requirement — real value beyond structured data. It omits reversibility/restore details, backup location, and permission requirements, so the disclosure is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler, and the destructive action is front-loaded before the safety behavior. Efficient, though each sentence is quite terse relative to what an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive bulk operation with no annotations and no output schema, the description covers the essentials (confirmation, backup) but leaves open the blast radius (all tasks across all lists/projects?), what the backup contains, and what the caller receives back. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for the single 'confirm' boolean, so the description must carry the meaning. 'After confirmation' hints that confirm gates the destructive action, but it never states that confirm defaults to false or whether omitting it aborts the operation. Partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource with scope ('clear all tasks'), which is clearly distinguishable from the sibling delete_task (singular). The second sentence adds the backup behavior. It stops short of explicitly naming the single-task alternative, but the bulk scope is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'After confirmation' implies a prerequisite gate for invocation, and the bulk scope implies this is for wiping everything at once. However, there is no explicit when-to-use vs. delete_task guidance and no statement of when NOT to use it (e.g., partial cleanup).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_taskB

Delete a task. Completed tasks are archived, others are removed.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose a genuinely useful behavioral trait — completed tasks are archived while others are permanently removed — which is exactly the kind of context a caller needs to avoid accidental data loss. However, it says nothing about whether the deletion is reversible, what permissions are required, or what happens to subtasks/dependents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and immediately followed by the one behavioral fact that matters. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive single-id operation with no annotations and no output schema, the description covers the core case but omits irreversibility, permission requirements, and error/not-found behavior. It is adequate but leaves real gaps for a delete tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has one parameter (taskId, required, uuid) with 0% schema description coverage, so the description must compensate. It mentions no parameter at all — not which task identifier to pass, nor what happens if the id is unknown. Only the intuitive name 'taskId' makes this callable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Delete a task') and adds the key behavioral distinction that completed tasks are archived whereas incomplete ones are removed. It is clear what the tool does, though it never names or contrasts with the obvious sibling clear_all_tasks, so an agent gets no help choosing between them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no exclusions. Given siblings like clear_all_tasks (bulk removal) and update_task, the description never says whether this is the single-task path or when clearing all tasks is preferable. Usage must be inferred entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_taskB

Mark a task as in progress and build an execution prompt with dependencies and implementation guide.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions state mutation (marking in progress) and prompt construction, but does not disclose critical side effects: whether the task status change is permanent, what happens to existing prompts or dependencies, whether this requires the task to be in a specific state, or what the generated prompt contains. For a mutation tool with zero annotation coverage, this is inadequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, tightly written sentence that front-loads the primary action and then details the secondary output. No wasted words, and the compound purpose is clear without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter mutation tool with no annotations and no output schema, the description covers the basic what but omits behavioral details (state transitions, side effects, permissions) and parameter semantics. It is minimally viable but has clear gaps that could lead to incorrect invocations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter 'taskId' is undocumented in both schema and description. The description implies the tool operates on a specific task but provides no additional parameter details, format expectations, or constraints beyond what is structurally visible. It fails to compensate for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific compound action: 'Mark a task as in progress' and 'build an execution prompt with dependencies and implementation guide.' It clearly distinguishes itself from siblings like plan_task, verify_task, or reflect_task by its unique execution-prompt-building purpose. However, it lacks a direct sibling reference or statement of what it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: call this when a task is ready to start executing. There's no explicit when-to-use/when-not-to-use guidance nor mention of alternative tools like plan_task or verify_task that might precede or follow this step. The description leaves the agent to infer the workflow position.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_server_infoB

Get server metadata, version, and health status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With zero annotations and no output schema, the description carries the full burden of behavioral disclosure. It does not state whether this requires authentication, what the health status format looks like, or any side effects. A server info check is likely read-only, but this is not stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single concise sentence with zero waste, front-loading the verb and resource.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read tool, the description is adequate but leaves behavioral gaps (e.g., auth, output format) that would be covered by annotations or an output schema, neither of which exist.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, so the baseline is 4. The description correctly implies no configuration is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('server metadata, version, and health status'). The purpose is unmistakable and clearly distinct from the task-management siblings like reflect_task or list_tasks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance or context. The description does not indicate when an agent should call this versus other tools, though the purpose is self-evident.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_task_detailB

Get complete details for a specific task including dependencies and related files.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the return scope (dependencies and related files), and 'Get' implies a read-only operation, but it says nothing about permissions, error behavior for missing/invalid IDs, or response format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the core action front-loaded and no wasted words. It is appropriately sized for a one-parameter read tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-ID lookup with no output schema, the description covers the essentials by naming what the response includes (dependencies, related files). The main gap is the absence of any note about error handling or access requirements.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter at 0% schema description coverage, so the description must compensate. 'A specific task' confirms taskId identifies a single task, but adds no format or sourcing detail beyond the self-documenting name and uuid type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (details for a specific task) plus the scope of returned content. It does not, however, distinguish itself from siblings like query_task or list_tasks, which could plausibly retrieve task information.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance, and no sibling alternatives are named. The only hint is that the tool targets 'a specific task', implying a single-ID lookup, but this is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

init_project_rulesB

Initialize or update project coding rules and guidelines.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing. It does not say whether existing rules are overwritten, where the rules are stored, whether permissions are required, or whether the operation is idempotent. For a mutation tool with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with no filler, and the action is front-loaded. It is arguably too terse for a mutation tool, but there is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and no output schema, the structural surface is small, so the description need not explain return values. However, it is a state-changing operation with no annotations, and the description never explains what actually gets created or modified, leaving the agent under-informed about effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema declares zero parameters, so there is nothing for the description to disambiguate; the baseline for a no-parameter tool is 4. The description adds no parameter detail, but none is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States specific verbs (initialize, update) and a concrete resource (project coding rules and guidelines), so the agent knows the general outcome. It lacks sibling differentiation, though none of the listed task/thought siblings overlap with this resource, so that gap is minor. The dual 'initialize or update' phrasing leaves the actual mode ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus alternatives, no prerequisites, and no explanation of how the tool decides between initializing and updating. The agent must infer the trigger condition entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tasksC

List tasks by status. Returns task overview with counts and filtering.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNoall

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read operation and hints at an aggregate view ("counts"), but never states read-only safety, whether results are paginated or limited, or how counts scope to the filter — significant gaps for a listing tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the operation, with no filler. The second sentence is slightly vague ("overview with counts and filtering") but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-optional-param tool with no output schema or annotations, the description is minimally adequate but leaves the return shape undefined beyond "counts," and gives no pagination, ordering, or result-size expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but there is only one parameter and its enum values (all/pending/in_progress/completed/blocked) are largely self-explanatory. The description adds only "by status" and does not clarify what "all" means or how filtering interacts with the returned counts.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("List tasks") with a scope qualifier ("by status"), which is clearer than a bare name restatement. It does not distinguish itself from close siblings such as query_task or get_task_detail, so an agent has to guess which list/query tool applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to use this over query_task, get_task_detail, or the other task-inspection siblings, and no prerequisites or exclusions. The only guidance is the implicit "when you want a list," which the description never actually states.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plan_taskC

Plan tasks and construct a structured task list. Converts natural language descriptions into actionable task proposals with goals and expected outcomes.

ParametersJSON Schema
NameRequiredDescriptionDefault
descriptionYes
requirementsNo
existingTasksReferenceNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not say whether planning persists or mutates state, whether it is safe to re-run, what happens to existing tasks, or what the agent gets back — the existingTasksReference parameter hints at state interaction but the description never addresses it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler, front-loading the core action before the elaboration. Efficient, though the second sentence is somewhat redundant with the first rather than adding new operational detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no annotations, no output schema, and zero schema description coverage, the description is too thin. It should at minimum explain the parameters, whether the result is persisted, and how it relates to the execute/verify lifecycle.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 3 parameters. The phrase 'natural language descriptions' loosely maps to the required 'description' field, but 'requirements' and 'existingTasksReference' are undocumented in both schema and description, leaving a real gap the description should have closed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Plan tasks and construct a structured task list') and elaborates with the transformation it performs ('converts natural language descriptions into actionable task proposals with goals and expected outcomes'). The purpose is clear, but it never distinguishes itself from adjacent siblings like split_tasks or analyze_task, so an agent gets no help disambiguating within this crowded family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, prerequisites, or alternatives are given. With 15 sibling tools including split_tasks, analyze_task, and reflect_task, the absence of routing guidance leaves the agent to guess which planning-stage tool applies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_thoughtC

Record a structured thought step in a reasoning process with stage tracking.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNo
stageYes
thoughtYes
axiomsUsedNo
thoughtNumberYes
totalThoughtsYes
nextThoughtNeededNo
assumptionsChallengedNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies a stateful write ("record") but says nothing about whether a session must be initialized, whether thought numbers must be sequential, whether prior thoughts are overwritten or appended, or whether this state persists across calls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with no waste and the core action front-loaded. But for an 8-parameter stateful tool, the brevity reads as under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, four required parameters, and 0% schema description coverage, the one-sentence description leaves an agent unable to determine sequencing rules, required context, or what a call returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 8 parameters, so the schema offers no help. The description only gestures at one of them ("stage tracking"); it adds no meaning for thought, thoughtNumber, totalThoughts, nextThoughtNeeded, tags, axiomsUsed, or assumptionsChallenged, leaving most of the tool's semantics unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Record a structured thought step in a reasoning process with stage tracking" gives a verb and a resource, so it exceeds a pure restatement of the name. However, "reasoning process" and "thought step" are vague and the description does nothing to separate this from the task-oriented siblings (reflect_task, plan_task, analyze_task), which also manipulate reasoning artifacts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling it replaces. Given twelve task-related siblings, the absence of any routing guidance is a real gap rather than a minor omission.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_taskC

Search tasks by keyword or ID with pagination support.

ParametersJSON Schema
NameRequiredDescriptionDefault
isIdNo
pageNo
queryYes
pageSizeNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It mentions pagination support but does not describe return format, permissions, case sensitivity, or any other runtime behavior beyond the bare capability.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It communicates the core operation and two key capabilities without redundant phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, no annotations, and four undocumented parameters, the description is too thin for a tool with several retrieval siblings. It omits return structure, pagination behavior, and any differentiation from list_tasks/get_task_detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all four parameters. It only implies keyword-vs-ID and pagination; it does not explain the isId boolean, query length limits, page defaults, or pageSize bounds.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Search') and resource ('tasks'), and scopes it to keyword/ID plus pagination. It does not explicitly differentiate from siblings such as list_tasks or get_task_detail, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this search tool versus list_tasks, get_task_detail, or other task retrieval siblings. The description only states the capability, not the context or exclusions for its use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reflect_taskC

Critically review analysis results and propose optimizations. Identifies potential issues, alternative approaches, and improvements.

ParametersJSON Schema
NameRequiredDescriptionDefault
summaryYes
analysisYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It says the tool identifies issues and improvements, but does not disclose whether it is read-only, whether proposals are applied, what permissions are needed, or what it returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences, front-loaded with the core action and then the specific outputs. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two required parameters, no annotations, no output schema, and no parameter descriptions, this description is incomplete. It gives a high-level purpose but omits usage context, parameter meaning, and operational behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the two parameters. It only vaguely references 'analysis results' and says nothing about the required 'summary' parameter or expected formats and limits.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Critically review analysis results and propose optimizations.' It is clear what the tool does, but it does not mention task context or distinguish itself from siblings like analyze_task or verify_task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. The implied sequence after analysis is weak, and no exclusions or sibling routing are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

research_modeC

Enter research mode to explore a programming topic in depth with iterative state tracking.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicYes
nextStepsYes
currentStateYes
previousStateYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It implies a stateful mode change via 'enter research mode' and 'iterative state tracking,' but does not disclose persistence, side effects, required state format, auth, or what happens to previous state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler or redundancy. It is concise, though its brevity leaves no room for the richer guidance a stateful four-parameter tool needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given four required parameters, 0% schema description coverage, no annotations, and no output schema, the description is far from complete enough to invoke the tool correctly. It communicates the general purpose but omits parameter semantics and behavioral specifics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for four required parameters. The description vaguely maps 'programming topic' to topic and 'state tracking' to previousState/currentState/nextSteps, but it does not explain what each parameter should contain or how they relate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action (enter research mode) and resource/purpose (explore a programming topic in depth with iterative state tracking). It is clear what the tool is for, but it does not explicitly differentiate from siblings such as reflect_task, plan_task, or process_thought.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: use this tool to explore a programming topic in depth with iterative state tracking. However, the description gives no explicit when-to-use guidance, no exclusions, and no alternatives among the sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

split_tasksC

Split a complex task into a structured set of smaller tasks. Supports bulk creation with dependency graphs and various update modes.

ParametersJSON Schema
NameRequiredDescriptionDefault
tasksYes
updateModeYes
globalAnalysisResultNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it omits the most important trait: the updateMode enum includes destructive values ('overwrite', 'clearAllTasks') that can replace or wipe existing tasks. Bulk creation of up to 100 tasks and dependency-graph handling are mentioned, but the destructive modes and their consequences are never disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core purpose front-loaded and no filler. It is efficient, though the brevity is partly a symptom of missing detail rather than disciplined editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a complex tool with nested task objects, dependency arrays, and a destructive update-mode enum, plus no annotations and no output schema. A two-sentence description leaves the agent without the mode semantics or safety context it needs to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the parameters, yet it only alludes to 'dependency graphs' and 'various update modes'. The four enum values for updateMode (append, overwrite, selective, clearAllTasks) are never explained, and the per-task fields (agent, notes, relatedFiles, verificationCriteria, implementationGuide) go entirely unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Split a complex task into a structured set of smaller tasks.' That is clear enough for an agent to understand the operation. It does not, however, distinguish this tool from nearby siblings like plan_task or analyze_task, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use split_tasks versus plan_task, analyze_task, or update_task, and no prerequisites or exclusions are stated. The phrase 'various update modes' gestures at behavior but never tells the agent which mode to pick or when.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_taskC

Update task details, dependencies, or related files.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
notesNo
taskIdYes
descriptionNo
dependenciesNo
relatedFilesNo
implementationGuideNo
verificationCriteriaNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the entire behavioral burden. It implies mutation but never says whether this is a partial/patch update or a full replace (critical with 8 optional fields), whether changes are reversible, what permissions are needed, or how dependencies/files arrays are merged rather than overwritten.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with no filler and the core action front-loaded. It is efficient, though it errs toward under-specification rather than being genuinely well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An 8-parameter mutation tool with no annotations, no output schema, and 0% schema description coverage needs substantially more from the description, and it delivers almost none of it. An agent has no way to know required permissions, merge semantics, or how the complex relatedFiles array should be populated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 8 parameters, so the schema offers no semantics at all. The description names dependencies and related files but omits name, description, notes, implementationGuide, and verificationCriteria, leaving most parameters undocumented by either source.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Update) and resource (task) and enumerates the categories it can change: details, dependencies, related files. However, it does nothing to distinguish this from the many sibling mutation tools (delete_task, split_tasks, execute_task, verify_task), so an agent must still infer where it fits.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives despite a crowded sibling set containing delete_task, split_tasks, and verify_task. The agent gets no signal about which task state or workflow stage this tool is appropriate for.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_taskC

Verify task completion with a score and summary. Marks task as completed.

ParametersJSON Schema
NameRequiredDescriptionDefault
scoreYes
taskIdYes
summaryYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the key side effect ('Marks task as completed'), but says nothing about whether this state change is reversible, whether it requires the task to already be in a particular state, or whether repeated verification overwrites a prior score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero padding, with the primary action front-loaded before the side effect. It is appropriately sized, though being that terse is part of why behavioral detail is missing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-required-parameter mutation tool with no annotations and no output schema, the description leaves too much open: the meaning of the score, the state transitions implied by completion, and any auth or preconditions. It is not adequate for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and does not. It names 'score' and 'summary' but never explains the 0–100 scale, what a passing score means, what the summary should contain, or that taskId must be a UUID identifying the task to verify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Verify task completion' plus the side effect 'Marks task as completed.' An agent can distinguish it from read siblings like get_task_detail or list_tasks, though it is not explicitly differentiated from update_task, which could plausibly perform the same state change.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites (e.g., must the task be in-progress or executed first?), and never names an alternative such as update_task for less formal status changes. The agent must infer usage entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 16 tool updatesv0.1.3
    • First observedanalyze_task
    • First observedclear_all_tasks
    • First observeddelete_task
    • First observedexecute_task
    • First observedget_server_info
    • First observedget_task_detail
    • First observedinit_project_rules
    • First observedlist_tasks
    • First observedplan_task
    • First observedprocess_thought
    • First observedquery_task
    • First observedreflect_task
    • First observedresearch_mode
    • First observedsplit_tasks
    • First observedupdate_task
    • First observedverify_task

TDQS

B3/5.0

Scored across 16 tools

Disambiguation3/5

The CRUD-oriented task tools (delete_task, clear_all_tasks, list_tasks, get_task_detail, query_task, update_task, execute_task, verify_task) are reasonably distinct, but the cognitive cluster—reflect_task, analyze_task, plan_task, process_thought, research_mode—overlaps heavily and an agent could easily pick the wrong reasoning tool. list_tasks vs query_task also blur slightly.

Naming Consistency4/5

Strong verb_noun snake_case pattern throughout (reflect_task, plan_task, delete_task, split_tasks, update_task, execute_task, verify_task). Minor deviations: research_mode is noun-style rather than verb-led, and get_task_detail adds a third token, but overall consistency is high.

Tool Count4/5

16 tools is just above the comfortable range but each covers a plausible distinct operation within a task-management workflow. Slightly heavy, with some reasoning tools that could arguably be consolidated.

Completeness4/5

Good lifecycle coverage: task creation (plan_task/split_tasks), read (list_tasks, get_task_detail, query_task), update, delete, execute, and verify are all present, plus project rules and server info. Minor gaps around explicit single-task creation semantics and no dedicated dependency-graph query beyond get_task_detail.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers