idea-base-mcp-server
OfficialManage IDEA Base projects, tasks, time tracking, products, and AI verification through MCP tools.
List, create, update, and get projects and sub-projects.
List, search, get, create, update, and change status of tasks; create subtasks and manage dependencies.
Log time, run timed work sessions (start/pause/resume/stop) with pause reasons that affect worked time.
Add work notes, comments, and resume context to tasks.
List, create, and link products to projects.
Assign or remove task assignees.
Run AI verification against PR diffs and record feedback.
Quickly log completed work with quick_log.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@idea-base-mcp-serverList my projects"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@idea-base/mcp-server
MCP (Model Context Protocol) server for IDEA Base — AI-powered project management. Manage projects, tasks, and time tracking directly from Claude Code, Cursor, or any MCP-compatible AI tool.
Quick Setup
1. Get your API key
Sign in to IDEA Base, go to Settings > API Keys, and create a key.
2. Add to your MCP config
Claude Code (~/.claude/claude_desktop_config.json):
{
"mcpServers": {
"idea-base": {
"command": "npx",
"args": ["-y", "@idea-base/mcp-server"],
"env": {
"IDEA_BASE_API_KEY": "ib_your_api_key_here"
}
}
}
}Cursor (.cursor/mcp.json in your project):
{
"mcpServers": {
"idea-base": {
"command": "npx",
"args": ["-y", "@idea-base/mcp-server"],
"env": {
"IDEA_BASE_API_KEY": "ib_your_api_key_here"
}
}
}
}Or via Claude Code CLI:
claude mcp add idea-base -- npx -y @idea-base/mcp-server \
--env IDEA_BASE_API_KEY=ib_your_api_key_here3. Start using it
Ask Claude to manage your projects:
"List my projects"
"Create a task in project 1: Implement login page"
"Log 2 hours on task 42 — built the auth flow"
"What tasks are in progress?"
Related MCP server: wootech-jira-mcp
Available Tools
Projects
Tool | Description |
| List all projects with task counts and progress |
| Get project details and statistics |
| Create a new project or sub-project |
| Update project name, description, or status |
Tasks
Tool | Description |
| List tasks for a project (filter by status). Compact rows by default ( |
| Get task details, acceptance criteria, and time entries |
| Create a task with title, description, estimate, priority, start/due dates, assignee; |
| Update task details, priority, start/due dates, assignee; |
| Change task status (todo/in_progress/blocked/done) |
| Search tasks across all projects, ranked title-first then description then recency. Compact rows by default; filter by |
| Create + complete + log time in one step |
Time Tracking
Tool | Description |
| Log time against a task with notes |
| Open a timed work session on a task, and mark yourself actively working (surfaces the saved resume context + recent work notes so a cold session re-orients). Calling it twice never opens a second session — it resumes or reports the one you have |
| Pause the open session without ending it, with a required |
| Close the pause and carry on in the same session |
| Close the session with a UTC end instant and drop the active flag (optionally capture a |
Sessions are measured, and the pause reason decides the number. start_working
records a UTC start instant, stop_working a UTC end instant, and each pause in
between records both its boundaries plus why work stopped:
| Effect on worked time |
| Excluded. A person has to act before you can continue |
| Counted. A sub-agent blocked on another active sub-agent is still working |
| Counted |
Only waiting on the human is not working. Use pause_working rather than
stop_working whenever you intend to carry on — stop_working ends the session
and the reason is lost.
Activity & Audit Trail
For AI agents, these leave a durable trail of what was done and why on each task — so a future session (or a human reviewer) can see the reasoning, not just the final state.
Tool | Description |
| Append a timestamped progress note to a task's activity log (append-only journal) |
| Add a comment to a task's discussion thread (customer-visible by default) |
| Overwrite the task's pinned "where I left off" block, read first on |
Products
Tool | Description |
| List products (top-level containers) |
| Get product details with linked projects |
| Create a new product |
| Link a project to a product |
Environment Variables
Variable | Required | Description |
| Yes | Your API key from Settings > API Keys |
| No | Custom API URL (default: |
Real-time Notifications
When you use the MCP server to update tasks or log time, changes are broadcast to all connected users. Team members viewing the project in their browser see live toast notifications — when Claude updates a task, everyone sees it immediately.
Security
All data access is scoped to your account via API key
API keys support read/write permissions
No data is stored locally — all operations go through the IDEA Base API
Cross-account access is blocked server-side
Rate limited per API key
Development
# Run the server directly
IDEA_BASE_API_KEY=your_key npm start
# Watch mode
IDEA_BASE_API_KEY=your_key npm run devLicense
MIT - IDEA Management LLC
Available Tools
27 toolsadd_assigneeA
Assign a user to a task WITHOUT disturbing anyone already assigned. task_assignments holds one row per (task_id, user_id) — POST /api/tasks/:id/assignments upserts exactly that one row and leaves every other assignee alone, unlike update_task's assignee_user_id (which REPLACES the whole set — use that only when you actually want a single owner). This is the fix for the defect where a second MCP assignment silently dropped the first.
Idempotent and is_active-safe: this checks the task's current assignees first, and if the user already holds a row (assigned OR actively working) it does nothing and says so, rather than re-POSTing. That matters because the REST endpoint's upsert OVERWRITES is_active on conflict — a blind re-assign of someone with a running timer (is_active=1, set by start_working) would silently stop their clock. Skipping the no-op re-POST is what keeps assignment and active-work presence independent.
Rejects a user who is not a member of this account (404, surfaced as an error) rather than creating a dangling assignment.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | The ID of the task to assign the user to. | |
| user_id | Yes | User ID to add as an assignee. Must be a member of the same account. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses idempotency, the is_active overwrite hazard on conflict, that a blind re-POST would stop a running timer, and the 404 rejection path for non-members. That is exactly the mutation/auth/side-effect context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the decisive contrast (additive vs replacing) in the first sentence. The remaining explanation of the is_active hazard and the historical defect is relevant but somewhat repetitive ('silently dropped'/'silently stop their clock'), costing a little density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param mutation with no annotations and no output schema, the description covers success semantics, no-op semantics, error semantics (404), and side-effect boundaries. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented in the schema. The description adds the composite-key model (one row per task_id/user_id) but the account-membership requirement it repeats is already in the user_id schema description. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Assign a user to a task') and immediately scopes it against the sibling update_task's assignee_user_id, so an agent can distinguish the additive semantics from the replacing one without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative (update_task's assignee_user_id) and the condition that selects it ('use that only when you actually want a single owner'), plus the when-not for this tool's re-POST behavior. Routing is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_commentA
Add a comment to a task's discussion thread. Unlike work notes (progress journal), comments are for communication and are customer-visible by default.
| Name | Required | Description | Default |
|---|---|---|---|
| comment | Yes | The comment text | |
| task_id | Yes | The ID of the task to comment on | |
| is_internal | No | Whether the comment is internal-only (hidden from customers). Defaults to false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose the key non-obvious trait: comments are customer-visible by default, which is the main behavioral risk. It does not cover notification behavior, editability, or whether the comment can be removed later, so it is strong but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, with the core action front-loaded and the disambiguation immediately following. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a small three-parameter tool with no output schema, the description covers the action, the target resource, and the crucial audience distinction. It leaves minor gaps around side effects (notifications) and what the call returns, but nothing critical to invoking it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents task_id, comment, and is_internal (including its default). The description's visibility remark reinforces the is_internal parameter but adds no syntax, format, or length constraints beyond what the schema provides, matching the baseline for fully-documented schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Add a comment to a task's discussion thread') and explicitly distinguishes itself from the sibling concept of work notes. An agent can tell it apart from add_work_note without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Draws a clear line against alternatives by contrasting comments (communication) with work notes (progress journal), which routes the agent correctly. It stops short of an explicit when-not-to-use rule or naming the add_work_note tool directly, so it is clear context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_work_noteA
Append a timestamped work note (status update) to a task's activity log. Use this to record progress so a future session can pick up where you left off (e.g. "finished auth handler, tests green, next: wire the callback"). Notes are append-only and never overwrite each other.
| Name | Required | Description | Default |
|---|---|---|---|
| note | Yes | The work note / status update text | |
| task_id | Yes | The ID of the task to add a work note to | |
| is_internal | No | Whether the note is internal-only (hidden from customers). Defaults to true for work notes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that notes are append-only and never overwrite each other, that entries are timestamped, and that they land in the task's activity log. It does not state permission requirements or what the call returns, which are the remaining gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action, then the motivating use case, then the immutability guarantee. The parenthetical example earns its place by showing the expected note shape; no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter append tool with no output schema and no annotations, the description covers what it does, when to use it, and its key behavioral guarantee. It could have noted whether the writer is attributed or whether the note is visible to customers by default beyond the is_internal hint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so task_id, note, and is_internal are already documented, including the internal-only default. The description adds no further syntax, format, or length constraints on the note text, so it merely matches the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (append) and resource (timestamped work note on a task's activity log), and the parenthetical '(status update)' plus 'future session can pick up' clearly separates it from a plain comment. It never names the sibling add_comment however, so differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use scenario ('record progress so a future session can pick up where you left off') with a concrete example. It does not mention when NOT to use it or point at alternatives like add_comment or set_resume_context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_productA
Create a new product. Products are top-level containers for projects. Every product belongs to a customer (customer_id is required) — use an internal/own-company customer for internal work. Find customer ids via list_products (each shows its customer_id).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Name of the product | |
| customer_id | Yes | Customer this product is for. REQUIRED — every product belongs to a customer. | |
| description | No | Description of the product | |
| github_repo_url | No | GitHub repository URL for the product |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the domain model and a required precondition (customer_id), but does not describe return value, side effects, permissions, or error behavior. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three focused sentences. It front-loads the action ('Create a new product') and adds only relevant conceptual and lookup guidance without unnecessary filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool with 4 parameters, complete schema descriptions, and no output schema, the description provides sufficient context about the product–customer relationship and required fields. It would benefit from mentioning return behavior or error cases, but these are not critical given the schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, providing the baseline of 3. The description adds practical semantic value by explaining that every product belongs to a customer, customer_id is required, and how to find customer IDs via list_products, going beyond the schema's field labels.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool creates a new product, and adds context that products are top-level containers for projects, which distinguishes it from sibling tools like create_project. The verb-resource pair is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use this tool (creating a top-level container) and provides explicit guidance on required customer_id, including using an internal/own-company customer for internal work. It also directs users to list_products for finding customer IDs, though it does not explicitly contrast with create_project.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_projectA
Create a new project. A TOP-LEVEL project requires product_id; a sub-project (parent_project_id set) inherits its parent's product. Returns the created project with its ID.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Name of the project | |
| product_id | No | Product this project belongs to. Required for a top-level project (no parent); omit for a sub-project. Find ids via list_products. | |
| description | No | Description of the project | |
| parent_project_id | No | Parent project ID if this is a sub-project (inherits the parent's product) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that a new project is created, states the inheritance behavior for sub-projects, and mentions the return value. It doesn't cover permission requirements or error conditions, but the core behavioral traits of a create operation are adequately disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with the primary action front-loaded in the first sentence. The second sentence packs crucial conditional logic without waste. Every phrase earns its place, maintaining clarity without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a creation tool with no output schema, the description is complete: it covers required vs optional parameters, the top-level/sub-project distinction, return value, and references sibling tool list_products for ID lookup. Minor gaps like explicit conflict handling (e.g., setting both product_id and parent_project_id) are implied but not stated, but overall it suffices for an agent to act correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with each parameter documented, but the description adds cross-parameter logic: product_id is required for top-level, omitted for sub-projects, and inheritance from parent_project_id. This goes beyond individual parameter descriptions and clarifies the relationship, adding significant value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Create a new project' with a specific verb and resource. It distinguishes from sibling tools by focusing on project creation and provides meaningful detail about top-level vs sub-project semantics, making its purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the tool and how to configure it: top-level projects require product_id, sub-projects inherit from parent and omit product_id. It also references list_products for finding IDs. While it doesn't explicitly mention alternatives like update_project, the context is clear enough for an agent to decide when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_taskB
Create a new task in a project. Returns the created task with its ID.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Title of the task | |
| due_date | No | Date the task is due, as YYYY-MM-DD. | |
| priority | No | Priority level (0-5, higher is more important) | |
| project_id | Yes | The ID of the project to add the task to | |
| sort_order | No | Position among the project's tasks (lower sorts first). NOT wired on the REST create path — POST /api/projects/:id/tasks always assigns MAX(sort_order)+1 and ignores this field on create. Set it after creation with update_task instead. | |
| start_date | No | Date work is planned to start, as YYYY-MM-DD. | |
| description | No | Detailed description of the task | |
| is_milestone | No | Mark this task as a milestone, drawn differently on the Gantt view. NOT wired on the REST create path — POST /api/projects/:id/tasks does not accept this field; the column is left at its default (0). Set it after creation with update_task instead. | |
| github_pr_url | No | URL of the GitHub Pull Request implementing this task. Used to fetch the PR diff for AI verification and to match incoming CI status (check_run/check_suite/workflow_run webhooks) back to this task. | |
| parent_task_id | No | Make this a SUBTASK of the given task. The parent must be in the same project, and a subtask cannot itself have subtasks (one level only). A parent with subtasks takes its status from them: any subtask in progress makes the parent in progress, and all subtasks done makes it ready to complete. | |
| assignee_user_id | No | User ID to assign the task to. Must be a member of the same account. Omit to leave the task unassigned. For a SECOND (or third, ...) assignee, do not call this again — use add_assignee after creation, since this field only ever sets a single owner at create time. | |
| estimated_minutes | No | Estimated time to complete in minutes | |
| verification_mode | No | Gate this task's own completion. "manual" (default) allows a plain status change to done. "ai_review" or "all" BLOCK marking the task done (403 VERIFICATION_REQUIRED) until ai_completion_score reaches 0.8 — run "Verify with AI" first. "ci_required" or "all" BLOCK it until a PR is linked (github_pr_url) with passing CI. Setting this arms a real check against your own future attempt to close the task. | |
| required_ci_checks | No | Which named CI checks (matched case-insensitively as a substring of the check name) must pass for "ci_required"/"all" verification_mode to consider CI satisfied — e.g. ["build","test"]. Read by GET /api/tasks/:id/ci-status's resolveRequiredChecks (app/functions/api/tasks/[id]/ci-status.js:13-32), NOT by the GitHub webhook handler, which sets ci_status unconditionally from whichever single check/workflow event just arrived. Omit to leave unset (falls back to the project's own required_ci_checks, then to requiring every check). NOT wired on the REST create path — POST /api/projects/:id/tasks does not accept this field. Set it after creation with update_task instead. | |
| acceptance_criteria | No | Criteria for task completion |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the return behavior ('Returns the created task with its ID'), which is genuinely useful since there is no output schema. But it says nothing about permissions, side effects, or the fact that several declared parameters are silently ignored on this path.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, purpose front-loaded and return behavior immediately after. No filler, nothing that fails to earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose and return shape, which compensates for the absent output schema. But for a 15-parameter mutation tool with no annotations and many sibling competitors, it lacks the routing and behavioral context an agent needs to choose it confidently over update_task or the other create_* tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the per-parameter docs are unusually detailed, so the schema does the heavy lifting. The tool description adds no parameter information, which is the baseline 3 when coverage is this high.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Create a new task in a project' — which is unambiguous and cannot be confused with read tools like get_task or list_tasks. However, it does not differentiate itself from the other create_* siblings (create_project, create_product) or position itself relative to update_task.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use context, no prerequisites, and no mention of alternatives. The schema does reference update_task for several fields that are not wired on the create path, but that routing guidance lives in the parameter descriptions, not the tool description where an agent deciding whether to call create vs update would look.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_productA
Get details of a specific product including linked projects, team members, and statistics.
| Name | Required | Description | Default |
|---|---|---|---|
| product_id | Yes | The ID of the product to retrieve |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. The verb 'Get' implies a read-only operation, and the description adds context about the return payload (linked projects, team members, statistics). However, it does not disclose potential errors, authorization requirements, or any side effects, which would be useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one concise sentence that directly conveys the purpose and expected return content without unnecessary fluff. It is front-loaded and effectively communicates the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (one parameter, no output schema, no annotations), the description is adequately complete. It specifies the key return components (linked projects, team members, statistics) which helps set expectations, though it omits edge-case behavior or explicit usage context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides 100% coverage for the single parameter product_id with a clear description. The tool description does not add further semantics about the parameter, and none is needed because the schema adequately defines it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Get') and resource ('product'), and differentiates from list_products by emphasizing 'specific product' and enumerating the returned details (linked projects, team members, statistics). This clearly distinguishes it from the get_project and list_products siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for retrieving a single product (via 'specific product'), but does not explicitly name alternatives like list_products or provide when-to-use versus when-not-to-use guidance. It gives context but lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_projectA
Get details of a specific project including name, description, status, and task statistics.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | The ID of the project to retrieve |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It states only that the tool fetches details and lists fields, but it does not disclose behavior around not-found errors, permissions, or any side effects beyond being a read operation. However, the nature of 'get' makes it clearly non-destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single concise sentence with no filler, completely front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter get tool without an output schema, the description provides the key information: what it does and what details are included. It could mention return format or error handling, but the core is covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema fully documents the only parameter (project_id) with a clear description, so the description adds no additional semantic value. Baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Get') and identifies the resource ('specific project') and the key output components ('name, description, status, and task statistics'), clearly distinguishing it from sibling tools like list_projects or create_project.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for a single project via 'specific project', but it does not explicitly mention when to use this instead of list_projects or alternatives. No exclusions or conditions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_taskA
Get detailed information about a specific task including description, acceptance criteria, time entries, and status.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | The ID of the task to retrieve |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations are provided, the description itself must convey behavior. It clearly indicates a read operation (via 'Get') and lists the returned fields, which is useful. However, it does not disclose error behavior (e.g., 404 on missing task), permission requirements, or any potential side effects, leaving gaps in full behavioral transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that leads with the verb and resource, directly states the purpose, and lists the specific information returned. There is no redundant or extraneous wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that the tool has only one parameter and no output schema, the description adequately conveys what the user gets by listing the returned components. It lacks details on edge cases (e.g., missing task, authentication) but is sufficiently complete for a straightforward retrieval tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides a complete description for the single parameter task_id ('The ID of the task to retrieve'), so the tool description adds no additional semantic meaning. With 100% schema coverage, the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get'), the resource ('specific task'), and enumerates the returned information ('description, acceptance criteria, time entries, and status'). This unambiguously distinguishes it from sibling tools like list_tasks (which lists tasks) and get_project/get_product (which target different resources).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies that this tool is for retrieving details of a single task, but it does not explicitly state when to use it over alternatives like list_tasks or search_tasks. It also does not mention any exclusions or prerequisites beyond the obvious 'need a task ID.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
link_project_to_productB
Link an existing project to a product.
| Name | Required | Description | Default |
|---|---|---|---|
| is_primary | No | Whether this is the primary project for the product | |
| product_id | Yes | The ID of the product | |
| project_id | Yes | The ID of the project to link |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of disclosing behavioral traits. It does not state whether links are idempotent, whether it overwrites existing links, what happens if the project is already linked, or what the response looks like. The description is too sparse to convey any side effects, permissions, or other operational details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loaded with the action, and contains no filler. It is concise, but slightly under-specified for the broader context, making it somewhat too terse. Still, it earns its place and is well-structured for a simple linking operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description must compensate by explaining what the tool returns or how it behaves. It only states the action without addressing outcomes, error conditions, or the effect on existing relationships. For a tool that modifies an association, this is a significant gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides descriptions for all three parameters (product_id, project_id, is_primary), achieving 100% schema coverage. The description itself does not add further meaning beyond the schema, so a baseline score of 3 is appropriate. No additional parameter clarification is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action: linking an existing project to a product. It uses a precise verb ('link') and identifies both the object and the target, distinguishing it from sibling tools like create_project, create_product, or update_project. No ambiguity about the tool's core function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, nor does it mention any prerequisites or constraints. There is no mention of when linking is appropriate, whether the project/product must already exist, or any exclusions. For a tool that associates two existing entities, this leaves the agent with little context for decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_productsA
List all products accessible to the authenticated user. Products are top-level containers for organizing related projects.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Filter by product status. Defaults to all. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It adds context about authentication scope ('accessible to the authenticated user') and the nature of products as containers, but does not disclose pagination, sorting, or result structure. For a simple list operation, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary purpose, and the second sentence adds useful domain context. No redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one optional parameter and no output schema, the description gives sufficient context about the resource and scope. It could mention pagination or result format, but the simplicity of the tool makes it nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides 100% coverage for the single 'status' parameter with a clear description. The tool description adds no additional parameter context, but none is needed. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb+resource: 'List all products accessible to the authenticated user.' It explicitly distinguishes products as 'top-level containers for organizing related projects,' differentiating from sibling tools like list_tasks and list_projects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by naming the function, but does not explicitly state when to use it versus alternatives like get_product or create_product. No when-not or alternative recommendations provided, leaving the agent to infer from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_projectsA
List all projects accessible to the authenticated user. Returns project names, IDs, status, and task counts.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Filter by project status. Defaults to all. | |
| parent_id | No | Filter by parent project ID to get sub-projects. Omit for top-level projects. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It adds useful context by stating that it lists projects accessible to the authenticated user and specifies the return fields (names, IDs, status, task counts). However, it does not explicitly confirm read-only behavior, mention pagination, sorting, or any rate limits, leaving some transparency gaps. This is a modest but not exhaustive disclosure profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, front-loaded with the core purpose, and contains zero extraneous information. Every word contributes meaning: scope, return fields, and access context. This is an exemplary concise description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity list tool with two optional parameters and no output schema, the description covers the essential points: what it lists, for whom, and what it returns. It does not mention pagination or sorting, which are common for list tools, but given the simplicity and the fact that the return fields are named, it is sufficiently complete. A score of 4 reflects that it meets the needs without over-specifying.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for both parameters (status and parent_id), so the description does not need to compensate. The description itself adds no extra parameter details beyond what the schema already provides, which matches the baseline score of 3 for full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and resource ('projects') with clear scope ('accessible to the authenticated user'), and distinguishes itself from sibling tools like list_tasks and get_project by focusing on project listing with project-specific fields. It is unambiguous and immediately tells the agent what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: it lists all projects the user can access. It does not explicitly mention alternatives or exclusions (e.g., 'for a single project, use get_project'), but the context is sufficient for an agent to infer typical use cases. A 4 is appropriate because the context is clear, though no alternatives are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tasksA
List all tasks for a specific project. Returns compact rows (id, title, status, priority, estimated_minutes, time_spent, due_date, snippet) by default to keep payloads small; pass verbose:true for full rows including description, acceptance_criteria, tags, and assignments.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Filter by task status. Omit for all tasks. Note this filters the STORED status; a task can also be showing as blocked because of an unmet dependency (see effective_status on each row). | |
| verbose | No | Return full task rows (description, acceptance_criteria, tags, assignments) instead of compact rows. Default false. | |
| project_id | Yes | The ID of the project whose tasks to list |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the burden. It explicitly calls out the default compact row shape, lists the exact compact fields, and explains the verbose:true switch for full fields. This is meaningful return-format transparency for a read-only list operation. Pagination or rate limits are not mentioned, but nothing suggests hidden side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The default behavior and the verbose alternative are front-loaded, and both field lists are presented compactly. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter list tool with no output schema and no annotations, the description covers the required project_id, the default compact representation, and the full-field verbose alternative. The main missing context is pagination or ordering behavior, but this is unlikely to block correct use for typical calls.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents all 3 parameters with 100% coverage. The description adds value by enumerating the default compact output fields and clarifying what verbose:true changes, effectively compensating for the missing output schema. The status filter nuance is already well covered in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and scope: 'List all tasks for a specific project.' The plural 'all tasks' plus the project-scoping clearly differentiates it from get_task (single task) and search_tasks (cross-project search) without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the main use case—enumerating tasks for one project—but does not explicitly say when to prefer this tool over search_tasks or get_task, nor when not to use it. No exclusions or alternatives are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
log_timeB
Log time spent on a task. This triggers a real-time notification showing time was logged.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | Date for the time entry (YYYY-MM-DD). Defaults to today. | |
| notes | No | Notes about what was done during this time | |
| minutes | Yes | Number of minutes to log | |
| task_id | Yes | The ID of the task to log time against |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does add one meaningful behavioral trait: 'This triggers a real-time notification showing time was logged.' However, it doesn't mention permissions, reversibility, or effects on task status, leaving significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long and immediately states the core action. The second sentence adds a relevant side-effect without any unnecessary fluff. Every word contributes value and the structure is ideally front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with fully documented parameters, but it lacks usage guidance, output details, and deeper behavioral context. The description explains the notification side effect but doesn't cover when to use it or what the response contains. Given the simplicity, a score of 3 reflects adequate but incomplete context for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides 100% coverage of parameter descriptions, so the schema handles the meaning of task_id, minutes, date, and notes. The description adds no additional parameter-level detail beyond what the schema already states, meeting the baseline expectation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Log time spent on a task.' It uses a specific verb and resource, making the primary purpose unambiguous. However, it doesn't distinguish itself from the sibling tool 'quick_log', which likely serves a similar quick-logging purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'quick_log' or 'start_working'. No context, prerequisites, or exclusions are given. Users are left to infer the appropriate use case from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pause_workingA
Pause your open work session without ending it, and record WHY. Use this whenever work stops but is not finished — you asked the human a question, you are waiting on another agent, you are blocked on a build.
THE REASON IS THE POINT, and it decides the number: waiting_on_human — you are blocked on a person and CANNOT proceed. The paused interval is EXCLUDED from worked time. waiting_on_agent — you are blocked on another agent or an automated run that is itself active. The interval IS COUNTED as worked time: a sub-agent blocked on another active sub-agent is still working. other — anything else. Counted as worked.
Only waiting on the human is not working. Choose waiting_on_human ONLY when a person has to act before you can continue; if a machine or another agent is doing the work, it is waiting_on_agent.
Resume with resume_working. Do NOT use stop_working for a pause — that ends the session and the reason is lost.
| Name | Required | Description | Default |
|---|---|---|---|
| detail | No | Optional free text: what specifically you are waiting on (e.g. "asked whether to merge PR #280"). | |
| reason | Yes | Why work stopped. waiting_on_human EXCLUDES this interval from worked time; waiting_on_agent and other INCLUDE it. Required — there is no default, because the whole worked-time figure turns on this value. | |
| task_id | Yes | The ID of the task whose session to pause |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: it discloses that the session is not ended, that the reason value determines whether the interval counts as worked time, and that stop_working would lose the reason. This is exactly the behavioral context an agent needs to choose correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the reason taxonomy, then the resume/stop routing rule — a logical order. The all-caps emphasis and repeated worked-time explanation cost a little economy, but every block is load-bearing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations exist, so the description must stand alone — and it does: action, prerequisites, enum semantics, and the boundary against the two neighbouring session tools are all covered. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description goes beyond the schema by explaining the decision rule behind the reason enum (a sub-agent blocked on another active sub-agent is still 'working') and why there is no default. That is genuine added semantic value over the enum's own text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Pause your open work session without ending it') and immediately differentiates itself from stop_working, which it explicitly names. An agent can distinguish it from the three sibling working-session tools without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use triggers ('you asked the human a question, waiting on another agent, blocked on a build'), names the correct follow-up tool (resume_working), and states an explicit when-NOT-to-use rule ('Do NOT use stop_working for a pause'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quick_logA
Quick workflow for logging completed work: creates a task, marks it done, and logs time in one call. Ideal for recording work that has already been completed.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | Notes about what was done during this time (optional, defaults to description if not provided) | |
| title | Yes | Title of the completed task | |
| minutes | Yes | Number of minutes spent on this work | |
| project_id | Yes | The ID of the project to add the task to | |
| description | No | Detailed description of the task (optional) | |
| acceptance_criteria | No | Criteria for task completion (optional) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the three main side effects (creates a task, marks it done, logs time) but does not go deeper into specifics like permission requirements, reversibility, or what happens to optional fields. This is adequate but minimal for a composite operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, front-loaded with the primary purpose, and every word earns its place. It clearly states the composite behavior and the ideal use case without any fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's composite nature and full parameter schema, the description is nearly complete. It explains the main behavior and when to use it. It could add details about return values or prerequisites, but the absence of an output schema and the clarity of the description keep it sufficient. Slightly incomplete because it doesn't mention that the task is created as 'done' (though that is implied by 'marks it done'), but still strong.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not add parameter-level detail beyond what the schema already explains; it only frames the overall purpose. It does not explain how optional fields like notes, description, or acceptance_criteria relate to the composite workflow, but that is not required given full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly and specifically states what the tool does: 'creates a task, marks it done, and logs time in one call.' This distinguishes it from sibling tools like create_task, update_task_status, and log_time by framing it as a composite workflow for completed work. The verb 'quick workflow for logging completed work' is specific and action-oriented.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: 'Ideal for recording work that has already been completed.' This implies when to use it (for completed work) and contrasts with the alternative of using separate tools (create_task, log_time, etc.). It does not explicitly name alternatives or say when not to use it, but the context is strong enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_verification_feedbackA
Record whether the most recent verify_task verdict on this task was accurate — calibration data, not a correction. Attaches to the latest ai_generations row of type completion_review for this task, so verify_task must have run at least once first. Does NOT change the stored ai_completion_score or re-run any model — it costs nothing (no Anthropic call, just an audit-log row) and is safe to call as often as useful.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | Optional: why you agreed or disagreed. | |
| agreed | Yes | Whether you agree the AI verdict (score + summary) was an accurate read of the work. | |
| task_id | Yes | The ID of the task whose most recent verification you are giving feedback on. | |
| actual_score | No | Optional: the score (0.0-1.0) you believe was actually correct, if you disagree with the AI's score. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it states it does NOT change the stored ai_completion_score, does not re-run any model, costs nothing (no Anthropic call, only an audit-log row), and is idempotent-safe. This is exactly the behavioral context an agent needs before calling a mutation-adjacent tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and its not-a-correction caveat, then the prerequisite, then the no-cost assurance. Dense with em-dashes but every clause earns its place; only slight compression loss in readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, yet the description fully covers purpose, prerequisite, mutation scope, cost, and safety. An agent has everything needed to call this correctly on the first attempt.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so agreed, notes, task_id, and actual_score are already documented in the schema. The description adds context (task_id refers to the task whose most recent verification is being judged) but no syntax or constraint beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (record) and resource (verification feedback) and immediately disambiguates intent: 'calibration data, not a correction.' It distinguishes itself from the sibling verify_task by explaining it attaches to an ai_generations row rather than producing a verdict.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear prerequisite — verify_task must have run at least once first — and notes it is safe to call as often as useful. It does not name a competing sibling to route away from, but the context of use is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_assigneeA
Remove exactly one user's assignment from a task, leaving every other assignee untouched. DELETE /api/tasks/:id/assignments?user_id=... removes that one task_assignments row outright (both the assignment and any is_active flag on it) — this is a full unassign, not stop_working's "still assigned, no longer active" (use stop_working to end a work session without unassigning).
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | The ID of the task to remove the assignment from. | |
| user_id | Yes | User ID to remove from the task's assignees. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and mostly does: it discloses the underlying DELETE endpoint and that the task_assignments row (plus its is_active flag) is removed outright, making the destructive, non-reversible nature clear. It does not mention required permissions or auth, which is the only notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and the count constraint, followed by the mechanism and the sibling contrast. The endpoint path is slightly verbose but it earns its place by clarifying the deletion semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param mutation tool with no output schema, the description covers action, scope, mechanism, and sibling disambiguation well. Only auth/permission expectations are missing, which is a minor gap rather than a blocking one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented in the schema. The description adds only mild value by showing user_id as a query parameter and task_id in the path, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (remove) and resource (one user's assignment on a task) with precise scope — exactly one assignee, others untouched. It distinguishes itself from the sibling stop_working by contrast, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative (stop_working) and the condition that selects it: use stop_working to end a work session without unassigning. This is a clear when-to-use-this vs when-to-use-that statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resume_workingA
Resume your paused work session — closes the pause by recording the UTC instant you came back, so the paused interval has both boundaries and can be measured. The session itself was never closed and keeps its original start.
Fails if the session is not paused, rather than pretending to resume something. start_working also resumes a paused session, so use whichever reads better; this one does not re-fetch the orientation block.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | The ID of the task whose session to resume |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses that the pause is closed by recording a UTC instant, that the session keeps its original start, that non-paused sessions cause failure, and that no orientation block is re-fetched. This is unusually complete for a simple mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short paragraphs are front-loaded with the core action and effect, then failure behavior and sibling routing. Every sentence earns its place by clarifying semantics or routing, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There are no annotations and no output schema, so the description must supply the behavioral context, and it does: mutation effect, failure mode, session state preservation, and alternative tool guidance. Nothing an agent needs in order to call it correctly appears missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single task_id parameter, so the schema already documents its meaning. The description does not add format, constraints, or usage nuance for task_id beyond what the schema provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: resuming a paused work session by recording the return instant. It also explicitly distinguishes itself from start_working, which also resumes a paused session but re-fetches the orientation block. An agent can select it correctly without opening the sibling schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives the key condition for success (session must be paused) and states what happens otherwise: it fails rather than pretending to resume. It also names start_working as an alternative and clarifies the practical difference, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_tasksA
Search for tasks across all projects by title or description. Returns compact rows (id, title, status, priority, project_id, project_name, product_name, customer_name, estimated_minutes, time_logged, due_date, snippet) ranked with title matches first, then description matches, then recency. Pass verbose:true for full rows.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results to return. Default 20, capped at 100. | |
| query | Yes | Search query to match against task titles and descriptions | |
| status | No | Filter results by status | |
| verbose | No | Return full task rows instead of compact rows. Default false. | |
| product_id | No | Restrict results to this product | |
| project_id | No | Restrict results to this project | |
| customer_id | No | Restrict results to this customer |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does a solid job: it discloses matching fields, output row shape, ranking order (title matches, then description, then recency), and the verbose option. It stops short of noting empty-result behavior or other edge cases, but the core behavioral contract is well communicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences carry all essential information with no filler. The main purpose is front-loaded, and the second sentence packs output details, ranking, and the verbose switch into one efficient, readable unit. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, describing the exact compact row fields is essential and well handled. The description also covers ranking and the verbose toggle. The main gap is that 'full rows' fields are not enumerated and no usage recommendation is given for choosing between compact and verbose output, but the tool is still safely invocable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers all 7 parameters at 100%, so the baseline is 3. The description adds value beyond the schema by explaining how query matches are ranked and what verbose:true changes, which directly informs parameter usage. It does not repeat parameter descriptions, appropriately relying on the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Search'), identifies the resource ('tasks'), and scopes the operation ('across all projects by title or description'). It also details the output format and ranking behavior, making the tool's purpose unmistakable even without inspecting the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies use for finding tasks by text across projects, and the 'across all projects' scope helps distinguish it from more narrow listing tools. However, it does not explicitly name any alternative sibling tool (e.g., list_tasks) or state when not to use this tool, leaving routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_resume_contextA
Set (overwrite) the task's resume context — a single pinned block describing the current state, what is done, what is next, and where to look. This is read first when a work session restarts cold (e.g. via start_working) so you can re-orient immediately. Update it at the end of a work session.
| Name | Required | Description | Default |
|---|---|---|---|
| context | Yes | The resume-context text (markdown ok): current state / next steps / where to look. | |
| task_id | Yes | The ID of the task to set resume context for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose the critical mutation trait — that this overwrites rather than appends, and that only a single pinned block exists. It omits permission/auth requirements and any error behavior, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the verb and overwrite semantics, then the rationale for the data, then the actionable timing. No padding or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema and no annotations, the description covers purpose, content shape, consumption path (start_working) and when to update. Additional notes on the effect of overwriting existing context or on failure modes would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema. The description reinforces what the context text should contain (state / next steps / where to look), but adds no syntax or format detail beyond what the schema already states, matching the baseline for full coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Set (overwrite) the task's resume context') and defines exactly what the resource is — a single pinned block covering current state, done, next, and where to look. It also names the sibling flow (start_working) that consumes this data, so an agent can distinguish it from generic note/update tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use guidance: 'Update it at the end of a work session,' and explains it is read first on a cold restart via start_working. It does not name explicit exclusions or alternatives (e.g. how this differs from add_work_note), so it falls short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_workingA
Start work on a task. Two things happen: you are shown to the team as actively working on it (and a todo task moves to in_progress), and a TIMED WORK SESSION is opened with a UTC start instant, so how long the task was actually worked becomes measurable instead of guessed.
Calling it a second time on the same task NEVER opens a second session. If your session is paused it is resumed; if it is already running you are told so, with how long it has been going. The response always says which of the three happened.
It also orients a cold session: the response carries the task's resume context and its recent work notes, so read the whole thing before you start. Pair it with stop_working, and use pause_working when you have to wait on something.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | The ID of the task to start working on |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does so well: it discloses the two side effects (team visibility + todo to in_progress, plus a timed session with a UTC start instant), the idempotent/resume semantics, the three possible outcomes, and that the response carries resume context and recent work notes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and side effects, and the trailing guidance about reading the response and pairing tools is genuinely useful. Slightly long, but every sentence carries actionable information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation-style tool with no annotations and no output schema, the description covers behavior, idempotency, and what the response contains (which of three outcomes, resume context, work notes). Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (task_id) and schema description coverage is 100%, so the schema already fully documents it. The description adds no syntax or format detail beyond the schema; baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (start work on a task) and immediately scopes it against siblings stop_working, pause_working and resume_working. An agent can distinguish it from those alternatives without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes usage: pair with stop_working, use pause_working when waiting on something, and it clarifies that repeat calls never open a second session (resume vs already-running). When-to-use and when-not-to-expect-a-new-session are both stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_workingA
Stop working on a task: close your open work session with a UTC end instant, and drop the "actively working" flag (you stay assigned). The closed session is what makes the task's worked duration reportable.
Requires an open session — if you have none this FAILS and writes nothing, rather than inventing a session with no start. Call start_working first.
This ENDS the session. If you are only waiting on something and intend to carry on afterwards, use pause_working instead: a pause keeps the session open and records why you stopped, which is what decides whether the interval counts as worked time.
Optionally capture your state on the way out: pass note to append a work note and/or resume_context to overwrite the resume block, so the next session can pick up where you left off.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Optional work note to append before stopping (what you did, what is next). | |
| task_id | Yes | The ID of the task to stop working on | |
| resume_context | No | Optional resume-context block to overwrite before stopping (current state / next steps / where to look). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that the operation mutates session state, that it FAILS and writes nothing when no session exists, that the flag is dropped while assignment persists, and that the closed session is what makes duration reportable. Optional side effects (note appended, resume_context overwritten) are also disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is longer than most definitions but well structured: core action first, then prerequisite/failure, then the pause_working distinction, then optional state capture. Every paragraph earns its place, though the middle paragraph could be tightened by a sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers purpose, prerequisites, failure semantics, alternatives, and side effects of the optional parameters. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: note is 'append[ed]... on the way out' and resume_context 'overwrite[s] the resume block' so the next session can continue. It clarifies sequencing (capture state before stopping) that the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Stop working on a task: close your open work session') and states the concrete effects (UTC end instant, dropping the 'actively working' flag, task stays assigned). It is immediately distinguishable from start_working and pause_working.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the prerequisite ('Requires an open session... Call start_working first'), the failure behavior, and explicitly names the alternative for the adjacent case: 'If you are only waiting on something and intend to carry on afterwards, use pause_working instead.' When-to-use, when-not, and the alternative are all present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_projectB
Update project details like name, description, or status.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | New name for the project | |
| status | No | New status for the project | |
| project_id | Yes | The ID of the project to update | |
| description | No | New description |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of explaining behavior. It only states the action without disclosing whether updates are partial, what is returned, or any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that gets straight to the point. It avoids unnecessary words and clearly says what the tool does.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool description is minimal but sufficient for a simple update operation, especially with full schema descriptions. However, it lacks any mention of the return value or update semantics, and there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides descriptions for all four parameters (100% coverage), so the description adds little beyond restating three of the fields. It does not clarify the required project_id parameter or the meaning of enum values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function with a specific verb ('Update') and resource ('project details'), and lists concrete fields (name, description, status). This distinguishes it from sibling tools like list_projects or create_project.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool relative to alternatives. It does not mention any prerequisites, conditions, or alternative tools for updating projects.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_taskC
Update task details like title, description, estimates, or acceptance criteria.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | New title for the task | |
| task_id | Yes | The ID of the task to update | |
| due_date | No | New due date as YYYY-MM-DD. Pass an empty string to clear it. | |
| priority | No | New priority level (0-5, higher is more important) | |
| blocked_by | No | REPLACES the full set of tasks this task is blocked by. Pass [] to clear all of them. Every id must be a task in the same product; a circular chain (A blocked by B blocked by A) is rejected. While any of these is not done, the task shows as blocked whatever its stored status says. | |
| sort_order | No | Position among the project's tasks (lower sorts first). Drives the order returned by list_tasks and the board/list views. | |
| start_date | No | New start date as YYYY-MM-DD. Pass an empty string to clear it. | |
| description | No | New description | |
| is_milestone | No | Mark this task as a milestone, drawn differently on the Gantt view. Pass false to unmark it. | |
| github_pr_url | No | URL of the GitHub Pull Request implementing this task. Used to fetch the PR diff for AI verification and to match incoming CI status back to this task. Pass an empty string to clear it. | |
| blocked_reason | No | Free-text reason shown next to a manually-set Blocked status ("waiting on the client"). Cleared automatically when the task leaves the blocked status. | |
| assignee_user_id | No | User ID to assign the task to, REPLACING THE WHOLE assignee set (every non-active assignment row is deleted, active workers spared) — use this only when you want exactly one owner. Must be a member of the same account. Pass an empty string to unassign everyone (active workers still spared). To ADD a second assignee without touching the first, use add_assignee instead; to remove just one, use remove_assignee. | |
| estimated_minutes | No | New time estimate in minutes | |
| verification_mode | No | Gate this task's own completion. "manual" (default) allows a plain status change to done. "ai_review" or "all" BLOCK marking the task done (403 VERIFICATION_REQUIRED) until ai_completion_score reaches 0.8 — run "Verify with AI" first. "ci_required" or "all" BLOCK it until a PR is linked (github_pr_url) with passing CI. Setting this arms a real check against your own future attempt to close the task. | |
| required_ci_checks | No | Which named CI checks (matched case-insensitively as a substring of the check name) must pass for "ci_required"/"all" verification_mode to consider CI satisfied — e.g. ["build","test"]. Read by GET /api/tasks/:id/ci-status's resolveRequiredChecks (app/functions/api/tasks/[id]/ci-status.js:13-32), NOT by the GitHub webhook handler, which sets ci_status unconditionally from whichever single check/workflow event just arrived regardless of this field. Stored as a JSON array. Pass an empty array to require every check; this tool has no way to clear the column back to "inherit from project" (the REST API reserves that for an empty string, which this array-typed field cannot send). | |
| acceptance_criteria | No | New acceptance criteria |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does none of it. It never states that this is a partial update (only supplied fields change), that assignment changes delete other assignment rows, or what permissions are required — all of that lives only in scattered property descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste, and the verb+resource lead. It is appropriately sized for a description, though at 16 parameters that brevity starts to look like under-specification rather than tight writing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter mutation tool with no annotations and no output schema, one sentence is not enough. The schema covers parameters, but the description omits update semantics, sibling routing (update_task_status, add_assignee, remove_assignee), and any failure/permission context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter in depth (clearing semantics, REPLACES wording, verification gates). The description's field list ('title, description, estimates, acceptance criteria') adds nothing beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (task) and enumerates the kinds of fields it touches. It is clear on its own, but it never distinguishes itself from the sibling update_task_status, which an agent must choose between for status changes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no exclusions. The schema hints at routing to add_assignee/remove_assignee for assignment edits and update_task_status exists as a sibling, but the description says nothing about which tool to pick or when.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_task_statusA
Change the status of a task (todo, in_progress, blocked, done). This triggers a real-time notification to users viewing the project. Note that a task with subtasks or dependencies also has a DERIVED status which can differ from the one set here — see effective_status / status_reason on get_task.
| Name | Required | Description | Default |
|---|---|---|---|
| status | Yes | New status for the task. "ready_to_complete" is NOT settable — it is derived when every subtask is done. | |
| task_id | Yes | The ID of the task to update | |
| blocked_reason | No | Why the task is blocked. Only meaningful with status "blocked"; cleared automatically on any other status. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It goes beyond the basic mutation by noting that a real-time notification is triggered and that a derived status may differ from the explicitly set one, which are meaningful side effects. It does not cover permissions or reversibility, but the most important behavioral nuances are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words. The action is front-loaded, followed by the notification side effect and then the important derived-status caveat. Every sentence adds necessary information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter update tool with full schema coverage, the description is largely complete: it covers the core action, side effects, and the key edge case of derived status, while pointing to get_task for the effective values. It does not describe the return value, but there is no output schema and none is strictly necessary for a status update; a brief note on permissions would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by explaining the distinction between the settable status and the derived status, and by pointing to effective_status/status_reason on get_task. It does not discuss blocked_reason, but the schema already fully documents that parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Change'), a clear resource ('status of a task'), and enumerates the allowed statuses, so the tool's purpose is immediately obvious. It does not explicitly distinguish itself from the sibling 'update_task', which is a general-purpose update tool, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for changing task status and warns about derived status differing from the set value, which is useful context. However, it does not explicitly say when to use this tool versus the sibling 'update_task' or any alternative, nor does it mention prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_taskA
Run AI verification: grades this task's acceptance_criteria against the diff of its linked PR (a Haiku relevance pre-screen, then a Sonnet deep review) and PERSISTS the result to ai_completion_score / ai_completion_notes. This is exactly what update_task_status checks when the task's verification_mode is "ai_review" or "all" — call this to satisfy that gate yourself rather than repeatedly hitting 403 VERIFICATION_REQUIRED with no way through it.
COSTS MONEY — DO NOT LOOP THIS. Every call is metered via recordAIUsage and recorded to ai_generations (Sonnet, large PR-diff context) for audit. It runs FREE while your account is within its plan's included monthly verification cap; once over that cap each call is billed as AI credit overage — 2 credits per call on this account. Calling this repeatedly hoping for a higher score spends real credits for no guaranteed gain; if a genuine score lands below the 0.8 gate, that is a finding about the work, not a reason to retry.
GRADES THE DIFF, NOT THE RUNNING SYSTEM. A high score means the PR's code changes look like they satisfy the criteria on paper — it is NOT proof the deployed/running behavior actually works. Never report this score as end-to-end verification; that is a human's job or an independent task-verify sub-agent's, not this tool's.
Requires the task to already have non-empty acceptance_criteria (set via update_task) and a way to see the work: either the task's own github_pr_url (set via update_task) or a pr_url/work_description passed here. Fails with a clear message naming what is missing rather than grading an empty contract — refusal is free, no credits spent.
| Name | Required | Description | Default |
|---|---|---|---|
| pr_url | No | GitHub PR URL to grade against, for this call only — does not overwrite the task's stored github_pr_url. Omit to use the task's own linked PR. | |
| task_id | Yes | The ID of the task to verify. Must already have non-empty acceptance_criteria. | |
| work_description | No | Free-text description of the completed work to grade against. Used ONLY as a fallback when no PR diff is available (no pr_url given here and the task has no github_pr_url); ignored whenever a PR diff can be fetched. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: it discloses real cost (metered via recordAIUsage, 2 credits/call over cap, audit-logged to ai_generations), persistence behavior, prerequisites, and graceful failure (refusal is free, no credits spent). It also bounds what the score means (diff on paper, not deployed behavior).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is longer than typical, but every paragraph earns its place — cost warning, scope caveat, prerequisites — and the core action is front-loaded in the first sentence. The caps-emphasis blocks aid scanning rather than pad the text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-annotation, no-output-schema tool this is unusually complete: it covers cost, side effects, prerequisites, failure modes, and the meaning of the result. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: pr_url applies 'for this call only' and does not overwrite the stored github_pr_url, and work_description is a fallback used only when no PR diff exists. The task_id prerequisite (non-empty acceptance_criteria) is reinforced in prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (AI verification grading acceptance_criteria against a linked PR diff) and describes the mechanism (Haiku pre-screen, Sonnet deep review) plus the persisted side effect (ai_completion_score / ai_completion_notes). It is clearly distinguishable from siblings like update_task_status and record_verification_feedback.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the trigger condition (task verification_mode of 'ai_review' or 'all', the same gate update_task_status checks) and gives strong when-not-to guidance: do not loop it hoping for a higher score, and do not treat it as end-to-end verification. It even routes the agent to the correct alternative for running-system proof.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v2.2.0- Added
add_assignee - Changed
add_comment1 field changed- removed
Input schema / properties / author_kindRemoved value: -{ - "description": "Who authored this comment: \"ai\" (default) or \"human\" (relaying the user).", - "enum": [ - "ai", - "human" - ], - "type": "string" -}
- Changed
add_work_note1 field changed- removed
Input schema / properties / author_kindRemoved value: -{ - "description": "Who authored this note: \"ai\" (default, your own update) or \"human\" (you are relaying something the user said).", - "enum": [ - "ai", - "human" - ], - "type": "string" -}
- Changed
create_task6 fields changed- changed
Input schema / properties / assignee_user_id / descriptionPrevious value: -"User ID to assign the task to. Must be a member of the same account. Omit to leave the task unassigned."New value: +"User ID to assign the task to. Must be a member of the same account. Omit to leave the task unassigned. For a SECOND (or third, ...) assignee, do not call this again — use add_assignee after creation, since this field only ever sets a single owner at create time." - added
Input schema / properties / github_pr_urlAdded value: +{ + "description": "URL of the GitHub Pull Request implementing this task. Used to fetch the PR diff for AI verification and to match incoming CI status (check_run/check_suite/workflow_run webhooks) back to this task.", + "type": "string" +} - added
Input schema / properties / is_milestoneAdded value: +{ + "description": "Mark this task as a milestone, drawn differently on the Gantt view. NOT wired on the REST create path — POST /api/projects/:id/tasks does not accept this field; the column is left at its default (0). Set it after creation with update_task instead.", + "type": "boolean" +} - added
Input schema / properties / required_ci_checksAdded value: +{ + "description": "Which named CI checks (matched case-insensitively as a substring of the check name) must pass for \"ci_required\"/\"all\" verification_mode to consider CI satisfied — e.g. [\"build\",\"test\"]. Read by GET /api/tasks/:id/ci-status's resolveRequiredChecks (app/functions/api/tasks/[id]/ci-status.js:13-32), NOT by the GitHub webhook handler, which sets ci_status unconditionally from whichever single check/workflow event just arrived. Omit to leave unset (falls back to the project's own required_ci_checks, then to requiring every check). NOT wired on the REST create path — POST /api/projects/:id/tasks does not accept this field. Set it after creation with update_task instead.", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / sort_orderAdded value: +{ + "description": "Position among the project's tasks (lower sorts first). NOT wired on the REST create path — POST /api/projects/:id/tasks always assigns MAX(sort_order)+1 and ignores this field on create. Set it after creation with update_task instead.", + "type": "number" +} - added
Input schema / properties / verification_modeAdded value: +{ + "description": "Gate this task's own completion. \"manual\" (default) allows a plain status change to done. \"ai_review\" or \"all\" BLOCK marking the task done (403 VERIFICATION_REQUIRED) until ai_completion_score reaches 0.8 — run \"Verify with AI\" first. \"ci_required\" or \"all\" BLOCK it until a PR is linked (github_pr_url) with passing CI. Setting this arms a real check against your own future attempt to close the task.", + "enum": [ + "manual", + "ai_review", + "ci_required", + "all" + ], + "type": "string" +}
- Added
pause_working - Added
record_verification_feedback - Added
remove_assignee - Added
resume_working - Changed
set_resume_context1 field changed- removed
Input schema / properties / author_kindRemoved value: -{ - "description": "Who authored this context: \"ai\" (default) or \"human\".", - "enum": [ - "ai", - "human" - ], - "type": "string" -}
- Changed
update_task6 fields changed- changed
Input schema / properties / assignee_user_id / descriptionPrevious value: -"User ID to assign the task to, replacing any existing assignee. Must be a member of the same account. Pass an empty string to unassign."New value: +"User ID to assign the task to, REPLACING THE WHOLE assignee set (every non-active assignment row is deleted, active workers spared) — use this only when you want exactly one owner. Must be a member of the same account. Pass an empty string to unassign everyone (active workers still spared). To ADD a second assignee without touching the first, use add_assignee instead; to remove just one, use remove_assignee." - added
Input schema / properties / github_pr_urlAdded value: +{ + "description": "URL of the GitHub Pull Request implementing this task. Used to fetch the PR diff for AI verification and to match incoming CI status back to this task. Pass an empty string to clear it.", + "type": "string" +} - added
Input schema / properties / is_milestoneAdded value: +{ + "description": "Mark this task as a milestone, drawn differently on the Gantt view. Pass false to unmark it.", + "type": "boolean" +} - added
Input schema / properties / required_ci_checksAdded value: +{ + "description": "Which named CI checks (matched case-insensitively as a substring of the check name) must pass for \"ci_required\"/\"all\" verification_mode to consider CI satisfied — e.g. [\"build\",\"test\"]. Read by GET /api/tasks/:id/ci-status's resolveRequiredChecks (app/functions/api/tasks/[id]/ci-status.js:13-32), NOT by the GitHub webhook handler, which sets ci_status unconditionally from whichever single check/workflow event just arrived regardless of this field. Stored as a JSON array. Pass an empty array to require every check; this tool has no way to clear the column back to \"inherit from project\" (the REST API reserves that for an empty string, which this array-typed field cannot send).", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / sort_orderAdded value: +{ + "description": "Position among the project's tasks (lower sorts first). Drives the order returned by list_tasks and the board/list views.", + "type": "number" +} - added
Input schema / properties / verification_modeAdded value: +{ + "description": "Gate this task's own completion. \"manual\" (default) allows a plain status change to done. \"ai_review\" or \"all\" BLOCK marking the task done (403 VERIFICATION_REQUIRED) until ai_completion_score reaches 0.8 — run \"Verify with AI\" first. \"ci_required\" or \"all\" BLOCK it until a PR is linked (github_pr_url) with passing CI. Setting this arms a real check against your own future attempt to close the task.", + "enum": [ + "manual", + "ai_review", + "ci_required", + "all" + ], + "type": "string" +}
- Added
verify_task
5 tool updates
v1.3.0- Changed
create_task4 fields changed- added
Input schema / properties / assignee_user_idAdded value: +{ + "description": "User ID to assign the task to. Must be a member of the same account. Omit to leave the task unassigned.", + "type": "string" +} - added
Input schema / properties / due_dateAdded value: +{ + "description": "Date the task is due, as YYYY-MM-DD.", + "type": "string" +} - added
Input schema / properties / parent_task_idAdded value: +{ + "description": "Make this a SUBTASK of the given task. The parent must be in the same project, and a subtask cannot itself have subtasks (one level only). A parent with subtasks takes its status from them: any subtask in progress makes the parent in progress, and all subtasks done makes it ready to complete.", + "type": "number" +} - added
Input schema / properties / start_dateAdded value: +{ + "description": "Date work is planned to start, as YYYY-MM-DD.", + "type": "string" +}
- Changed
list_tasks3 fields changed- changed
Input schema / properties / status / descriptionPrevious value: -"Filter by task status. Omit for all tasks."New value: +"Filter by task status. Omit for all tasks. Note this filters the STORED status; a task can also be showing as blocked because of an unmet dependency (see effective_status on each row)." - changed
Input schema / properties / status / enumPrevious value: -[ - "todo", - "in_progress", - "done" -]New value: +[ + "todo", + "in_progress", + "blocked", + "done" +] - added
Input schema / properties / verboseAdded value: +{ + "description": "Return full task rows (description, acceptance_criteria, tags, assignments) instead of compact rows. Default false.", + "type": "boolean" +}
- Changed
search_tasks6 fields changed- added
Input schema / properties / customer_idAdded value: +{ + "description": "Restrict results to this customer", + "type": "number" +} - added
Input schema / properties / limitAdded value: +{ + "description": "Max results to return. Default 20, capped at 100.", + "type": "number" +} - added
Input schema / properties / product_idAdded value: +{ + "description": "Restrict results to this product", + "type": "number" +} - added
Input schema / properties / project_idAdded value: +{ + "description": "Restrict results to this project", + "type": "number" +} - changed
Input schema / properties / status / enumPrevious value: -[ - "todo", - "in_progress", - "done" -]New value: +[ + "todo", + "in_progress", + "blocked", + "done" +] - added
Input schema / properties / verboseAdded value: +{ + "description": "Return full task rows instead of compact rows. Default false.", + "type": "boolean" +}
- Changed
update_task6 fields changed- added
Input schema / properties / assignee_user_idAdded value: +{ + "description": "User ID to assign the task to, replacing any existing assignee. Must be a member of the same account. Pass an empty string to unassign.", + "type": "string" +} - added
Input schema / properties / blocked_byAdded value: +{ + "description": "REPLACES the full set of tasks this task is blocked by. Pass [] to clear all of them. Every id must be a task in the same product; a circular chain (A blocked by B blocked by A) is rejected. While any of these is not done, the task shows as blocked whatever its stored status says.", + "items": { + "type": "number" + }, + "type": "array" +} - added
Input schema / properties / blocked_reasonAdded value: +{ + "description": "Free-text reason shown next to a manually-set Blocked status (\"waiting on the client\"). Cleared automatically when the task leaves the blocked status.", + "type": "string" +} - added
Input schema / properties / due_dateAdded value: +{ + "description": "New due date as YYYY-MM-DD. Pass an empty string to clear it.", + "type": "string" +} - changed
Input schema / properties / priority / descriptionPrevious value: -"New priority level"New value: +"New priority level (0-5, higher is more important)" - added
Input schema / properties / start_dateAdded value: +{ + "description": "New start date as YYYY-MM-DD. Pass an empty string to clear it.", + "type": "string" +}
- Changed
update_task_status3 fields changed- added
Input schema / properties / blocked_reasonAdded value: +{ + "description": "Why the task is blocked. Only meaningful with status \"blocked\"; cleared automatically on any other status.", + "type": "string" +} - changed
Input schema / properties / status / descriptionPrevious value: -"New status for the task"New value: +"New status for the task. \"ready_to_complete\" is NOT settable — it is derived when every subtask is done." - changed
Input schema / properties / status / enumPrevious value: -[ - "todo", - "in_progress", - "done" -]New value: +[ + "todo", + "in_progress", + "blocked", + "done" +]
21 tool updates
v1.2.0- First observed
add_comment - First observed
add_work_note - First observed
create_product - First observed
create_project - First observed
create_task - First observed
get_product - First observed
get_project - First observed
get_task - First observed
link_project_to_product - First observed
list_products - First observed
list_projects - First observed
list_tasks - First observed
log_time - First observed
quick_log - First observed
search_tasks - First observed
set_resume_context - First observed
start_working - First observed
stop_working - First observed
update_project - First observed
update_task - First observed
update_task_status
TDQS
Scored across 27 tools
Most tools have clearly distinct purposes, and the descriptions explicitly disambiguate overlapping areas like add_assignee vs. update_task's assignee replacement, and work-session tools vs. update_task_status. However, some conceptual proximity remains between start_working and update_task_status, and between log_time and quick_log, which could cause occasional misselection.
Tool names consistently use lowercase snake_case with action-oriented verbs, and most follow a predictable verb_noun pattern. Minor deviations like quick_log and gerund forms (start_working, pause_working, resume_working) keep it from being perfectly uniform.
With 27 tools, the server exceeds the typical well-scoped range and includes clear redundancy, such as quick_log duplicating create_task + update_task_status + log_time. While the domain is broad, several tools could be consolidated without losing capability.
Core task, project, and product lifecycle operations are largely present, but there are notable gaps: no delete operations for tasks, projects, or products, no update_product, and no customer management despite customer_id being required to create products. These missing operations create dead ends for common administrative workflows.
Maintenance
Related MCP Connectors
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
MCP Server for Slima - AI Writing IDE for Novel Authors with AI Beta Reader.
- mcpOAuthnet.todoist
Official Todoist MCP server for AI assistants to manage tasks, projects, and workflows.
MCP server for Linear project management and issue tracking
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceA comprehensive MCP server for time tracking, project management, and AI-powered memory storage using semantic search. It enables users to log time, manage client billing, and capture shared or personal ideas through integrated tools and team collaboration features.1-
- FlicenseNot gradedqualityDmaintenanceMCP Server to log time and view Jira statistics directly from Cursor IDE.-
- AlicenseBqualityFmaintenanceMCP server for Panel Todo — lets AI assistants manage your tasks, issues, and sprints directly in VS Code.3818 npmMIT
- FlicenseNot gradedqualityDmaintenanceMCP server that enables AI assistants in IDEs to interact with RefBase webapp for managing conversations, bugs, features, and project context via natural language.3 npm3-