Rhizm MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Rhizm MCP Serverstart a timer for my 'finish project report' task"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@rhizmapp/mcp
MCP (Model Context Protocol) server for Rhizm. Allows LLM clients (Claude Desktop, Claude Code, etc.) to operate tasks, timers, and notes via natural language.
Install
npm install -g @rhizmapp/mcpOr run directly with npx:
npx @rhizmapp/mcpRelated MCP server: Taskwarrior MCP Server
Setup
Generate an API key at https://rhizm.app/ → Settings → API Keys
Configure your MCP client (see below)
Claude Desktop
~/Library/Application Support/Claude/claude_desktop_config.json (macOS):
{
"mcpServers": {
"rhizm": {
"command": "npx",
"args": ["@rhizmapp/mcp"],
"env": {
"RHIZM_API_KEY": "rz_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"
}
}
}
}Claude Code
.claude/settings.json:
{
"mcpServers": {
"rhizm": {
"command": "npx",
"args": ["@rhizmapp/mcp"],
"env": {
"RHIZM_API_KEY": "rz_..."
}
}
}
}Environment Variables
Variable | Required | Default | Description |
| yes | — | Your API key ( |
| no |
| API base URL |
Tools
Tool | Required permission | Description |
| read | List tasks (date range / tag filter) |
| write | Create a task |
| write | Update a task |
| delete | Delete a task |
| timer | Start timer for a task |
| timer | Stop timer for a task |
| read | Get running timers |
| read | Aggregated time summary |
| read | List notes |
| read | Get a single note |
| write | Create a note |
Documentation
For full API documentation, see docs/api.md.
Available Tools
11 toolscreate_noteC
Create a new note.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | ||
| content | No | Plain text or Tiptap JSON string. | |
| folderId | No | ||
| tags | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. 'Create a new note' implies a write/mutation operation but reveals nothing about permissions, side effects, error conditions, rate limits, or what happens on success. This is inadequate for a tool that presumably modifies system state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise—just three words—and front-loaded with the core action. There's zero wasted language, though this brevity comes at the cost of completeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation tool with no annotations, no output schema, and low schema description coverage, this description is completely inadequate. It doesn't explain what a note is, how creation works, what parameters mean, or what to expect in return. The agent would be flying blind.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (only the 'content' parameter has a description). The tool description adds no information about any parameters—it doesn't mention title, content format specifics, folderId purpose, or tags usage. It fails to compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Create a new note' clearly states the action (create) and resource (note), but it's vague about what constitutes a note in this system and doesn't distinguish it from similar tools like 'create_task'. It's functional but lacks specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'create_task' or how it relates to sibling tools like 'list_notes' or 'get_note'. The description offers no context about prerequisites, constraints, or appropriate scenarios for note creation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_taskC
Create a new task.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Task title (required). | |
| scheduledDate | No | Scheduled date (YYYY-MM-DD) or null for inbox. | |
| tags | No | Tag names without leading #. | |
| estimatedMin | No | Estimated duration in minutes. | |
| priority | No | Priority: 1=low, 2=medium, 3=high. | |
| memo | No | Free-form memo. | |
| status | No | Initial status. Default: backlog. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. 'Create a new task' implies a write operation, but it doesn't specify required permissions, whether the operation is idempotent, what happens on failure, or what the response looks like. For a mutation tool with zero annotation coverage, this leaves significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is maximally concise with a single sentence containing only essential information. There's no wasted verbiage or unnecessary elaboration. It's front-loaded with the core action, making it immediately scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with 7 parameters, no annotations, and no output schema, the description is insufficient. It doesn't explain what happens after creation, what the return value is, or provide behavioral context. The agent must rely entirely on the input schema for understanding, which is inadequate for a write operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema fully documents all 7 parameters. The description adds no parameter information beyond what's in the schema. According to scoring rules, when schema coverage is high (>80%), the baseline is 3 even with no param info in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Create') and resource ('a new task'), making the purpose immediately understandable. It doesn't differentiate from sibling tools like 'update_task' or 'delete_task', but the basic action is unambiguous. The description avoids tautology by not just restating the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention when to choose 'create_task' over 'update_task' or 'list_tasks', nor does it specify prerequisites or appropriate contexts. The agent must infer usage from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_taskA
Delete a task permanently. ⚠ Cannot be undone — confirm with the user before calling. Time entries linked to this task will also be deleted (cascade).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Task ID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well by disclosing critical behavioral traits: the permanent deletion ('permanently'), the cascade effect ('Time entries linked to this task will also be deleted'), and the need for user confirmation. It doesn't cover all possible aspects like error conditions or permissions, but it adds substantial value beyond basic function.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and warning, followed by additional context in a second sentence. Every sentence earns its place by providing essential information without waste, making it highly efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (destructive operation with cascade effects) and lack of annotations or output schema, the description does a good job covering key aspects like permanence and cascading deletions. It could be more complete by mentioning potential errors or response format, but it's largely adequate for the context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, clearly documenting the single parameter 'id' as 'Task ID.' The description doesn't add any parameter-specific details beyond what the schema provides, so it meets the baseline for high schema coverage without extra value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Delete a task permanently') and resource ('task'), distinguishing it from sibling tools like 'update_task' or 'list_tasks' which perform different operations on the same resource type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit guidance on when to use this tool ('confirm with the user before calling') and highlights the irreversible nature ('Cannot be undone'), which helps differentiate it from safer operations like 'update_task' or 'list_tasks' among the siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_noteA
Get a single note with full content (Tiptap JSON or plain text).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Note ID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions the content format (Tiptap JSON or plain text) but does not cover other behavioral aspects such as error handling (e.g., what happens if the ID is invalid), authentication needs, rate limits, or response structure. The description adds minimal context beyond the basic operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the key information (action, resource, content details) with zero waste. Every word earns its place, making it highly concise and well-structured for quick understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (single parameter, no output schema, no annotations), the description is adequate but has clear gaps. It covers the purpose and content format but lacks details on behavioral traits, error handling, and output expectations, which are important for a read operation without annotations. It meets minimum viability but could be more complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, with the single parameter 'id' documented as 'Note ID.' in the schema. The description does not add any further meaning or details about the parameter beyond what the schema provides, such as format examples or constraints. Baseline 3 is appropriate given the high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Get') and resource ('a single note'), including the scope ('with full content') and format details ('Tiptap JSON or plain text'). It distinguishes from sibling tools like 'list_notes' by specifying retrieval of a single note rather than listing multiple notes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when needing a specific note's full content, but does not explicitly state when to use this tool versus alternatives like 'list_notes' or other sibling tools. No exclusions or prerequisites are mentioned, leaving usage context somewhat open-ended.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_running_timersB
Get all currently running timers (with task title).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool retrieves running timers but doesn't cover aspects like whether it requires authentication, how it handles errors, if it's read-only (implied by 'Get'), or what the return format looks like. This is a significant gap for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose without any wasted words. It's appropriately sized for a simple tool with no parameters, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (0 parameters, no output schema, no annotations), the description is minimally adequate. It states what the tool does but lacks details on behavioral traits, usage context, and output format. This makes it incomplete for full agent understanding, though not severely lacking given the simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so the schema fully documents the lack of inputs. The description doesn't add parameter details beyond what the schema provides, but with no parameters, a baseline of 4 is appropriate as it doesn't need to compensate for gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as 'Get all currently running timers (with task title).' It specifies the verb 'Get' and resource 'running timers' with the additional detail 'with task title.' However, it doesn't explicitly distinguish this from sibling tools like 'time_summary' or 'stop_timer,' which prevents a score of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools like 'time_summary' for aggregated data or 'stop_timer' for active timer management, nor does it specify prerequisites or exclusions. This leaves usage context implied at best.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_notesA
List notes (excerpt only). Use get_note for full content.
| Name | Required | Description | Default |
|---|---|---|---|
| folderId | No | Filter by folder ID. Pass "none" for unfiled. | |
| tag | No | Filter by tag. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool returns 'excerpt only' (a key behavioral trait), but does not mention other aspects like pagination, rate limits, authentication needs, or error handling. The description adds some value but lacks comprehensive behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with zero waste. The first sentence states the purpose and key limitation, and the second provides explicit usage guidance, making it front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (list operation), no annotations, no output schema, and 100% schema coverage, the description is mostly complete. It covers purpose, differentiation, and usage, but lacks details on output format (e.g., structure of excerpts) and behavioral traits like pagination, which would enhance completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters ('folderId' and 'tag') with clear descriptions. The description does not add any parameter-specific information beyond what the schema provides, meeting the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('List') and resource ('notes'), and specifically distinguishes this tool from its sibling 'get_note' by noting it returns 'excerpt only' versus full content. This provides precise differentiation from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('List notes (excerpt only)') and when to use an alternative ('Use get_note for full content'), providing clear guidance on tool selection based on the need for excerpts versus full content.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tasksC
List tasks within a date range. Returns tasks with id, title, status, scheduledDate, tags, etc.
| Name | Required | Description | Default |
|---|---|---|---|
| from | No | Start date (YYYY-MM-DD). Inclusive. | |
| to | No | End date (YYYY-MM-DD). Inclusive. | |
| includeInbox | No | Include tasks with no scheduled date. | |
| tag | No | Filter by tag (comma-separated for multiple). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions the tool 'Returns tasks with id, title, status, scheduledDate, tags, etc.', which adds some behavioral context about output format. However, it lacks critical details like whether this is a read-only operation, pagination behavior, rate limits, authentication needs, or error handling. For a list operation with no annotation coverage, this is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with zero waste. The first sentence states the core action and scope, and the second sentence describes the return format. It's appropriately sized and front-loaded, though it could be slightly more structured (e.g., separating purpose from output details).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 4 parameters with full schema coverage, the description is minimally adequate. It covers the basic purpose and output format but lacks behavioral transparency, usage guidelines, and details on error cases or limitations. For a list tool with no structured safety hints, it should do more to compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 4 parameters thoroughly. The description adds minimal value beyond the schema by implying date-range filtering ('within a date range') and hinting at tag filtering ('tags' in return values). It doesn't provide additional syntax, format details, or usage examples. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('List') and resource ('tasks'), and specifies scope ('within a date range'). It distinguishes from siblings like 'get_running_timers' or 'time_summary' by focusing on tasks with date filtering. However, it doesn't explicitly differentiate from 'list_notes' or other list operations beyond the resource type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_running_timers' for active tasks or 'time_summary' for aggregated data. It mentions date range filtering but doesn't specify prerequisites, exclusions, or typical use cases. The agent must infer usage from the tool name and parameters alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_timerA
Start a timer for the specified task. Stops any currently running timer first.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | Task ID to start tracking. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes a key behavioral trait: 'Stops any currently running timer first,' which is crucial for understanding its side effects beyond just starting a timer. However, it lacks details on permissions, rate limits, or error conditions, leaving some behavioral aspects uncovered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and front-loaded, consisting of only two sentences that directly convey the tool's purpose and key behavior. Every word earns its place, with no redundancy or unnecessary information, making it highly efficient and clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (a mutation tool with side effects), no annotations, and no output schema, the description is somewhat complete but has gaps. It covers the core action and a critical side effect, but lacks information on return values, error handling, or prerequisites, making it adequate but not fully comprehensive for safe agent use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, with the parameter 'taskId' fully documented in the schema. The description does not add any additional meaning or context about the parameter beyond what the schema provides, such as format examples or usage tips. Thus, it meets the baseline score of 3 for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Start a timer') and the target resource ('for the specified task'), with the additional behavioral detail that it 'Stops any currently running timer first.' This distinguishes it from sibling tools like 'stop_timer' and 'get_running_timers' by specifying its unique start-and-stop behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by mentioning it stops any currently running timer, which suggests it should be used when switching tasks or starting fresh timing. However, it does not explicitly state when to use it versus alternatives like 'stop_timer' or 'get_running_timers,' nor does it provide exclusions or prerequisites, keeping it from a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_timerC
Stop the running timer for the specified task.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | Task ID whose timer to stop. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but lacks behavioral details. It doesn't disclose whether stopping a timer is reversible, if it requires specific permissions, what happens to the recorded time, or what the response looks like (e.g., confirmation, error if no timer exists). This is inadequate for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with zero waste. It's appropriately sized and front-loaded, directly stating the tool's purpose without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (a mutation with no annotations and no output schema), the description is incomplete. It doesn't explain behavioral traits, error conditions, or return values, leaving significant gaps for an agent to understand how to use it effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the single parameter 'taskId' fully. The description adds no additional meaning beyond implying the parameter identifies a task with a running timer, which aligns with the schema. Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Stop') and target ('the running timer for the specified task'), providing a specific verb+resource combination. However, it doesn't explicitly differentiate from sibling tools like 'get_running_timers' or 'time_summary', which prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., a timer must be running), exclusions, or relationships with sibling tools like 'start_timer' or 'get_running_timers', leaving usage context unclear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
time_summaryC
Get aggregated time tracking summary for a date range.
| Name | Required | Description | Default |
|---|---|---|---|
| from | Yes | Start date (YYYY-MM-DD). | |
| to | Yes | End date (YYYY-MM-DD). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It mentions aggregation but doesn't specify what data is included (e.g., total hours, by project), whether it's read-only or has side effects, or the format of the summary. For a tool with no annotation coverage, this leaves significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It's appropriately sized and front-loaded with the core functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description is insufficient. It doesn't explain what the aggregated summary includes, the return format, or any behavioral constraints. Given the complexity of time tracking data and lack of structured output documentation, more context is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents both parameters ('from' and 'to' as date strings). The description adds no additional parameter semantics beyond implying date-range filtering, which is already clear from the schema. Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Get aggregated time tracking summary') and resource ('time tracking summary') with scope ('for a date range'), making the purpose immediately understandable. It doesn't explicitly differentiate from siblings like 'get_running_timers' or 'list_tasks', but the aggregation aspect provides some implicit distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_running_timers' or 'list_tasks', nor does it mention prerequisites such as needing existing time tracking data. It only states what the tool does, not when it should be selected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_taskC
Update an existing task. Only provided fields are changed.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Task ID. | |
| title | No | ||
| scheduledDate | No | ||
| tags | No | ||
| estimatedMin | No | ||
| priority | No | ||
| memo | No | ||
| status | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. While it mentions 'Only provided fields are changed' (which is useful partial update behavior), it doesn't address critical aspects like authentication requirements, error conditions, whether the operation is idempotent, what happens with invalid inputs, or what the response looks like. For a mutation tool with zero annotation coverage, this leaves significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise at just two sentences with zero wasted words. It's front-loaded with the core purpose and follows with important behavioral context about partial updates. Every sentence earns its place, making it efficient for an AI agent to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given this is a mutation tool with 8 parameters, no annotations, no output schema, and only 13% schema description coverage, the description is incomplete. It doesn't explain what values are acceptable for parameters like 'priority' or 'estimatedMin', doesn't describe the response format, and doesn't address error handling or side effects. For a tool that modifies data, this level of documentation is insufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds minimal parameter semantics beyond the schema. With only 13% schema description coverage (only the 'id' parameter has a description), the description doesn't explain what the other 7 parameters represent or how they should be used. The phrase 'Only provided fields are changed' provides some context about partial updates but doesn't clarify individual parameter meanings. The baseline is 3 since the schema provides structure, but the description doesn't adequately compensate for the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Update') and resource ('an existing task'), making the purpose immediately understandable. It distinguishes from sibling tools like 'create_task' by specifying it works on existing tasks. However, it doesn't explicitly differentiate from other update-like operations that might exist in the system.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (like needing the task ID), doesn't specify when to use this versus 'create_task' or 'delete_task', and offers no context about appropriate use cases or limitations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.1.0- First observed
create_note - First observed
create_task - First observed
delete_task - First observed
get_note - First observed
get_running_timers - First observed
list_notes - First observed
list_tasks - First observed
start_timer - First observed
stop_timer - First observed
time_summary - First observed
update_task
TDQS
Scored across 11 tools
Each tool has a clearly distinct purpose with no ambiguity. Tools are cleanly separated by resource (note, task, timer) and action (create, get, list, update, delete, start, stop, summary), making it easy for an agent to select the right one. For example, create_note vs. create_task, or start_timer vs. stop_timer.
All tools follow a consistent verb_noun naming pattern throughout, such as create_note, list_tasks, start_timer, and time_summary. There are no deviations in style or convention, making the set predictable and easy to navigate.
With 11 tools, the count is well-scoped for a productivity/note-taking/time-tracking server. Each tool earns its place by covering essential operations without bloat, such as CRUD for notes and tasks, plus timer management and summaries.
The tool set provides near-complete coverage for the domain, including full CRUD for tasks (create, list, update, delete) and notes (create, get, list), plus timer control and summaries. A minor gap is the lack of update_note and delete_note tools, which agents might need to work around, but core workflows are well-supported.
Maintenance
Related MCP Connectors
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
- mcpOAuthnet.todoist
Official Todoist MCP server for AI assistants to manage tasks, projects, and workflows.
An MCP server that provides an API to LLMs to manage their JumpCloud resources.
MCP server for generating rough-draft project plans from natural-language prompts.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceA task management MCP server that provides tools to create, list, complete, and delete tasks using pluggable storage backends. It enables users to interact with their task lists through natural language using MCP-compatible clients like Claude Desktop.-
- AlicenseAqualityDmaintenanceAn MCP server that enables AI assistants to interact with the Taskwarrior command-line task management tool. It allows users to list, create, modify, and organize tasks using projects, tags, and annotations through natural language.132MIT
- AlicenseNot gradedqualityFmaintenanceMCP server for TickTick task management. Enables AI assistants to create, update, list, search, and complete tasks via natural language.53 npm29MIT
- FlicenseNot gradedqualityDmaintenanceA task manager MCP server that demonstrates all three MCP primitives (tools, resources, prompts). Enables users to manage tasks, read task summaries and details, and run structured planning/review prompts through natural language.-