Skip to main content
Glama

Server Details

Task tracker for product teams: projects, features, tasks, bugs, team chats and code context.

Ownership verified
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

A3.6/5.0

Scored across 73 tools

Disambiguation3/5

Many tools target distinct resources, but several read/context tools overlap, such as branch_list vs repo_branches, task_context vs task_code_context vs task_commits, and feature_context vs feature_code_context. Descriptions help, but with 73 tools the boundaries are not always crisp.

Naming Consistency4/5

Names mostly follow a snake_case resource_action or resource_noun pattern with clear domain prefixes. A few outliers like create_rules, evaluation_rules, iteration_key, and file_upload_complete deviate slightly, but the overall convention is readable and predictable.

Tool Count1/5

73 tools far exceeds the well-scoped 3-15 range and crosses the extreme threshold of 50. Even for a broad platform, this creates an overwhelming surface that makes tool selection and navigation difficult for an agent.

Completeness4/5

The surface covers core lifecycle operations for tasks, bugs, features, products, versions, repos, chat, memory, files, and users. Minor gaps exist, such as no explicit delete for versions, modules, components, or products beyond archiving, but core workflows are well supported.

Available Tools

73 tools
branch_listA
Read-only
Inspect

Lists repository branches with the task key from each name (taskKey); taskID keeps only that task's live branches, search filters by a name substring.

ParametersJSON Schema
NameRequiredDescriptionDefault
searchNoBranch name substring.
taskIDNoOnly branches with this task's key.
repositoryIDYesRepository ID.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered. The description usefully adds that each branch name is parsed for a task key and that taskID restricts to 'live' branches, but says nothing about pagination or result size, leaving moderate behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence that front-loads the core action (listing branches) and then the two filter behaviors. No wasted words, though the parenthetical taskKey note is slightly crammed in.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with no output schema and full parameter coverage, the description conveys enough to call it correctly. A brief note on result limits or ordering would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters; the baseline is 3. The description adds mild clarifying detail (taskKey extraction, substring matching) but no syntax or format beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Lists) and resource (repository branches), plus what each entry carries (taskKey). It does not explicitly differentiate itself from the nearby repo_branches sibling, which keeps it just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains how taskID and search behave but gives no guidance on when to reach for branch_list versus repo_branches or other branch-oriented siblings. Usage is only implied from the filter semantics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

branch_nameA
Read-only
Inspect

Returns the branch name for a task by the server's convention and its live branch, if one exists (existing; a task has one live branch). Creates nothing.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIDYesTask ID.
repositoryIDYesRepository ID.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the description adds value by explicitly reinforcing that it 'Creates nothing' and explaining the conditional return of the live branch (only if one exists). It doesn't cover rate limits or auth needs, but for a simple read this is sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Highly concise: two sentences that front-load the core purpose and the key behavioral constraint. The parenthetical and trailing phrase feel slightly awkward but do not waste space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter read-only tool with no output schema, the description covers purpose and key behavior. It could specify the return type or format, but with annotations and full schema support, nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both taskID and repositoryID are fully documented in the schema. The description mentions neither parameter, so it adds no semantic value beyond the structured fields – baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: returns the branch name for a task. It distinguishes the two outputs (server-convention name and live branch) and explicitly clarifies that this is a lookup, not a creation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implicitly indicates when to use: when you need the convention branch name or the task's live branch. However, no alternatives are named and no exclusions or prerequisites are given, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bug_createAInspect

Creates a bug as a draft when the user asks to log a defect; the token owner reviews and publishes it. Placement is one of: a feature (featureID or featureKey), a version (versionID) or the product backlog (productID); type is the criticality (the nature of the defect is the separate bugType field), and bugs always have a testing stage. Returns the key, link, status and allowed transitions.

ParametersJSON Schema
NameRequiredDescriptionDefault
ordNoPosition of the new bug; tasks are ordered by `ord` descending. Omitted - end of the feature's plan, so bugs created one by one keep their order.
gistYesGist: one or two sentences, "problem → solution", up to 300 characters of plain text (no Markdown, links or task keys), in the language of the workspace data.
specYesBug description in Markdown (no HTML) following the bug creation template, in the language of the workspace data.
typeYesCriticality by actual impact: bug_trivial - cosmetic; bug_minor - inconvenience with a workaround; bug_major - a function works wrong or part of a scenario is unavailable, the workaround is costly; bug_critical - a scenario is impossible, data is lost or corrupted, a section crashes; bug_block - blocks the product or all testing. In doubt between two levels - bug_major.
buildNoBuild in which the defect was reproduced, from the deployed version; QA reproduces on it.
titleYesBug title, up to 200 characters, in the language of the workspace data; states the defect itself, not just the area (like "Settings").
featureIDNoFeature ID - the preferred placement
moduleIDsNoIDs of the modules (screens) where the defect is visible, even if its cause is in a service.
productIDNoProduct ID - the product backlog, last resort
versionIDNoVersion ID - when no feature fits
featureKeyNoFeature key #F12 (or F12, 12); not with featureID
componentIDsNoIDs of the components where the defect is actually visible; empty if unsure.
testAssigneeNoTester: employee UUID or "any" (any employee); default "any" - bugs are always tested.
workAssigneeNoAssignee: employee UUID, "any" (any employee) or null (no work stage); default "any".
workTimePlanNoPlanned fix time in minutes, for the assignee's role level.
rulesRevisionYesRevision of the bug creation template the spec follows.
reviewAssigneeNoReviewer: employee UUID, "any" (any employee) or null (no review stage); default: no review.
approvalAssigneeNoApprover: employee UUID, "any" (any employee) or null (no approval stage); default: no approval.
workIterationKeyNoIteration key for work outside the current iteration; default - current

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnlyHint=false / destructiveHint=false annotations by disclosing that the object is created as an unpublished draft requiring owner review, that bugs always have a testing stage, and that the response returns key, link, status and allowed transitions. These are meaningful lifecycle facts an agent cannot infer from annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action, trigger and draft lifecycle, then placement, the `type` caveat and return values — every clause carries information. It is dense and would read better with light structuring, but there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 19-parameter write tool with no output schema, the description covers the creation workflow, placement choices, criticality semantics and even the return payload, while the 100%-covered schema carries individual field detail. Complete enough to call correctly, though the review/approval stage behavior is left entirely to field descriptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantic value by disambiguating `type` (criticality) from the separate `bugType` (nature of the defect) and by framing the placement parameters as a preference order. It stops short of explaining the remaining timing/assignee parameters, which the schema handles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Creates a bug") and immediately scopes it as a draft that a token owner later publishes, which clearly separates it from sibling write tools like task_create. The distinct 'bug' resource plus the draft/review workflow makes the tool's identity unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ("when the user asks to log a defect") and lays out the placement options (feature, version, product backlog). It stops short of naming an alternative tool or stating when not to use this one, so it is clear context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chat_message_deleteA
Destructive
Inspect

Deletes a chat message. Irreversible: the token owner can delete their own message within 60 minutes, chat administrators and the product lead any message. Intended for an explicit user request to remove an unneeded message (duplicate, wrong chat).

ParametersJSON Schema
NameRequiredDescriptionDefault
messageIDYesMessage ID

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond annotations by disclosing irreversibility, the permission model (token owner within 60 minutes; chat administrators and product lead any message), and the operational scope. Annotations only flag destructiveHint=true; the description provides actionable constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the action and irreversibility warning, then permissions and intended use. No filler or repetition; every clause adds useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter destructive tool with annotations covering safety hints and no output schema requirement, the description supplies irreversibility, permission boundaries, timing, and intended use. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single messageID parameter, including format and description. The description adds no parameter-specific syntax or constraints, so the baseline of 3 applies for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Deletes') and resource ('chat message'), clearly distinguishing it from sibling operations like chat_message_update and chat_send. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the intended context: an explicit user request to remove an unneeded message, with examples ('duplicate, wrong chat'). This implies when not to use it (unsolicited/autonomous deletion), though it does not name an alternative tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chat_message_updateAInspect

Edits the token owner's own regular message within 60 minutes of sending: text replaces the old one, mentions are recalculated from mention tokens, passed attachments replace the old ones ([] removes all). The chat shows an "edited" mark; later edits are not possible.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesFull new text (Markdown), in the language of the workspace data
mentionsNoNot needed: mentions come from mention tokens in text
messageIDYesMessage ID
attachmentsNoNew attachment set, up to 5; replaces the old one entirely

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond annotations: text fully replaces the old content, mentions are recalculated from mention tokens, attachments are wholesale replaced ([] clears all), an 'edited' mark appears, and only one edit is permitted. These replacement semantics and the edit-window limit are exactly the facts an agent needs before committing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the edit constraint and then enumerates replacement behavior with no filler. Every clause carries operational meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering the safety profile, a 100%-covered schema, and no output schema expected to be explained, the description supplies the remaining gaps: ownership, time window, replacement semantics, and irreversibility. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3, but the description adds the '[] removes all' convention for attachments and clarifies that mentions are derived from tokens rather than passed explicitly. The full-replacement semantics for text and attachments also give context the schema only implies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Edits) and resource (the token owner's own regular message) with an explicit scope restriction that separates it from chat_send and chat_message_delete. An agent can identify the operation and its boundary without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete use conditions: only the owner's own message, only within 60 minutes of sending, and later edits being impossible tells the agent this is a one-shot operation. It does not name alternative sibling tools for the 'too late' or 'not my message' cases, so it stops short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chat_question_inboxA
Read-only
Inspect

Returns messages addressed to the token owner (direct chat, mention, reply to their message) still without their answer, regardless of age: chat and link, why it is addressed, the related item, the message with attachments. Useful when the user asks what awaits their reply; pages continue with nextCursor while hasMore.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoPage size, 1-100 (default 20)
cursorNoPrevious page's nextCursor, as is

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered. The description adds real behavioral context beyond that: results are age-independent ('regardless of age'), and it discloses the returned payload shape (chat/link, why addressed, related item, message with attachments). Pagination continuation via nextCursor/hasMore is also stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is one dense run-on sentence, but it is front-loaded with purpose, then usage intent, then paging behavior. Every clause earns its place, though the parenthetical enumeration makes it heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description explicitly describes what a returned item contains (chat and link, why it is addressed, related item, the message with attachments), compensating for the missing schema. Combined with the paging behavior and the read-only annotation, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so limit and cursor are already documented, and the baseline is 3. The description's pagination note refers mostly to the response fields (nextCursor/hasMore) rather than adding syntax or constraints to the parameters themselves, so it does not clearly exceed baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Returns) and resource (messages addressed to the token owner still unanswered), and enumerates the trigger types — direct chat, mention, reply — that qualify. An agent can distinguish this from chat_read and chat_send without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear triggering condition: 'Useful when the user asks what awaits their reply.' This maps the tool to an intent. However, it names no alternative (e.g., chat_read) and gives no exclusions, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chat_reactionAInspect

Adds or removes the token owner's reaction to a chat message - a textless reply when a reaction says it unambiguously: agreement, "got it, will do", "thanks", "looking", "agreed". Only Kosmodrom web reactions are available (meanings in reaction); an existing reaction is not added twice, system messages take no reactions, and message reactions are keyed by who added them.

ParametersJSON Schema
NameRequiredDescriptionDefault
removeNotrue - remove the token owner's reaction
reactionYesthumbs-up - 👍 agree, accepted; thumbs-down - 👎 disagree; ok - 👌 got it, will do; handshake - 🤝 agreed (deal); hands-folded - 🙏 thanks; 100p - 💯 fully agree; face-smile - 🙂 glad; face-lol - 😂 funny; face-thinking - 🤔 need to think, doubtful; face-grimacing - 😬 awkward; face-screaming - 😱 alarm; face-angry - 😠 displeased; fire - 🔥 excellent; eyes - 👀 looking, taking a look; call - 📞 let's have a call; lady-bug - 🐞 this is a bug; crutch - hack, temporary workaround; no-bicycles - don't reinvent the wheel, use an existing solution
messageIDYesMessage ID

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false, leaving the real behavior to the description. It adds substantial rules beyond them: no duplicate reaction, system messages take no reactions, reactions are keyed by the reacting user, and only Kosmodrom web reactions are supported.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the add/remove purpose and then packs the constraints efficiently. It is a single long compound sentence rather than cleanly separated, but every clause earns its place and none is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and full annotation coverage of only the safety flags, the description supplies the edge cases (duplicate suppression, system-message exclusion, keying by user) an agent needs. It could go slightly further on remove-specific behavior, but is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters, including the full reaction enum meanings that the description merely points to ('meanings in reaction'). The description adds little parameter-level detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (adds/removes), resource (the token owner's reaction), and target (a chat message). It clearly distinguishes itself from chat_send/chat_message_update by being the textless-reaction tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the appropriate context for a reaction (a textless reply when the meaning is unambiguous) with concrete examples. It stops short of naming alternative siblings (e.g. chat_send) or giving explicit when-not-to-use conditions, so it is strong but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chat_readA
Read-only
Inspect

Returns chat messages from a cursor, oldest to newest; without a cursor - the latest ones. hasMore tells whether more messages exist in the requested direction; an after cursor returns newer messages, such as replies to a sent one.

ParametersJSON Schema
NameRequiredDescriptionDefault
afterNoRead messages newer than chat.after.
limitNoMessages to return, 1-100 (default 30)
beforeNoRead messages older than chat.before.
chatIDYesChat ID; for a task's chat - the task ID

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply readOnlyHint=true, so the description carries real additional weight: it discloses ordering (oldest to newest), pagination direction semantics, and the meaning of the hasMore response field. It adds no rate-limit, auth, or size caveats, keeping it short of 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the primary behavior and then the hasMore semantics; nothing is padded. The semicolon-heavy phrasing is slightly dense but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully covers the key return signal (hasMore) and pagination behavior, which is what an agent needs to iterate. It does not describe the message object shape, so it is not fully complete, but the essential pagination contract is there.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description genuinely adds meaning by explaining that after returns newer messages in the requested direction and how the boundaries relate to ordering. This clarifies the cursor parameters beyond their terse schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Returns chat messages') and qualifies it by cursor direction, so an agent can immediately tell it apart from write-side siblings like chat_send or chat_message_update. It stops short of naming any sibling explicitly, so it does not reach the top of the scale.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operational context: no cursor returns the latest messages, an after cursor walks forward to newer ones (e.g. replies), implying before for older. It never states when NOT to use it or points to an alternative tool, so it lacks the exclusions a 5 would require.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chat_sendAInspect

Sends a message on behalf of the token owner to a chat (a task's chat ID is the task ID) or, with userID, a direct message to an employee; replyMessageID makes it a reply. Intended for blocking questions, replies to messages addressed to the token owner and messages the user asks for - not for status updates, summaries or links to created items. Returns chatID, the message and an after cursor; the message can be edited or deleted within 60 minutes.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesMessage in the language of the workspace data: Markdown without headings, tables or images; bare URLs become links. Tasks and features go in as their link value, employees as their mention token where addressed
chatIDNoChat ID; for a task's chat - the task ID
userIDNoEmployee ID for a direct message, instead of chatID
mentionsNoNot needed: mentions come from mention tokens in text
attachmentsNoUp to 5 attachments: each is an uploaded file's fid or name with dataBase64
idempotencyKeyNoRetry-safety key: one key per chat yields one message
replyMessageIDNoID of the message in this chat to reply to

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the write-but-non-destructive profile is covered. The description adds genuinely new behavior: it returns chatID, the message and an after cursor, and the message is editable/deletable within a 60-minute window. It omits rate limits or failure modes, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence front-loads the action, then the targeting variants, then the negative guidance and return contract. No sentence is wasted, though the packed clauses around replies and returning cursors make it slightly heavy to parse in one pass.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 7 parameters, 100% schema coverage and no output schema, the description supplies the missing return shape (chatID, message, after cursor) and the 60-minute edit window. Nothing an agent needs to select or invoke this tool correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including the idempotencyKey retry semantics is already documented in the schema. The description's notes that chatID equals a task ID and that userID is the DM alternative restate what the schema already says, rather than adding syntax or format meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (sends a message on behalf of the token owner) and immediately distinguishes the three targeting modes: chatID, userID direct message, and replyMessageID. An agent can tell this apart from siblings like chat_read, chat_reaction, or chat_message_update without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('blocking questions, replies to messages addressed to the token owner, and messages the user asks for') and explicit when-not ('not for status updates, summaries or links to created items'). This is precisely the routing guidance an agent needs among ~70 siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

component_createAInspect

Creates a product component. Only at the user's request or with their consent: the product structure belongs to the team.

ParametersJSON Schema
NameRequiredDescriptionDefault
descNoComponent description, Markdown
titleYesComponent name, in the language of the workspace data
productIDYesProduct ID

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the mutation/safety profile is covered. The description adds a meaningful behavioral constraint beyond the annotations: it must only be invoked at the user's request or with consent, because the product structure belongs to the team. It does not cover validation errors, permissions, or what happens to existing structure, but it adds real context the annotations do not.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and followed by the consent constraint. No wasted words and nothing important buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter create tool with full schema coverage and no output schema, the description is largely sufficient, and the consent rule fills the main behavioral gap. It does not mention required fields, the language/workspace convention for title, or what a successful call returns, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% – title, desc, productID each carry their own description in the schema. The description adds no parameter-level syntax, format, or constraint information, so the baseline of 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Creates a product component'), and the sibling set contains both component_update and module/feature create tools, so the resource is clearly distinguished. An agent can immediately tell this is a creation tool for the component entity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit precondition: 'Only at the user's request or with their consent,' which is real guidance about when to invoke it, not just what it does. It does not name a sibling alternative (e.g., component_update) or state when not to use it, so it falls short of the top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

component_updateAInspect

Updates a component's name and description or removes its main Figma frame; omitted fields stay unchanged, module membership is not affected. Only at the user's request or with their consent: the product structure belongs to the team.

ParametersJSON Schema
NameRequiredDescriptionDefault
descNoNew description, Markdown
titleNoNew name
designNonull - remove the component's main Figma frame
componentIDYesComponent ID

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds genuine behavioral context beyond annotations: partial-update semantics (omitted fields stay unchanged) and that module membership is unaffected. However, the description says the tool can remove the component's main Figma frame while annotations declare destructiveHint=false, a tension that is not reconciled; no permission or reversibility detail is given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the primary operation and the effects on omitted fields; only two clauses before the consent caveat. Tight, though the second sentence about product-structure ownership is slightly editorial rather than operational.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers scope (what changes, what does not) and the consent requirement, which is enough for correct invocation. It falls short only on permission prerequisites and on clarifying the destructiveHint tension around frame removal.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every field is documented, including the design=null removal semantics, so the schema does the heavy lifting. The description adds only the general 'omitted fields unchanged' rule, not per-parameter meaning. Baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (updates) plus resource (component) and enumerates exactly what is mutable: name, description, or removal of the main Figma frame. It also clarifies what is NOT touched (module membership), which cleanly separates it from module_update and component_create.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides one explicit usage condition – only act at the user's request or with their consent – which is real guidance, but never names alternatives or says when a sibling (component_create, module_update) should be chosen instead. Implied rather than complete routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_rulesA
Read-only
Inspect

Returns the workspace's conventions and description template for a new feature, task or bug (content) and the template revision (rulesRevision) that creation requests carry. For a task, the template depends on its work type.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYesWhat is being created: feature, task or bug.
taskTypeNoThe task's work type; required when entity is "task", not passed for feature and bug.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already declares this is a safe read, so the safety burden is lifted. The description adds the useful detail of exactly what is returned (content + rulesRevision) and the task/work-type dependency, but omits any format, failure mode, or authorization context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences that lead with the return payload and follow with the conditional task rule. No filler, though the phrasing is dense enough to read slightly awkwardly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully names the returned fields (content, rulesRevision) and the work-type conditional. For a 2-param read tool this is close to sufficient; a pointer to the consuming create tools would make it complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both enums are self-documenting, so the schema carries the load. The description's note that the task template depends on work type adds marginal motivation for taskType but no syntax beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it returns workspace conventions, a description template, and a rulesRevision for feature/task/bug creation, which is a specific verb+resource. This notably clarifies a misleading tool name ('create_rules' implies a write). However, it names no sibling to distinguish itself from, e.g., evaluation_rules or the *_create tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is fetched before creating a feature/task/bug ('templates that creation requests carry'), giving usable context. But it never explicitly says when to call it versus the create tools, nor any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

currency_convertA
Read-only
Inspect

Converts an amount between RUB, USD, EUR and CNY at the Bank of Russia rate for a date: non-business days use the last business day's rate, cross rates go via RUB, dates after tomorrow are rejected. Returns rate (units of to per 1 from) and result (the amount in to, 2 decimals).

ParametersJSON Schema
NameRequiredDescriptionDefault
toYesTarget currency
dateNoRate date, YYYY-MM-DD; defaults to today in the company's time zone
fromYesCurrency of the amount
amountYesAmount in the from currency

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only carry readOnlyHint=true, but the description adds the rate source (Bank of Russia), holiday/weekend fallback behavior, cross-rate routing via RUB, and an explicit error condition (future dates rejected). That is rich behavioral disclosure well beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Core purpose is front-loaded in the first clause, edge-case rules follow, and the return contract is stated last. No filler; every clause conveys a distinct rule or behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description compensates by documenting the return shape (rate and result, with result rounded to 2 decimals). Combined with the date and error rules, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning: it defines rate as 'units of to per 1 from', removing ambiguity about rate direction, and reinforces the date fallback rule that the schema only alludes to. The 'from'/'to' enums are already in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (converts) and resource (an amount between RUB, USD, EUR, CNY) with the exact rate source. It is instantly distinguishable from every sibling tool, none of which touch currency.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete operating context: which currency set is supported, that non-business days fall back to the last business day, that cross rates route via RUB, and that dates after tomorrow are rejected. No sibling overlaps so no alternative needs naming, but it stops short of an explicit 'use this when...' framing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_comment_replyAInspect

Replies in a Figma comment thread through the company's integration, visible to everyone with access to the file; the server adds an "employee and agent" signature. Source: url or designID; commentID - the thread or a reply in it. Suitable when a remark is fixed or the answer is unambiguous; disputed points are left to the user.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoFigma link to the file, page or frame with the thread
messageYesReply in the language of the workspace data, up to 4000 characters: what was fixed and where; the server adds the signature
designIDNoImported design ID, instead of url
commentIDYesThread ID (thread.ID) or a reply ID in it

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare write (readOnlyHint=false) and non-destructive; the description adds real behavioral context beyond them: the reply is visible to everyone with file access and the server appends an 'employee and agent' signature. It does not cover failure/error behavior or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and its side effect, then usage guidance, in three dense clauses. Information is efficient, though the semicolon-packed phrasing is slightly cramped.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with no output schema, it covers visibility, the server-added signature, source options, and usage conditions. Missing only edge-case handling such as invalid commentID or file-access failures.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters. The description reinforces the url-or-designID source choice and that commentID may be a thread or a reply, but adds no syntax or format detail the schema lacks, so baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (replies) and resource (Figma comment thread) plus the integration context. It is clearly distinguishable from sibling read tools like design_comments without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly scopes when to use it ('when a remark is fixed or the answer is unambiguous') and when not to ('disputed points are left to the user'). It stops short of naming the alternative surface (e.g., design_comments for reading), so it falls just below a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_commentsA
Read-only
Inspect

Returns Figma design comments through the company's integration: threads with author, date, text, replies and pinned node. Source: url (a frame or page returns threads inside it, a file - all) or designID of an imported design; author.isIntegration marks replies already sent via Kosmodrom. A later reply may cancel a remark; the Figma API cannot mark threads resolved.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoFigma link to a file, page or frame
statusNoopen - unresolved threads only (default); all - including resolved
designIDNoImported design ID: a designs key or the uuid in design:<uuid>

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=true already covering safety, the description adds genuinely non-obvious context: author.isIntegration flags replies sent through Kosmodrom, a later reply can cancel a remark, and the Figma API cannot mark threads resolved. These are operational caveats an agent could not derive from annotations, though return format/pagination is untouched.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, front-loaded with purpose and result shape, then source semantics and caveats. No filler; the telegraphic 'Source: ...' clause is terse but carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description compensates by listing the fields of each returned thread and by documenting the integration caveats. The one gap is that with zero required parameters, it does not explicitly state that url or designID must be supplied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so baseline is 3, but the description adds real semantics: for url it explains that a frame or page returns only threads inside it while a file returns all, and for designID it indicates an imported design identifier. Only the status enum default is left entirely to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns Figma design comments') and enumerates the returned structure (threads with author, date, text, replies, pinned node). It also implicitly separates this read tool from the design_comment_reply sibling by framing the returned data rather than any write action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the reader can infer 'read comments' vs 'reply to comments' from sibling names, but the description never states when to use this versus design_comment_reply or design_import. It does clarify the source-selection rule (url vs designID), which is useful but is parameter guidance rather than when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_importAInspect

Imports Figma frames into a task, bug or feature through the company's Figma integration, like pasting a design in the web app (snapshot, layers, assets): one entity, up to 10 frame links with node-id per request; a failed frame does not stop the others. Per frame returns status (imported, updated, unchanged, linked - linked but not refreshed, error), designID and a designToken marker, or error with action - how to fix it; the designToken in the entity's spec or result is its only link to the design. moduleID / componentID set the module's or component's main frame (one frame, no marker); re-import refreshes the snapshot and keeps markers.

ParametersJSON Schema
NameRequiredDescriptionDefault
framesYesFrames to import
taskIDNoTask or bug ID
taskKeyNoTask or bug number, when the ID is unknown
moduleIDNoModule ID - its main frame (one frame)
featureIDNoFeature ID
featureKeyNoFeature key #F12 (or F12, 12); not with featureID
componentIDNoComponent ID - its main frame (one frame)

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only flag readOnlyHint=false and destructiveHint=false. The description adds substantial context beyond that: partial-failure semantics ('a failed frame does not stop the others'), the full per-frame status set, the designToken as the only design link, and re-import refresh behavior that preserves markers. This is rich, non-obvious behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action is front-loaded, but the description is a single dense compound sentence with semicolon chains that is hard to scan. Most clauses earn their place, yet the structure could be broken into clearer segments.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by documenting per-frame return values (status enum values, designID, designToken, error/action). For a 7-parameter tool it covers targets and behavior well, though it omits prerequisites such as whether the target entity must already exist.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: moduleID/componentID set the module's or component's main frame and are constrained to one frame with no marker. That explains the otherwise opaque parameters beyond their schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Imports) and resource (Figma frames) plus the target entity types (task, bug or feature), with an analogy to pasting a design in the web app. An agent can distinguish this from sibling design_comment_reply and design_comments without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives context ('into a task, bug or feature through the company's Figma integration') which implies when the tool applies, but never names alternatives or states when NOT to use it. Usage is inferable rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluation_rulesA
Read-only
Inspect

Returns the work evaluation rules for a task type and stage (procedure, rating scale, comment requirements) and their rulesRevision; the server picks the variant. Useful when the user asks to evaluate work on a task, e.g. "evaluate #12345" or "evaluate the review of #12345".

ParametersJSON Schema
NameRequiredDescriptionDefault
stageYesStage: work - execution (default), review, or testing.
taskTypeYesTask type; for a bug, its criticality.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered; the description then adds a genuinely useful non-obvious behavior: 'the server picks the variant', telling the agent the caller cannot select which rule set comes back. It also notes rulesRevision is returned, though it says nothing about error cases when no rules exist for a type/stage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, with the return payload front-loaded before the usage example. The second sentence is slightly dense but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly enumerates the return contents (procedure, rating scale, comment requirements, rulesRevision). Combined with 100% schema coverage and readOnlyHint, an agent has what it needs, though the missing-rule case is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with both parameters enum-documented, so the baseline is 3. The description only restates the two axes (task type, stage) already fully described in the schema and adds no format or edge-case semantics beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Returns) and resource (work evaluation rules) plus the key discriminator (task type + stage) and names the returned contents. It is clear what the tool does, though it never explicitly contrasts itself with close siblings such as task_evaluate or task_evaluation_context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete trigger conditions with user-phrase examples ('evaluate #12345', 'evaluate the review of #12345'), which is clear usage context. It stops short of naming alternatives or when NOT to use it, so it falls short of the 5 bar.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

feature_code_contextA
Read-only
Inspect

Returns a feature's code: its tasks, the product's repositories, live branches and commits by task keys. Useful for checking a feature's spec against the implementation.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoFeature key #F12 (or F12, 12); not with featureID
featureIDNoFeature ID; not with key

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safe-read profile is covered. The description adds genuine value by disclosing the returned payload categories (tasks, repos, live branches, commits), but says nothing about auth needs, result size, or how the two identifiers are resolved. With annotations carrying safety, 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the capability front-loaded and the use case appended; nothing is wasted or redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully sketches the return contents, and the read-only safety profile comes from annotations. It is complete enough to call correctly, though a note on how key vs featureID resolve to one feature would close the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (key and featureID, plus their mutual exclusivity and accepted key formats) are already fully documented. The description adds no parameter-level meaning, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (returns) and resource (a feature's code) and enumerates the contents: tasks, repositories, live branches, and commits by task keys. Assuming 'code' here does not literally mean source, it is distinguishable from plain feature_context. It does not explicitly contrast itself with the very similar task_code_context sibling, so it stops short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The sentence 'Useful for checking a feature's spec against the implementation' gives an implied usage scenario, but there is no explicit when-to-use guidance, no exclusions, and no routing among near-neighbors like feature_context, task_code_context, or repo_context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

feature_contextA
Read-only
Inspect

Returns a feature by ID or key (#F12): spec, gist, plan tasks (tasks, top to bottom by ord) and bugs as compact lines, employees (users), designs and attachments (attaches) with file IDs, and specRevision - the spec revision needed to edit the spec. Useful when the user asks about a feature, its plan or who works on it (workAssigneeID).

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoFeature key #F12 (or F12, 12); not with featureID
featureIDNoFeature ID; not with key

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already declares this a safe read, so the description's job is to add context — and it does: it discloses the exact field set returned and explains that specRevision is the revision needed to edit the spec, an implicit concurrency hint for a subsequent write. It does not describe failure behavior when neither key nor featureID is supplied, which is the remaining gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense but well-structured sentence that is front-loaded with the core action and then enumerates return fields in priority order. No filler, though the inline parenthetical aliases (workAssigneeID, attaches, ord) make the single sentence heavier than ideal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries full responsibility for return values and discharges it thoroughly. The only unaddressed case is that both parameters are optional with no stated default or error behavior when neither is provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are self-documenting ('not with featureID' / 'not with key'), so the schema already carries the semantics. The description repeats the #F12 key convention but adds no new parameter guidance, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns a feature by ID or key') and then enumerates the exact payload: spec, gist, plan tasks ordered by ord, bugs, employees, designs/attachments with file IDs, and specRevision. This level of return-content detail makes it distinguishable from siblings like feature_list or feature_search even though those names are not cited.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger condition ('Useful when the user asks about a feature, its plan or who works on it'), which tells the agent when this is the right call. It stops short of naming an alternative (feature_list vs feature_search vs feature_context) or stating exclusions, so it does not reach a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

feature_createAInspect

Creates a feature in a product version when the user asks to add one; it is visible to the team at once, with no approval step and no tasks. Returns the feature's link, product, version, ord and specRevision.

ParametersJSON Schema
NameRequiredDescriptionDefault
ordNoPosition in the version; features are ordered by `ord` ascending. Omitted - end of the version.
gistYesGist: one or two sentences, "problem → solution", up to 300 characters of plain text (no Markdown, links or task keys), in the language of the workspace data.
specNoFeature spec in Markdown (no HTML) following the feature creation template, in the language of the workspace data.
titleYesFeature title in the language of the workspace data, describing the user story rather than the implementation.
analyticNoFeature analyst: employee UUID or "any".
productIDYesProduct ID
versionIDNoVersion ID; omitted - the first version in development, or the backlog.
rulesRevisionYesRevision of the feature creation template the spec follows.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations establish the write/non-destructive profile, and the description adds real behavioral context: the feature is immediately team-visible, has no approval step, and spawns no tasks. It also discloses the return shape (link, product, version, ord, specRevision), compensating for the absent output schema. It does not cover permissions, idempotency, or the spec/rulesRevision relationship, keeping it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the action front-loaded and no filler. It is slightly run-on, packing side-effect disclosure and return values into one semicolon-joined clause, but every element earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-param mutation tool with no output schema, the description covers side effects and return values well. It is slightly short on the required-parameter relationship (title/gist/rulesRevision/productID) and how spec relates to rulesRevision, leaving minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 8 parameters (including defaults for ord and versionID) are already documented in the schema. The description adds no parameter-level meaning beyond what the schema provides; baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource with scope: 'Creates a feature in a product version.' This is unambiguous against siblings like feature_update, feature_delete, feature_list, and even the differently-scoped task_create. An agent can select this without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'When the user asks to add one' gives a clear triggering context for use. However, no alternatives or exclusions are named (e.g., when to use feature_update or branch/task creation instead), so it stops short of the explicit when/when-not routing a 5 requires.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

feature_deleteA
Destructive
Inspect

Deletes a feature with all its tasks and bugs (soft cascading deletion, no restore). Irreversible; only at the user's explicit request. Returns deletedTasks, the number of tasks and bugs deleted.

ParametersJSON Schema
NameRequiredDescriptionDefault
featureIDYesFeature ID.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only flag destructiveHint=true; the description adds the far more important specifics: this is a soft cascading deletion affecting child tasks and bugs, it is irreversible with no restore, and it returns a deletedTasks count. The soft-delete-plus-no-restore nuance is exactly the kind of detail annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: cascade scope first, then the irreversibility/usage constraint, then the return value. No filler or repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter destructive tool with no output schema, the description covers cascade semantics, irreversibility, usage constraint, and even the return value. Nothing an agent needs in order to call it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (featureID) and schema description coverage is 100%, so the schema already documents it fully. The description adds nothing about the parameter beyond the schema, which is the baseline 3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Deletes) and resource (a feature) plus the cascade scope (with all its tasks and bugs), which cleanly separates it from siblings like task_delete, bug_delete, or feature_update. An agent can distinguish it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"only at the user's explicit request" gives a clear activation condition and warns against autonomous use. It does not, however, name an alternative (e.g., feature_update or deactivating/archiving) for cases where deletion is too strong, so the routing guidance stops short of explicit alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

feature_listA
Read-only
Inspect

Lists the features of a product or one version in the web order (ord ascending), for an overview of what a product contains. spec and specRevision come only with includeSpec (much larger response), completed features only with allowDone; the next page is requested with nextCursor and the same parameters.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoFeatures per page, up to 200 (default 50).
cursorNonextCursor of the previous page of the same request.
allowDoneNoInclude completed features; default false
productIDYesProduct ID
versionIDNoVersion ID; omitted - all versions.
includeSpecNotrue - include each feature's spec and specRevision; default false.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only assert readOnlyHint=true, but the description adds real behavioral context: results are `ord`-ascending, spec/specRevision only appear with includeSpec which yields a 'much larger response', completed features are gated by allowDone, and pagination requires reusing the same parameters with nextCursor. It does not spell out the returned shape, but that is beyond a read-only list tool's usual burden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence front-loads the core action and ordering before covering the optional flags and pagination. No filler, though cramming purpose, options, and paging into one sentence makes it slightly dense to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still covers ordering, filtering, payload-size implications, and the paging loop needed to use the tool correctly. For a read-only list endpoint it is essentially complete, missing only a note on default result size or explicit non-use cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the description adds meaning the schema does not: it discloses the response-size tradeoff of includeSpec, that allowDone filters to completed-only, and crucially that the next page must be fetched with nextCursor using the same parameters. That pagination contract is genuine added value over the field definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Lists the features of a product or one version') plus ordering ('`ord` ascending'), which lets an agent distinguish a bulk-listing operation from the sibling `feature_search`. It stops short of explicitly naming the search sibling as the alternative, so it is clear but not fully disambiguated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"for an overview of what a product contains" implies the intended browse/overview use case, and it notes when to add includeSpec or allowDone. However, it never states when NOT to use this tool or names feature_search/feature_context as the alternative for targeted lookups, so routing guidance remains implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

feature_spec_syncA
Read-only
Inspect

Returns the context for updating a feature's spec from its tasks: feature.spec, tasks[].spec and tasks[].result in one batch plus the workspace's spec-update guidelines (instructions); writes nothing. includeTaskIDs adds full contexts with chats of tasks with discrepancies. mode=analyze - differences and a full new version without saving; mode=apply - when the user asks to apply the update.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoFeature key #F12 (or F12, 12); not with featureID
modeYesanalyze - propose without saving; apply - the user asks to apply
basisNoactual (default) - actual state from the tasks; requirements - confirmed in chats
chatLimitNoLatest messages per full context, 1-100; older via chat.before
featureIDNoFeature ID; not with key
taskTypesNoTask and bug types to include; default frontend, backend, fullstack, design
includeTaskIDsNoTasks with discrepancies to add with full context and chat, up to 20; requires taskTypes, the same as in the previous response

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, and the description reinforces it with 'writes nothing' while enumerating exactly what is returned (specs, task results, workspace guidelines). It adds useful return-shape context beyond the safety hint, though it doesn't discuss pagination or limits on the batch size.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the primary purpose, then details of return content and mode behavior in a compact block. Slightly dense and partially redundant with the mode enum in the schema, but every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully spells out the returned payload (feature.spec, tasks[].spec, tasks[].result, instructions) and the analyze/apply distinction, which is essential for a 7-parameter tool. Minor gaps remain around defaults for basis and chatLimit that live only in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented in the schema. The description largely restates mode semantics and the includeTaskIDs behavior ('full contexts with chats of tasks with discrepancies') with little added meaning beyond the schema text, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and resource: returns the context (feature.spec, tasks[].spec, tasks[].result, guidelines) needed to update a feature's spec. The feature-level scope clearly separates it from the sibling task_spec_sync, so an agent can pick between them without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit mode-based usage: 'mode=analyze - differences and a full new version without saving' versus 'mode=apply - when the user asks to apply the update.' That is a clear when-to-use rule, though it does not name an alternative tool or state exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

feature_updateAInspect

Updates a feature's title, spec, gist, version (moves it) or position (ord) at the user's request; only passed fields change. Changes shared team data; closing a feature is a human decision, not available here. Returns the link, product, version, ord, specRevision and changed fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoFeature key #F12 (or F12, 12); not with featureID
ordNoPosition in the version, as drag-and-drop in the web UI; features go by `ord` ascending (tasks the opposite). Between two neighbours - a value between their ord; first - below the smallest; 0 - end of the version. Neighbours closer than 100 - the whole version, closed included, is renumbered top down from Date.now() - (N - 1) * 10000 in steps of +10000 for N features.
gistNoGist: one or two sentences, "problem → solution", up to 300 characters of plain text (no Markdown, links or task keys), in the language of the workspace data; an empty string clears it.
specNoNew spec in Markdown (no HTML), in the language of the workspace data. Replaces the whole field: existing text, `::: doc` sections and `attach:`, `design:`, `color:` tokens stay only if included.
titleNoNew name
featureIDNoFeature ID; not with key
versionIDNoVersion ID to move the feature to.
designUnlinkNotrue - the new spec may drop design markers (`design:`), and designs left unlinked are deleted; otherwise losing a marker is rejected.
expectedSpecRevisionNoThe feature's current specRevision; required when spec changes.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say it is a non-read-only, non-destructive write; the description adds that it mutates shared team data and that partial updates are the norm, plus that closure is out of scope. It does not mention the design-deletion side effect of designUnlink, which is only visible in the schema, but it goes beyond the annotations on the mutation model.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the field list and the partial-update rule, and every sentence earns its place. The single long opening sentence is dense, which slightly hurts scanability, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter mutation with no output schema, the description covers the mutable fields, the partial-update semantics, the closure exclusion, and even the return payload (link, product, version, ord, specRevision, changed fields). Selection and invocation are fully supported; only the design-unlink side effect remains outside the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description contributes the partial-update contract ('only passed fields change') and the semantic framing of version as a move and ord as a position, which the schema does not state as plainly. That is real added meaning on top of the already thorough schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Updates) and resource (a feature) and enumerates the mutable fields including the non-obvious moves (versionID) and ordering (ord). An agent can separate this from feature_create, feature_delete, and feature_list without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit scope ('only passed fields change') and an explicit exclusion ('closing a feature is a human decision, not available here'), which is exactly the kind of when-not guidance that helps selection. It stops short of naming an alternative tool for the excluded operation, so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

feed_work_queueA
Read-only
Inspect

Returns a person's work queue in a product, as on their home page in the web UI: work-stage tasks assigned to the token owner or to userID (review, testing and approval are not included). Four lists in priority order - urgentBugs, urgentTasks, bugs, tasks - each with the next item first; includePool adds the shared "any" pool (poolBugs, poolTasks) matched to the person's role and level. Useful when the user asks what to work on next.

ParametersJSON Schema
NameRequiredDescriptionDefault
userIDNoWhose queue to show (employee ID); defaults to the token owner.
productIDYesProduct ID
includePoolNoAdd the shared "any" pool (poolBugs/poolTasks); default false.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true; the description adds real behavioral context beyond that — the four-list priority ordering, the next-item-first convention, and that includePool entries are matched to the person's role and level. It doesn't cover pagination or result-size limits, but it materially extends what the annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense paragraph, front-loaded with what is returned, then scoping rules, then the usage cue. Every sentence carries information, though the packing of list names and pool names into a single block is slightly heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates well by spelling out the four list names and their ordering plus the optional pool lists. It omits the shape of individual queue items and any size/limit behavior, so it is close to but not fully complete for a tool with no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds semantics the schema lacks: includePool pulls a shared 'any' pool (poolBugs/poolTasks) filtered by the person's role and level, and the priority/ordering meaning of the returned lists. userID's default-to-token-owner is already in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource and goes further by enumerating the exact return contents (urgentBugs, urgentTasks, bugs, tasks) and the exclusion of review/testing/approval items. An agent can distinguish this from task_search or task_context without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the trigger context: 'Useful when the user asks what to work on next.' The exclusion note also implicitly routes agents away from this tool for review/testing/approval work. No alternative tool is named by name, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_getA
Read-only
Inspect

Returns a file's content by fid: an image as an image, text as text, up to 5242880 bytes; larger files, video, archives and other binaries are not returned inline. An attach token without an entry in attaches means the attachment was deleted or the text copied.

ParametersJSON Schema
NameRequiredDescriptionDefault
fidYesFile ID: an attaches key and the fid in attach:<fid>
entityNoOptional owner entity

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true, but the description adds real behavioral context: the 5242880-byte inline cap, which content types are inlined, and the edge case where an attach token has no attaches entry (deleted attachment or copied text). The last sentence is terse but discloses a failure mode the schema does not.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with the core behavior (what comes back and the size/type limit) front-loaded, followed by the edge case. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly spends its budget describing the return shape (image as image, text as text, size cap, non-inline types) and one failure mode. It still leaves unclear what the caller actually receives for a non-inline file, which is the main remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the fid description already explains the 'attach:<fid>' relationship, so the schema carries the parameter burden. The description adds no format or constraint detail for the optional entity object beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns a file's content by fid') and immediately scopes it with return-format rules, which separates it in spirit from file_info and file_url. It never names those siblings, so the differentiation is inferable rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear when-not condition: larger files, video, archives and other binaries are not returned inline, so an agent knows this tool won't work for those cases. It does not route the agent to an alternative (file_url, file_info) for the non-inline cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_infoA
Read-only
Inspect

Returns file metadata by fid: name, MIME type, kind, size, upload date and uploader - for a fid without known metadata; task and feature attachments already carry it in attaches.

ParametersJSON Schema
NameRequiredDescriptionDefault
fidYesFile ID: an attaches key and the fid in attach:<fid>
entityNoOptional owner entity

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already establishes that this is a safe, non-mutating read. The description's one added behavior detail is the set of fields returned, which matters because there is no output schema; it says nothing about permissions, rate limits, or behavior when the fid is unknown.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded sentence with no filler and the purpose leads. The trailing clause about attaches is dense and jargon-heavy but still earns its place as routing guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, listing the returned fields compensates well, and the read-only annotation covers the safety profile. Both parameters are fully documented in the schema, so an agent has what it needs, though return format details (e.g. null when unknown) are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including the nested entity object, so the structured data already carries parameter meaning. The description only echoes 'by fid' and adds no format, constraint, or entity-scoping nuance beyond the schema, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource (returns file metadata by fid) and enumerates the returned fields (name, MIME type, kind, size, upload date, uploader). It does not explicitly differentiate itself from the close siblings file_get, file_url, or file_upload, so an agent must still infer which file tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear use condition ('for a fid without known metadata') and an exclusion ('task and feature attachments already carry it in attaches'), which is real routing guidance. It stops short of naming the alternative tools an agent would otherwise consider, e.g. file_get or file_url.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_uploadAInspect

Uploads a file for attachment: attachToken goes into a task, bug or feature spec / result, fid into chat message attachments; one fid can be attached in several places. With dataBase64 (up to 2097152 bytes) the file is saved at once; without it, a file of any size up to maxSize is PUT to uploadURL (curl -T "") and the upload is then completed. Files cannot be deleted through MCP.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesFile name with extension, e.g. "screenshot.png"; the extension sets the type
dataBase64NoFile content in base64, up to 2097152 bytes; omitted for upload via URL

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only mark it non-read-only and non-destructive; the description adds material behavior: max upload sizes, two-phase URL upload, attachment multiplicity, and a hard limitation that files cannot be deleted through MCP. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose and dense but purposeful sentences; every clause (token destinations, size limits, upload modes, deletion limitation) carries operational weight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-mode upload tool with no output schema, the description supplies the key attachment identifiers and uploadURL flow. It falls short by not naming file_upload_complete as the completion step and by not fully specifying the return shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for both parameters, and the schema already states the base64 size limit and the URL-upload alternative. The description reinforces the workflow but adds little parameter-level meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource, then immediately distinguishes the two attachment identifiers (attachToken vs fid) and the two upload modes. An agent can tell this is the upload primitive rather than file_upload_complete or file_get, though the sibling is not named.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the dataBase64 path for small files and the uploadURL PUT path for larger files, including size limits. However it does not explicitly name file_upload_complete as the final step or provide conditions for choosing among sibling file tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_upload_completeAInspect

Completes an upload via URL after a successful PUT to uploadURL: verifies the stored size and returns file and attachToken for a spec / result, and the fid for chat attachments. Repeating it returns the same response.

ParametersJSON Schema
NameRequiredDescriptionDefault
fidYesFile ID of the upload

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare a non-read, non-destructive write, but the description adds real behavior: it verifies stored size and is idempotent on repeat. Those traits beyond the annotations (verification step, safe re-invocation) are exactly the value-add expected here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the trigger condition front-loaded and the return/idempotency facts trailing. No wasted filler, though the return-value clause is packed tightly enough to require a careful read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the return values (file and attachToken for spec/result, fid for chat attachments) plus idempotency, which compensates for the missing output contract. Coverage is good for a one-param completion tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is a single documented 'fid' parameter, so the schema carries the semantics. The description adds no syntax or format detail for fid, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Completes an upload via URL') and pins its place in a multi-step flow ('after a successful PUT to uploadURL'). An agent can distinguish it from the upload-initiating sibling (file_upload) via the flow role, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear activation condition: call it only after a successful PUT to uploadURL. It also gives idempotency guidance ('Repeating it returns the same response'), reducing retry risk. No explicit when-not or named alternative, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_urlA
Read-only
Inspect

Returns a signed, time-limited download URL and metadata for a file by fid - for video, archives and other files not returned inline.

ParametersJSON Schema
NameRequiredDescriptionDefault
fidYesFile ID: an attaches key and the fid in attach:<fid>
entityNoOptional owner entity

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, while the description adds that the URL is signed and time-limited and returns metadata. It does not state expiration duration or permission requirements, but it adds meaningful behavioral context beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with purpose front-loaded and no filler. Every element, including return type, file scope, and exclusion case, earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only URL-retrieval tool with no output schema, the description covers the return object (signed URL and metadata) and the use case. The optional entity parameter's effect is left to the schema, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description says 'by fid' but adds no syntax or meaning beyond the schema, and it does not discuss the optional entity parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'Returns' and resource 'signed, time-limited download URL and metadata for a file by fid', and scopes it to video, archives, and other non-inline files. This distinguishes it from sibling tools like file_get and file_info by explaining that it serves files not returned inline.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: use for video, archives, and other files not returned inline. It implies file_get or file_info are alternatives for inline-accessible files, but does not explicitly name when not to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

iteration_keyA
Read-only
Inspect

Returns the iteration key for a date, today by default: the workIterationKey value for planning a task or bug outside the current iteration.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoISO date; defaults to today

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds one useful behavioral detail, the 'today by default' fallback, but says nothing about format of the returned key or any failure conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the core action front-loaded. The trailing colon clause is slightly dense but every element carries meaning; no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter, read-only lookup with no output schema, the description covers what it returns and the default date behavior. Nothing critical is missing, though naming the return format would round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter already documents 'ISO date; defaults to today'. The description only restates the default, adding no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns the iteration key for a date') and explains what the value is (workIterationKey). The scope is concrete enough that an agent can distinguish it from unrelated siblings, though there is no explicit sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'for planning a task or bug outside the current iteration' implies the usage context, giving the agent a scenario. However, it never states when not to use it or names an alternative, leaving usage to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_createAInspect

Saves a confirmed fact, rule or recipe as an agent memory record when the user asks to keep knowledge for future agent work; secrets only at the user's direct request. The record starts personal and unverified: it applies to the token owner's agents at once, and a moderator can make it shared. level is the scope: company (Git, cloud, CI), user (personal preferences), project (stands, addresses, project infrastructure) or repository (conventions, tests, code pitfalls).

ParametersJSON Schema
NameRequiredDescriptionDefault
jobsNoJobs the record is for; at least one for job and request, none for always.
loadYesalways - delivered with any work, short; job - with the jobs in `jobs` (prevention, silent pitfalls); request - only in the `memory` index, read by ID (symptom diagnostics, narrow areas).
textYesOne topic, 1.5-3 thousand characters (max 4,000), in the language of the workspace data: rules and facts with reasons. Author and date are record fields, no "Source" line needed.
levelYes
titleYes"<area> · <situation or symptom>: <what's inside>", up to 200 characters, in the language of the workspace data; agents judge relevance by it.
projectIDNoProject ID, for level = project.
repositoryIDNoRepository ID, for level = repository.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false. The description adds real behavioral context beyond that: the record starts personal and unverified, immediately applies to the token owner's agents, and a moderator can later make it shared. Auth/rate-limit details are absent, but the lifecycle disclosure is genuinely useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action before the visibility and `level` details. Dense but every clause (trigger, secret caveat, visibility lifecycle, level semantics) earns its place; only minor cost is that the level list is inline rather than scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter creation tool without an output schema, the description covers the trigger, the safety caveat, the visibility lifecycle, and the one ambiguous required parameter's semantics, while the schema already documents the remaining fields. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 86% (above the 80% baseline of 3), but the description goes further by spelling out the four `level` values with concrete scope examples (company/Git/cloud/CI, user/preferences, project/infrastructure, repository/conventions) — meaning the schema's bare enum does not carry. This is a real addition over structured data.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: "Saves a confirmed fact, rule or recipe as an agent memory record." Combined with the trigger condition (user asks to keep knowledge), an agent can distinguish this create tool from the sibling memory_list/read/update/delete tools without inspecting any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear when-to-use ("when the user asks to keep knowledge for future agent work") plus an explicit restriction ("secrets only at the user's direct request"). It does not name sibling alternatives such as memory_update for editing an existing record, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_deleteA
Destructive
Inspect

Deletes an agent memory record at the user's request: authors their own, moderators (memory:moderate) any but user-level ones. Rejected if the record changed since the given version.

ParametersJSON Schema
NameRequiredDescriptionDefault
IDYesRecord ID.
versionYesRecord version the change is based on.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say destructive/not-read-only; the description adds real behavioral context beyond that: a two-tier permission model (author vs memory:moderate) and a version-based rejection rule that implies optimistic concurrency control. These are the exact traits an agent needs to call a destructive tool safely.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the action front-loaded and no filler. The compressed phrasing ('authors their own, moderators ... any but user-level ones') is efficient but requires a second read, which costs it the top mark.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, fully documented, destructive operation with no output schema, the description covers authorization and the failure mode well. It stops short of stating whether deletion is permanent/recoverable and what the response returns, which would matter for an irreversible action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description gives the version parameter a functional role beyond the schema's 'Record version the change is based on' — it is a concurrency token whose mismatch causes rejection. That meaningfully upgrades what the agent knows about the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Deletes) and resource (agent memory record) and scopes it to a user-initiated action, which cleanly separates it from memory_update, memory_read, and memory_create. An agent knows exactly what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides the governing authorization context: authors may delete their own records, moderators with memory:moderate may delete any but user-level ones, and the call is rejected under optimistic-concurrency conflict. It does not explicitly route between memory_delete and memory_update (e.g. prefer update over delete-and-recreate), so it falls short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_listA
Read-only
Inspect

Lists agent memory records visible to the user, with filters and usage statistics: deliveredCount (delivered in full), readCount (read by ID), lastUsedDT. Useful for reviewing, moderating and cleaning up memory: with memory:moderate it covers all company records except personal ones, verified = false is the moderation queue, sort = usage with unusedDays finds rarely used records.

ParametersJSON Schema
NameRequiredDescriptionDefault
IDsNoRecord IDs, up to 50.
jobsNoRecords for at least one of these jobs.
loadNo
sortNousage - least used first.
levelNo
queryNoSubstring of the title or text, in the language of the workspace data.
sharedNo
verifiedNofalse - unverified records (the moderation queue).
authorUIDNo
projectIDNoProject ID: memory of the project and its connected repositories.
unusedDaysNoNot delivered to or read by agents for this many days.
repositoryIDNoRepository ID.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered. The description adds real behavioral context beyond that: the permission (memory:moderate) changes the visible scope, and it discloses the usage-stat fields returned (deliveredCount, readCount, lastUsedDT). This is meaningful for a list tool with no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then usage and parameter recipes. The second sentence is dense with several clauses, but each clause (moderate scope, moderation queue, cleanup pattern) contributes distinct, actionable information rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter list tool with no output schema, the description covers the conceptual model (record visibility), the key returned stats, and the two dominant workflows. It is close to complete; minor gaps are parameters like load/shared/level that neither schema nor description illuminate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, so several parameters are already documented. The description adds interpretation the schema lacks: verified=false identifies the moderation queue, and the sort=usage + unusedDays pairing is explained as a rarely-used-records finder. It leaves some params (load, shared, level) implicit, keeping it below a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Lists agent memory records visible to the user') and adds scope via 'with filters and usage statistics'. It is distinguishable from siblings like memory_read (single record by ID) and memory_create/delete, though it doesn't name those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete when-to-use context: 'reviewing, moderating and cleaning up memory', with the moderator pattern (memory:moderate scope, verified=false = moderation queue) and the cleanup pattern (sort=usage with unusedDays). It stops short of naming the alternative sibling (memory_read) for single-record retrieval.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_readA
Read-only
Inspect

Returns agent memory records and MCP rules in full by ID (rules by slug, e.g. git; memory records by UUID). Useful when a record from the memory index is relevant to the user's task; unavailable IDs come in notFound.

ParametersJSON Schema
NameRequiredDescriptionDefault
IDsYesRule slugs or memory record UUIDs from the `memory` index, up to 50.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint already declares the safe-read profile, but the description adds genuinely useful behavior beyond annotations: unavailable IDs surface in the `notFound` field rather than erroring, and rule lookups use slugs while records use UUIDs. That is real operational context not available in structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with what is returned and how it is keyed, followed by when it is useful. No filler; the notFound detail earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with no output schema, the description covers input formats, scoping, and the partial-failure behavior (notFound). Return shape beyond that is unspecified, but nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single `IDs` parameter is already documented, so baseline is 3. The description adds meaning above the schema by explaining the two ID formats (rule slugs vs. memory-record UUIDs) and the example `git`, clarifying what values are actually valid.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Returns) and resource (agent memory records and MCP rules) with the key differentiator: full retrieval by ID, contrasted with the `memory` index used for discovery. An agent can distinguish this from memory_list without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context ('Useful when a record from the `memory` index is relevant to the user's task'), which implicitly positions this as the follow-up to index/search tools. It does not explicitly name an alternative or state when not to use it, so it falls just short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_updateAInspect

Updates an agent memory record at the user's request; passed fields are replaced, and a new level, projectID or repositoryID moves it. Authors edit their own records, moderators (memory:moderate) any but user-level ones; content edits make the record personal and unverified again, and verified and shared (shared needs verified) are a human moderator's decision. Rejected if the record changed since the given version.

ParametersJSON Schema
NameRequiredDescriptionDefault
IDYesRecord ID.
jobsNoJobs the record is for; at least one for job and request, none for always.
loadNoalways - delivered with any work, short; job - with the jobs in `jobs` (prevention, silent pitfalls); request - only in the `memory` index, read by ID (symptom diagnostics, narrow areas).
textNoOne topic, 1.5-3 thousand characters (max 4,000), in the language of the workspace data: rules and facts with reasons. Author and date are record fields, no "Source" line needed.
levelNoNew level (moves the record).
titleNo"<area> · <situation or symptom>: <what's inside>", up to 200 characters, in the language of the workspace data; agents judge relevance by it.
sharedNoRequires memory:moderate: shared with all company agents (true) or personal to the author's (false).
versionYesRecord version the change is based on.
verifiedNoRequires memory:moderate: the record is verified.
projectIDNoNew project ID, for level = project.
repositoryIDNoNew repository ID, for level = repository.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations covering only readOnlyHint=false and destructiveHint=false, the description carries the real behavioral load: partial-update semantics, level changes relocating the record, permission tiers (author vs memory:moderate), the side effect that content edits reset the record to personal and unverified, the human-moderator gate on verified/shared, and optimistic-concurrency rejection on stale version. This is materially more than the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single dense semicolon-chained block, but the purpose is front-loaded and nearly every clause carries distinct information (semantics, permissions, side effects, concurrency). It could be broken into shorter sentences for scanability, but there is little filler to cut.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter mutation tool with no output schema, it covers permissions, side effects, partial-update semantics, and concurrency adequately. The main omission is what the call returns (e.g., the new version for subsequent updates), which matters for a version-guarded write.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every field, setting the baseline at 3. The description still adds cross-parameter meaning the schema does not: that passed fields are replaced (partial update), that level/projectID/repositoryID jointly move the record, and that shared/verified require moderator rights.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Updates an agent memory record') and immediately scopes the semantics ('passed fields are replaced, and a new level, projectID or repositoryID moves it'). It is unambiguous against memory_read/list/delete/create by name and content, though it never explicitly names a sibling to distinguish itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operating context: the update is performed 'at the user's request', authors may edit their own records, moderators (memory:moderate) may edit any but user-level ones, and the call is rejected on version conflict. It stops short of explicitly naming alternatives (e.g., memory_create for new records), so it is clear context without full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

module_createAInspect

Creates a product module. Only at the user's request or with their consent: the product structure belongs to the team.

ParametersJSON Schema
NameRequiredDescriptionDefault
descNoModule description, Markdown
titleYesModule name, in the language of the workspace data
parentIDNoParent module ID; omitted - a root module
productIDYesProduct ID

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the write-but-non-destructive profile is covered. The description adds genuinely non-structured behavior: a consent/authorization precondition tied to shared team ownership, which the agent could not infer from the schema or annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero redundancy, and the core action is front-loaded before the consent caveat. Nothing could be trimmed without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 4-parameter create tool with full schema coverage and no output schema, the description covers purpose and the consent gate adequately. It leaves minor gaps such as the post-create result and whether a parent module must already exist.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: title, desc, parentID and productID are each documented, including that omitting parentID yields a root module. The description adds nothing about parameter format or defaults beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Creates a product module'), which clearly distinguishes it from module_update and product_modules. It never names a sibling explicitly, so it stops just short of the 5 bar.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit usage condition: only at the user's request or with their consent, framing the product structure as team-owned. It gives clear context but names no alternative tool or exclusion (e.g., when to prefer module_update), so it is clear-but-incomplete rather than exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

module_updateAInspect

Updates a module's name, description, parent or components, or removes its main Figma frame; omitted fields stay unchanged. Only at the user's request or with their consent: the product structure belongs to the team.

ParametersJSON Schema
NameRequiredDescriptionDefault
descNoNew description, Markdown
titleNoNew name
designNonull - remove the module's main Figma frame
moduleIDYesModule ID
parentIDNoNew parent module ID; null - make it a root module
addComponentIDsNoExisting components to add to the module
removeComponentIDsNoComponents to remove from the module

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=false. The description adds real context beyond that: patch semantics ('omitted fields stay unchanged'), the consent requirement, and the fact that the main Figma frame can be removed. It does not clarify whether frame removal is reversible or what happens to components on the removed branch.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact clauses with the functional scope front-loaded and the governance caveat second, no wasted words. The policy sentence is slightly tangential to the mechanical behavior but is short and justified by the safety framing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Seven parameters, no output schema, and a mutation surface spanning rename, reparent, component add/remove, and frame removal. The description covers the scope and merge semantics adequately; the only unfilled gap is interaction between addComponentIDs and removeComponentIDs, which the schema also leaves unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description earns above baseline by stating the merge/patch contract ('omitted fields stay unchanged'), which is semantic information the per-field schema does not convey and which materially changes how an agent builds the call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (Updates) and resource (module) and enumerates the exact fields it can touch: name, description, parent, components, and removal of the main Figma frame. This is clearly distinguishable from module_create and product_modules in the sibling list without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit usage gate: only act at the user's request or with their consent, with the rationale that the product structure belongs to the team. It stops short of naming alternative tools or enumerating when-not cases, but the consent condition is a concrete, actionable constraint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

product_contextA
Read-only
Inspect

Returns a product's working context: project, spec excerpt (full spec with includeSpec), product analyst and product lead, unfinished versions with defaultVersionID (the default for new work), repositories and design files. With job, it also returns the workspace conventions and team notes for that kind of work (memoryInstructions) and the memory index of notes, as reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
jobNoKind of work: analysis - analysis and specs; dev - code and automated tests (local repositories only); qa - manual testing; devops - CI, builds, deployment; review - code review; design - mockups
productIDYesProduct ID
includeSpecNotrue - full product description in spec; default - first paragraph in specExcerpt
localRepositoriesNoIDs of the product's repositories cloned locally

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply readOnlyHint=true; the description goes further by disclosing conditional behavior — passing 'job' additionally returns workspace conventions, team notes (memoryInstructions) and the memory index — and by distinguishing specExcerpt from the full spec under includeSpec. It stops short of describing the shape/format of the returned context or any limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences that lead with what is returned and then layer the conditional job behavior. Every clause carries information, though the parenthetical-heavy second sentence is slightly packed for a single breath.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of describing returns, and it does so well for the default case plus the job-augmented case. Missing only guidance on when this tool is the right one versus the other *_context siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains that includeSpec toggles between excerpt and full spec, that job triggers extra return sections, and it labels defaultVersionID as the default for new work. It does not mention the localRepositories parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns a product's working context') and enumerates the concrete contents (project, spec excerpt, analyst/lead, unfinished versions, repositories, design files). It is clearly a retrieval tool, but it never distinguishes itself from siblings in the same family such as feature_context, task_context, or repo_context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer you call this to orient on a product before doing work. There is no explicit when-to-use, when-not-to-use, or comparison against the sibling context tools, and no stated prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

product_listA
Read-only
Inspect

Lists projects with their active products (archived products are excluded); projectID narrows the list to one project. Useful for choosing the product to work in; by default it is a project's first active product.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIDNoOptional: return only this project's products.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true. The description adds real behavioral context beyond that: archived products are excluded, and the default return is a project's first active product. This helps the agent predict output shape without contradicting the safe-read annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compound sentence that leads with what is listed and then qualifies scope and defaults. Efficient, though the semicolon-joined clauses are slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, single-parameter list tool with no output schema, the description covers the essentials: what is returned, what is filtered out, and the default selection. Sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so projectID is already fully documented as an optional filter. The description restates that it 'narrows the list to one project' but adds no extra syntax or format meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Lists projects with their active products') and adds scoping detail (archived products excluded). It does not explicitly differentiate itself from nearby siblings like product_context or product_modules, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Useful for choosing the product to work in' provides an implied use case, giving some context for when the tool applies. However there are no explicit conditions, exclusions, or named alternatives (e.g., when to prefer product_context), leaving usage largely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

product_modulesA
Read-only
Inspect

Returns the product structure for tagging tasks and bugs: modules (componentIDs link to the shared components list) and components; a module does not imply its components. desc and design (the main Figma frame: snapshotFID - snapshot, contextFID - layers) are returned only for items in IDs; query filters by name.

ParametersJSON Schema
NameRequiredDescriptionDefault
IDsNoModule and component IDs to return with desc and design
queryNoCase-insensitive module or component name substring, in the language of the workspace data
productIDYesProduct ID

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already covers the safety profile, so the description's value is in the extra rules it discloses: desc and design are returned only for items listed in IDs, a module does not imply its components, and query is a case-insensitive name substring in the workspace language. These are non-obvious behavioral constraints that the annotations do not carry. It omits list-size/pagination behavior beyond the schema's maxItems hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single dense block, front-loaded with the return contents, which is good. However it runs on with parenthetical field mappings ('snapshotFID - snapshot, contextFID - layers') that read like raw schema notes rather than agent-facing prose, hurting scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does carry the return-shape burden and describes modules, components, desc, and design, plus the IDs-conditioned inclusion rule. It leaves the pagination/limit story (maxItems 50 on IDs, any cap on returned modules) and the snapshotFID/contextFID meanings somewhat opaque.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds genuine semantics: IDs gates whether desc/design payloads are populated, and query filters by name case-insensitively in the workspace data language. That is meaningfully more than the schema's terse field text, though the cryptic 'snapshotFID - snapshot, contextFID - layers' shorthand is hard to parse.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource: 'Returns the product structure for tagging tasks and bugs: modules ... and components.' That clearly separates it from siblings like product_context, module_create, or component_create. It stops short of naming alternatives explicitly, but the resource scope is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for tagging tasks and bugs' implies the use context, and the IDs/query clauses hint at retrieval patterns, but there is no explicit when-to-use, when-not-to-use, or alternative tool named (e.g., product_context or component_update). Usage must be inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

product_updateAInspect

Changes a product (only the passed fields) and returns its updated context; isArchived moves it to the archive with its versions, features and tasks, taking them out of work selection. Changes shared product structure; only at the user's request or with their consent.

ParametersJSON Schema
NameRequiredDescriptionDefault
specNoNew product spec in Markdown (no HTML), in the language of the workspace data. Replaces the whole field: existing text, `::: doc` sections and `attach:`, `design:`, `color:` tokens stay only if included.
titleNoNew name
designNoProduct design files (Figma links); passed fields are replaced.
leadUIDNoNew product lead: employee UUID.
analyticNoNew product analyst: employee UUID or "any".
platformNoNew product platform.
productIDYesProduct ID
techStackNoNew technology stack; replaces the previous one.
isArchivedNotrue - archive the product, false - restore it.
isDeployEnabledNoWhether the product's tasks have a deployment stage.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false, so the description carries the burden of what actually happens. It discloses that only passed fields change and that isArchived pulls versions, features and tasks out of work selection. Archiving is described as reversible (restore), so this does not contradict destructiveHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the partial-update qualifier are front-loaded, and every clause (field scope, return context, archive side-effects, consent rule) adds information. It is a dense run-on sentence with semicolons rather than cleanly separated sentences, which slightly hurts scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter mutation with a nested design object, 100% schema coverage, and no output schema, the description covers the essentials: partial semantics, the returned context, the archive side-effect, and the consent constraint. It omits nothing an agent would need to call it correctly, though naming the read-only sibling would complete the picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and already documents per-field replacement semantics, so the baseline would be 3. The description earns the extra point by explaining the cross-system consequence of the isArchived parameter (versions/features/tasks leave work selection), which the schema does not state.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Changes a product') and qualifies the mutation as partial ('only the passed fields'), which is meaningful precision for an agent. It does not name or differentiate against siblings such as product_context or product_list, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit permission condition for the risky path: 'Changes shared product structure; only at the user's request or with their consent.' That is real when-to-use guidance. It names no alternatives (e.g., read-only product_context) and states no exclusions, so it is not a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_blameA
Read-only
Inspect

Returns line-by-line blame of a file at a ref: ranges (lines attributed to a commit sha) and commits (title, message with task keys, author, dates). Blame at the parent of a fix commit shows the lines as they were before the fix. Up to 204800 bytes (large files in ranges via startLine/endLine); renames are not tracked.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoBranch, tag or sha; defaults to the default branch.
pathYesFile path from the repository root.
endLineNoLast line of the range, inclusive.
startLineNoFirst line of the range (1-based).
repositoryIDYesRepository ID.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint, and the description meaningfully extends that: it discloses the output shape (ranges attributed to a commit sha plus commit metadata), a 204800-byte cap with the startLine/endLine workaround, and the 'renames are not tracked' limitation. Missing pagination/error behavior, but the added operational constraints are substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences with the core purpose front-loaded and no filler. The output-format clause is packed tightly, which slightly strains readability but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining returns and does so (ranges, commits with title/message/author/dates), plus byte limits and rename behavior. Auth/prerequisite context is absent but not critical for a read-only query.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter is already documented at baseline 3. The description adds value by explaining the purpose of startLine/endLine (range reads for large files) beyond the schema's mechanical definition, slightly clarifying their role.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns line-by-line blame of a file at a ref') and even outlines the return structure (ranges and commits). It is easy to distinguish from generic list/search tools, though it never names the closest sibling (repo_line_history) to draw a hard boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Offers one concrete usage scenario ('blame at the parent of a fix commit shows the lines as they were before the fix'), which implies when the tool is useful, but gives no explicit when-to-use/when-not guidance or routing against sibling tools like repo_line_history or repo_history.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_branchesA
Read-only
Inspect

Lists repository branches with their sha and default/protected flags. search filters by a name substring, e.g. "#26388" finds the live branch of that task.

ParametersJSON Schema
NameRequiredDescriptionDefault
searchNoBranch name substring.
repositoryIDYesRepository ID.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered. Beyond that, the description discloses return content (sha, default/protected flags) which is the only source of return information since no output schema exists, adding real behavioral value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with what the tool returns, followed by the filter semantics. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only listing tool with full schema coverage, the description covers purpose, filter semantics, and return fields. The one gap is that it never routes the agent away from the closely named branch_list/branch_name siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds meaning beyond the schema's terse 'Branch name substring' by showing that a task ID like '#26388' resolves to that task's live branch — a non-obvious matching semantic the schema does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Lists repository branches') plus the attributes returned (sha, default/protected flags). However, it never distinguishes itself from the sibling tools branch_list and branch_name, leaving the agent to guess which branch-listing tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the search example ('#26388' finds the live branch of that task) hints at when to use the filter, but there is no explicit when-to-use statement, no exclusions, and no mention of the alternative branch_list sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_compareA
Read-only
Inspect

Returns the diff between two refs without truncation: by default from the merge base of from and to, as in an MR; straight compares with from directly (e.g. a previous review's sha). Patch lines carry new-version line numbers; pages are up to 153600 bytes (next page: offset = nextOffset until null), and a file that does not fit comes with error plus baseSha and toSha.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYesRef to compare: task branch, tag or sha.
fromNoBase: branch, tag or sha; defaults to the default branch.
pathNoDirectory prefix or file glob.
offsetNoPage offset: nextOffset of the previous response.
contextNoContext lines in the patch; default 3.
straightNoCompare with from directly, not with the merge base.
extensionsNoFile extensions without the dot.
lineNumbersNoNew-version line numbers; default true.
repositoryIDYesRepository ID.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses critical runtime behavior: no truncation, pages capped at ~153600 bytes, the pagination contract (offset = nextOffset until null), that patch lines carry new-version line numbers, and that an oversized file yields an error together with baseSha and toSha. This is exactly the operational detail an agent cannot get from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense sentence front-loads the core purpose, then layers pagination and error semantics in the order an agent needs them. Every clause carries information, though the run-on structure makes it heavier to parse than a split would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining return shape, and it does so for pagination, line numbering and the oversized-file error case. Edge cases such as binary-content handling or a maximum page count are not covered, so it is strong but not exhaustive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is already 100%, so 3 is the floor, but the description adds meaning the schema does not: how from/to interact (merge base vs straight), what lineNumbers controls in the output, and how offset is driven by nextOffset. The path, context and extensions parameters get no extra prose, which keeps it short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns the diff between two refs'), immediately distinguishing it from read-oriented siblings such as repo_history, repo_blame or repo_search. The scope is further pinned down by naming the default comparison basis (merge base of from and to, as in an MR).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete guidance on when to flip the straight flag: 'as in an MR' for the merge-base default versus 'a previous review's sha' for direct comparison. It does not name alternative sibling tools or state exclusions, but the usage context for the key mode switch is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_contextA
Read-only
Inspect

Returns an overview of a product repository: git provider, URL, default branch and its sha (defaultBranchSha), linked projects, agent access level, the start of the README (up to 16384 bytes) and the technical passport (stack, architecture, conventions; stale - may be outdated). Useful for an unfamiliar repository.

ParametersJSON Schema
NameRequiredDescriptionDefault
repositoryIDYesRepository ID.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply readOnlyHint=true, so the description carries the rest of the burden and does so credibly: it discloses the README is truncated to 16384 bytes and that the technical passport is stale and may be outdated. That freshness/truncation context is exactly what an agent needs to interpret results, though nothing is said about behavior for missing repositories or access-denied cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense but front-loaded sentence with no filler; the payload contents come first and the use-case cue last. Slightly list-heavy, but every enumerated field earns its place in telling the agent what will come back.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must describe the return payload and it does so comprehensively, including the staleness caveat. It falls short only on error/edge-case behavior, which is a minor gap for a read-only overview tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with a single parameter, so the schema already documents repositoryID fully. The description adds no format, constraint, or scoping detail for the parameter, making 3 the correct baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Returns an overview') plus the exact contents of that overview: provider, URL, default branch and sha, linked projects, agent access level, README head, technical passport. This clearly separates it from granular siblings like repo_file, repo_tree, and branch_list, which return raw repository contents rather than a summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The closing clause 'Useful for an unfamiliar repository' implies when to reach for it, but no alternative is named (e.g. repo_file/repo_tree for targeted lookups) and there are no exclusions or prerequisites. Usage is hinted rather than guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_fileA
Read-only
Inspect

Reads files at a ref: one file (path) or up to 20 parts (files) - whole files, line ranges or declarations (symbol: a function, class, method or type with its leading comment, found heuristically). Up to 204800 bytes per response, so large files are read in ranges via startLine/endLine (totalLines is returned); a failed part carries its own error without affecting the others.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoBranch, tag or sha; defaults to the default branch.
pathNoSingle file path from the repository root.
filesNoSeveral parts: whole files, line ranges or declarations by name.
symbolNoDeclaration name in the path file.
endLineNoLast line of the range for path, inclusive.
startLineNoFirst line of the range for path (1-based).
lineNumbersNoInclude line numbers; default true.
repositoryIDYesRepository ID.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond readOnlyHint=true by disclosing the 204800-byte response cap, the resulting need to paginate large files via startLine/endLine, that totalLines is returned, that symbol lookup is heuristic, and that a failed part carries its own error without affecting siblings. These are precisely the behavioral traits an agent needs and that annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence, but front-loaded with the core capability and the mode choice before drilling into limits and error handling. Every clause carries information; only the nested parenthetical about 'symbol' is slightly heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully covers returns (totalLines) and partial-failure behavior, and it addresses the byte cap that governs pagination. It stops short of describing the response shape (e.g. how content and errors are structured per part), which is the main remaining gap for an 8-parameter read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real semantic value: it explains the path-vs-files duality (one file vs up to 20 parts), clarifies that symbol means a declaration 'with its leading comment' found heuristically, and ties startLine/endLine to the byte-limit workaround. That exceeds what the schema's terse field descriptions convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Reads files at a ref') and immediately enumerates the two modes (single 'path' vs up to 20 'files' parts) and the three part types (whole file, line range, declaration). This distinguishes it clearly from siblings like repo_search, repo_blame, and repo_history.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when each mode applies (single file vs batched parts; ranges when files are large), but never explicitly says when to prefer this over file_get, repo_context, or repo_search. Usage is inferable from the mechanics rather than stated as guidance or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_historyA
Read-only
Inspect

Returns the commit history of a ref, file or directory, like git log, paginated: message, author, dates, parents, task keys from the message (taskKeys) and linked tasks (tasks); hasMore marks a next page.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoBranch, tag or sha; defaults to the default branch.
pageNoPage number, from 1.
pathNoFile or directory from the repository root; defaults to the whole ref.
perPageNoCommits per page; default 30.
repositoryIDYesRepository ID.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered, yet the description adds real behavioral context the annotations lack: pagination behavior, the hasMore next-page marker, and the set of returned fields including taskKeys and linked tasks. This is genuinely additive for a read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence front-loads the purpose and then appends the return/pagination details with no filler. It is dense (the parenthetical field list) but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the returned fields and the pagination flag, and annotations cover the safety profile. Parameters are fully covered by the schema, leaving little an agent must guess to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents ref, path, page, perPage and repositoryID in detail. The description only echoes 'ref, file or directory' and adds no syntax or format guidance beyond the schema, so the baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns the commit history of a ref, file or directory') and reinforces it with the 'like git log' analogy, which cleanly separates it from line-level siblings such as repo_line_history and repo_blame. It stops short of naming a sibling it is not, so it sits at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'like git log' framing implies the usage context, but no explicit when-to-use/when-not or alternative routing is given (e.g., vs repo_line_history, repo_compare, or task_commits). Usage is reasonably inferable but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_line_historyA
Read-only
Inspect

Shows who changed lines startLine-endLine of a file and in which task, from commits synced to Kosmodrom across all branches, newest first (linkType: message - task key in the message, branch - commit on the task branch; match=path - no line ranges for that commit). Lines shift with later edits, so results are candidates; only task-linked commits are included, up to 100 (truncated beyond).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesFile path from the repository root in the current version.
limitNoCommits in the response; default 30.
endLineNoLast line of the range, inclusive.
startLineNoFirst line of the range (1-based); without a range, all changes to the file.
repositoryIDYesRepository ID.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=true already covering the safety profile, the description goes well beyond it: it discloses truncation ('up to 100 (truncated beyond)'), ordering ('newest first'), the candidate/approximate nature of results due to line shifting, and the semantics of linkType (message vs branch) and match=path. This is exactly the behavioral context an agent needs before trusting the output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the opening clause, followed by scoping constraints, then caveats about result reliability. It is a single dense sentence with stacked parentheticals, which costs a little readability, but there is essentially no filler text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must characterize returns — it does so reasonably (who, which task, newest first, cap of 100). It stops short of describing the returned fields or pagination/limit interplay with the default of 30, which is a minor remaining gap for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning the schema lacks — the linkType explanation (message = task key in message, branch = commit on task branch) and the match=path behavior where line ranges are ignored for that commit. The startLine-endLine relationship to range filtering is also clarified beyond the raw schema strings.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb+resource (show who changed lines startLine-endLine of a file and in which task) and pins down the data source and ordering (commits synced to Kosmodrom, all branches, newest first). This distinguishes it from the nearest sibling, repo_blame, by tying results to task-linked commits rather than raw attribution.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied through hard scoping rules — 'only task-linked commits are included' and 'results are candidates' — so an agent can infer this is for task-linked annotation, not exhaustive blame. However, no sibling alternative is named (e.g. repo_blame for non-task attribution), leaving selection partly to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_treeA
Read-only
Inspect

Lists a repository directory at a ref: up to 500 entries, 1-4 levels below path (truncated when cut off). The response includes the resolved sha.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoBranch, tag or sha; defaults to the default branch.
pathNoDirectory from the repository root, without a leading /; defaults to the root.
depthNoNesting depth, 1-4; default 1.
repositoryIDYesRepository ID.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, so the description carries the rest and does it well: it discloses the 500-entry cap, the 1-4 level depth window, that results are truncated when cut off, and that a resolved sha is returned. These are genuine behavioral traits beyond the annotation, though nothing is said about pagination or error/auth behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, front-loaded with the operation and followed by the operational limits. Every clause carries information (entry cap, depth, truncation, returned sha) and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description responsibly mentions the returned resolved sha, plus the truncation and depth constraints an agent needs to interpret results. It is close to complete for a read-only listing tool, missing only pagination and any hint about error conditions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents ref, path, depth, and repositoryID. The description restates the ref and depth semantics ('at a ref', '1-4 levels below path') without adding syntax or format details the schema lacks, which is the baseline-3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Lists) and resource (a repository directory at a ref), which is enough to tell it apart from repo_file, repo_blame, or repo_history by resource type. It does not explicitly name any sibling, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative routing guidance. With siblings like repo_search, repo_file, and repo_context available, an agent gets no help deciding when a directory listing is the right call versus those tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_contextB
Read-only
Inspect

Returns the token owner (signed-in user), the company and the language of its data, token scopes, the owner's privileges (allows) and token expiry (expiresAt; null - no expiry). Also includes the company's and user's conventions and notes for AI agents (memoryInstructions) with an index of further memory records (memory), as reference material.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already tells the agent this is a safe read. The description adds useful semantics beyond the annotation, such as expiresAt being null when there is no expiry and the memory index being reference material. It omits any note on auth requirements or caching, but for a read-only tool the burden is low.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence front-loads the core identity/auth fields before the memory-related clauses. Every clause maps to a returned field, though the trailing 'as reference material' phrasing is slightly redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining return values, and it enumerates the key fields (owner, company, language, scopes, allows, expiresAt) plus the memory index. It is reasonably complete, though it could clarify the relationship between memoryInstructions and the memory records.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The schema requires nothing and the description correctly implies there is no input to negotiate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb+resource: it returns the session/token context including owner, company, language, scopes, privileges, and expiry. This is a distinct resource from siblings like user_context and product_context. However, it never explicitly distinguishes itself from user_context, which likely overlaps on the owner information.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given. It does not say to call it at session start, to establish identity/permissions before other tools, or how it relates to the memory_list/memory_read siblings it indexes. The agent must infer the trigger condition entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_actionAInspect

Moves a task or bug through its lifecycle (shared team data); only the task's allowedTransitions are accepted: work (start), pause, wait (back to waiting), publish (publish a draft), accept / decline (approve or send back at approval or review), reject (send back from testing), done (submit the work), reopen (reopen a closed task for the token owner in the current iteration). done leads to review if reviewAssignee is set, else for dev and bug tasks to deployment when the product has it enabled (isDeployEnabled) or to testing when testAssignee is set; otherwise the task closes and only reopen returns it. result is not written here; the done response includes feature.plan (plan vs. actual).

ParametersJSON Schema
NameRequiredDescriptionDefault
IDNoTask ID; preferred over key.
keyNoTask number, if the ID is unknown.
commentNoComment posted to the task chat, in the language of the workspace data; required for decline and reject (what needs fixing).
transitionYesOne of the task's allowedTransitions.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false; the description goes far beyond by spelling out the state machine (done routes to review, deployment, testing, or close), the reopen restriction to the token owner in the current iteration, that result is not written here, and that done returns feature.plan. This is exactly the extra behavioral context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and dense with useful detail, and no sentence is pure filler. However it is delivered as one long run-on chain, which makes the transition list harder to scan than a shortened enumerated form would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter mutation tool with no output schema, it covers the key preconditions (allowedTransitions), the required comment, and the post-transition outcomes, and even hints at the done response. It stops short of describing the general response shape, which is the remaining gap given no output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains what each transition value does and states that comment is required for decline and reject, plus ID being preferred over key. This goes beyond the bare enum and field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (moves) and resource (a task or bug) scoped to lifecycle transitions in shared team data, which clearly separates it from siblings like task_update, task_create, and task_delete. An agent can tell what it does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Enumerates each accepted transition (work, pause, wait, publish, accept/decline, reject, done, reopen) with its meaning, giving strong per-case guidance. It also states the precondition that only the task's allowedTransitions are accepted, but never names alternative tools (e.g., task_update) or when-not-to-use this one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_code_contextA
Read-only
Inspect

Returns a task's code: product repositories (main one first), live branches with the task key, commits with stats and paths, the task's MRs (link; review threads stay at the git provider) and reading hints. Useful when the user asks how a task is implemented or wants it reviewed.

ParametersJSON Schema
NameRequiredDescriptionDefault
IDNoTask ID.
keyNoTask number without #, if the ID is unknown.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already tells the agent this is a safe read, so the bar is lower; the description still adds real scope boundaries beyond the annotation — MRs are returned as links with review threads deliberately left at the git provider, repositories are ordered main-first, and "reading hints" are included. It does not mention pagination, result limits, or behavior when the task has no code, which keeps it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the returned artifacts and closed with the use case; no filler or repetition. The first sentence is dense with parenthetical asides, which slightly taxes readability but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates well by enumerating the return contents (repos, branches, commits, MRs, hints) so the agent knows what to expect. Combined with the readOnly annotation and fully documented parameters, the definition is nearly self-sufficient; only pagination/empty-state behavior is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (ID, key) are documented in the schema, including the "key if the ID is unknown" fallback. The description says only "a task's code" and adds no syntax, format, or ID-vs-key guidance, so the baseline of 3 for schema-carried parameters is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Returns a task's code") and then enumerates the payload precisely: product repositories in main-first order, live branches carrying the task key, commits with stats and paths, the task's MRs, and reading hints. This scope enumeration implicitly separates it from siblings like task_commits or repo_context. It stops short of naming any alternative tool, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Useful when the user asks how a task is implemented or wants it reviewed" gives a clear positive trigger for invocation. There are no exclusions and no routing against the many adjacent tools (task_commits, task_context, feature_code_context), so it is clear context without the when-not guidance a 5 would require.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_commitsA
Read-only
Inspect

Returns a task's commits from the product's synced repositories: sha, message, link, author, branches, stats and file paths. linkType: message - "#" in the commit message; branch - a commit on the task branch without the key.

ParametersJSON Schema
NameRequiredDescriptionDefault
IDNoTask ID.
keyNoTask number without #, if the ID is unknown.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=true already covering the safety profile, the description adds genuinely non-structured context: commits come specifically from synced repositories, and it defines the linkType classification ('message' = '#<key>' in the commit message; 'branch' = commit on the task branch without the key). That linkType semantics is real behavioral/domain information absent from annotations and schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose and return shape before the linkType detail. The linkType notation is terse to the point of being slightly cryptic, but there is no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, output-schema-less tool, the description supplies the returned fields and the linkType semantics an agent needs to interpret results. The remaining gap is that it never states both parameters are optional or what happens if neither ID nor key is supplied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (both ID and key are documented, including the 'without #' nuance), so the schema carries the parameter burden and a 3 baseline applies. The description adds no parameter meaning of its own beyond the mention of '<key>' inside the linkType explanation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns a task's commits') and scopes the source ('from the product's synced repositories'), enumerating the returned fields. This clearly distinguishes it from generic repo tools, though it never names the closest siblings (repo_history, task_code_context) that an agent might confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance. Nothing tells the agent how this differs from repo_history, repo_compare, or task_code_context, and the schema (not the description) is the only place that notes key is a fallback when ID is unknown, so the selection logic is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_contextA
Read-only
Inspect

Returns a task or bug by ID or number (unique within the company): spec, gist (gistInfo.stale - the spec changed after the gist), category, type (severity for bugs; bugType - defect nature), feature (feature.spec with includeFeatureSpec; feature.plan - a window of the plan), relations (dependencies - what it waits for; blockedTasks - what waits for it), allowedTransitions, allowedActions, specRevision, productID and a page of the task chat. Imported designs are in designs (design: markers; snapshotFID, contextFID with contextSize, assetsFID) and attachments in attaches by fid; raw design links without a designs entry are not imported. Useful when the user mentions a task number or asks about a task.

ParametersJSON Schema
NameRequiredDescriptionDefault
IDNoTask ID; preferred over key.
keyNoTask number without a prefix, e.g. 25865, if the ID is unknown.
chatAfterNoRead messages newer than chat.after.
chatLimitNoChat messages to return (default 20, max 100).
chatBeforeNoRead messages older than chat.before.
includeFeatureSpecNotrue includes the full feature spec (feature.spec); omitted by default.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true, so safety is covered. The description goes further by disclosing that the task chat is returned as a single page, that designs are only surfaced via imported markers, and that raw design links without a designs entry are not imported. These are meaningful behavioral facts beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first clause, which is good. However, the body is one long run-on enumeration of returned fields and the usage hint is tacked on at the end, making it dense and hard to scan rather than concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of describing the return shape, and it does so thoroughly (spec, gist, relations, transitions, designs, attaches, chat). Combined with full schema coverage and a readOnly annotation, an agent has what it needs to call and interpret the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters are already documented in the schema. The description references includeFeatureSpec and the chat page only in passing, adding little semantics beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening states a specific verb and resource: 'Returns a task or bug by ID or number (unique within the company),' which is unambiguous. It then enumerates the returned context fields, reinforcing what the tool is for. It does not, however, explicitly contrast itself with siblings like task_search or task_code_context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It offers a trigger condition ('Useful when the user mentions a task number or asks about a task'), which is implied usage guidance. But 'asks about a task' is broad enough to overlap with task_search, and no alternative tool or when-not condition is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_createAInspect

Creates a task (not a bug) as a draft; the token owner reviews and publishes it. rulesRevision is the current revision of the workspace's task creation template for the type, and spec is checked against that template; placement is one of featureID / featureKey (product and version from the feature), versionID (no suitable feature) or productID (product backlog). Returns link, status and available transitions; without ord the task goes to the end of the feature plan.

ParametersJSON Schema
NameRequiredDescriptionDefault
ordNoPosition of the new bug; tasks are ordered by `ord` descending. Omitted - end of the feature's plan, so bugs created one by one keep their order.
gistYesGist: one or two sentences, "problem → solution", up to 300 characters of plain text (no Markdown, links or task keys), in the language of the workspace data.
specNoSpec (requirements) in the language of the workspace data, per the task creation template. Markdown (HTML is rejected); a write replaces the whole field, including `::: doc` sections and `attach:`, `design:`, `color:` tokens
typeYesType by the nature of the work: planning, analysis, design, research, education, documentation, other - general tasks; frontend, backend, fullstack, devops, tests - development (client and server together - one fullstack task); freetest, smoke, regression, testcase - testing
titleYesTask title in the language of the workspace data, up to 200 characters, conveying the substance of the work rather than an area: "Authorization", "List", "Settings" are poor titles that don't distinguish the task
featureIDNoFeature ID - the preferred placement
moduleIDsNoProduct modules the task affects
productIDNoProduct ID - the product backlog, last resort
versionIDNoVersion ID - when no feature fits
featureKeyNoFeature key #F12 (or F12, 12); not with featureID
componentIDsNoProduct components the work actually touches; a module does not imply them. May stay empty if unsure
testAssigneeNoTester: employee uuid, "any" (any employee takes the stage) or null (no stage); default - no testing
workAssigneeNoAssignee: employee UUID, "any" (any employee) or null (no work stage); default "any".
workTimePlanNoPlanned work time in minutes, per the assignee's role level
rulesRevisionYesRevision of the task creation template for this type; a stale one is rejected
reviewAssigneeNoReviewer: employee UUID, "any" (any employee) or null (no review stage); default: no review.
approvalAssigneeNoApprover: employee UUID, "any" (any employee) or null (no approval stage); default: no approval.
workIterationKeyNoIteration key for work outside the current iteration; default - current

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare a non-destructive write; the description adds real behavioral context beyond that: the task lands as a draft that the token owner must review and publish, a stale rulesRevision is rejected, and it names the returned link/status/transitions. Auth/permission requirements are not covered, keeping this below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose and packs placement, validation, and return info into one dense paragraph. At 18 parameters some density is justified, but the run-on clause chain costs readability versus short scannable sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 18-parameter tool with no output schema, the description usefully covers the return surface, the template-revision requirement, ord semantics, and the draft/publish workflow. It stops short of failure modes and permission requirements, but nothing essential for a correct call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description goes further by explaining cross-parameter placement logic (which key to choose and why) and the ord default behavior ('without ord the task goes to the end of the feature plan'), which the schema states only tersely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Creates a task') and immediately scopes it against the sibling it is not ('not a bug'), which is exactly the distinction an agent needs versus bug_create. The draft/review lifecycle is also stated up front.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing among placement options: featureID/featureKey preferred, versionID when no feature fits, productID as last resort. It does not explicitly name bug_create as the alternative for bugs (only the parenthetical 'not a bug'), so exclusions are implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_deleteA
Destructive
Inspect

Deletes a task or bug (shared team data). Irreversible: deletion is soft, but there is no way to restore the task; only at the user's direct request. Submitting work is the done transition, not deletion.

ParametersJSON Schema
NameRequiredDescriptionDefault
IDNoTask ID; preferred over key.
keyNoTask number, if the ID is unknown.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true, but the description adds the crucial nuance that the delete is soft yet effectively irreversible with no restore path, and imposes a consent precondition (user's direct request only). This is real behavioral context beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, action first, then the irreversibility warning, then the semantic disambiguation. Every sentence carries distinct, non-redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no output schema, the description covers the safety profile, reversibility, consent requirement, and the most likely semantic confusion. Nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already explains ID-vs-key preference and the 'if the ID is unknown' fallback. The description adds nothing about parameter handling, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (deletes a task or bug) and scopes it as shared team data. It further disambiguates against the adjacent 'done transition' semantic that task_action would cover, so an agent can separate it from siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance ('only at the user's direct request') and a when-not clarification (submitting work is the done transition, not deletion). It does not name the alternative tool (e.g. task_action/task_update) explicitly, so routing still requires light inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_evaluateAInspect

Saves a task stage evaluation when the user asks to evaluate work: a recommended bonus factor and a comment, visible to all who see the task. One evaluation per stage (a new one replaces the old); the server marks it final or interim by whether the task is closed, and the factor is a recommendation that does not change tracked time.

ParametersJSON Schema
NameRequiredDescriptionDefault
IDNoTask ID; preferred over key.
keyNoTask number without #, if the ID is unknown.
stageNoStage: work - execution (default), review, or testing.
factorYesRecommended bonus factor, 0 to 2 in steps of 0.1.
commentYesRationale in the language of the workspace data: up to 4 paragraphs, no rates or money amounts. Markdown (paragraphs, headings, lists, bold, italic, tables, links, emoji); attach:, design: and nested ::: doc are not supported.
rulesRevisionYesRevision of the evaluation rules for this task type and stage.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare only readOnlyHint=false and destructiveHint=false, and the description adds substantial behavior beyond that: one evaluation per stage with replacement semantics, server-side final/interim determination based on task closure, and the fact that the factor is a recommendation that does not alter tracked time. Visibility of the comment to all task viewers is also disclosed. This is exactly the kind of context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that leads with purpose and then layers the write/visibility semantics. It is dense but every clause is informative; slight overpacking into one sentence keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description adequately covers the mutation's outcome (final/interim marking, replacement of prior evaluation) and the non-effect on tracked time. Required rulesRevision is left to the schema, which is acceptable given full coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: the factor is a recommendation that does not change tracked time, and the comment is visible to all who see the task. It does not explain rulesRevision, but that is documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Saves a task stage evaluation') plus the data it carries (bonus factor, comment). It is clearly distinguishable from siblings like task_evaluation_context or evaluation_rules, which are read/config tools rather than the act of recording an evaluation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit triggering condition: 'when the user asks to evaluate work.' That is clear context for invocation. It does not, however, name an alternative tool or state when-not-to-use (e.g., versus reading task_evaluation_context), so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_evaluation_contextA
Read-only
Inspect

Returns the data for evaluating a task stage when the user asks for a work evaluation: spec and result, the stage assignee with level, rates, time and cost, plan and time by status, returns, bugs and regressions with fix cost, commits and the previous evaluation. Money requires task:costs for the task, otherwise null; null salary and bonus are hidden. Rates and amounts are for the assessment, not for the evaluation comment.

ParametersJSON Schema
NameRequiredDescriptionDefault
IDNoTask ID; preferred over key.
keyNoTask number without #, if the ID is unknown.
stageNoStage: work - execution (default), review, or testing.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already establishes this as a safe read, so the bar is lower, and the description still adds real disclosure: money fields require the task:costs permission and are otherwise null, null salary/bonus values are hidden from the payload, and rates/amounts are assessment-scoped rather than comment-scoped. These are permission and null-handling traits an agent could not derive from the annotations or schema. It stops short of describing payload shape or size limits, so it isn't a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single dense run-on with an embedded list, but the purpose is front-loaded in the first clause and every subsequent clause carries distinct information (payload contents, permission caveat, null handling, scope caveat). No sentence is filler, though the list-heavy packing makes it slightly harder to scan than an ideally structured definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly carries the burden of describing return contents and does so at field-group granularity, plus the permission precondition and null semantics. For a zero-required-parameter read tool this is close to sufficient; the remaining gaps (payload size, whether cost fields appear when permission is absent, meaning of "returns") are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the ID/key/stage parameters and the work-default are already fully documented in the schema, establishing a baseline of 3. The description confirms the work-stage default and ties cost fields indirectly to caller permissions, which is useful but does not add syntax or selection meaning beyond the schema's own enum and default notes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ("Returns the data for evaluating a task stage") and enumerates the payload precisely: spec and result, assignee with level, rates, time/cost, plan by status, returns, bugs and regressions with fix cost, commits, previous evaluation. That is far more informative than a tautology like "task_evaluate data". It does not, however, name or contrast against overlapping siblings such as task_context, task_commits, or task_evaluate, so an agent cannot fully discriminate without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"when the user asks for a work evaluation" gives a usable trigger condition, which is more than implied usage, but it does not exclude or route against clearly adjacent siblings (task_evaluate for actually filing an evaluation, task_context, task_commits, evaluation_rules). The closing note "for the assessment, not for the evaluation comment" hints at a distinction but never names the alternative tool, leaving the boundary to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_relation_createAInspect

Links two tasks or bugs with a role relative to the current task key, named as in the task's relations: dependencies - the current task waits for the other; blockedTasks - the other waits for it; duplicateOf / duplicates; regressionOf / regressions; foundDuring / findings; related - neutral. For a bug, foundDuring is the task it was found during, regressionOf the task that introduced the defect (from git history only); desc explains the relation.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesCurrent task number — the one the role refers to
descNoExplanation of the relation
roleYesRole relative to the current task
otherKeyYesNumber of the task or bug to link

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the safety profile is covered; the description adds directional semantics (which side waits for which) and the provenance constraint that regressionOf is derived from git history only. It says nothing about permissions, idempotency, or what happens when the same relation already exists — meaningful gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense paragraph with zero filler, and the linking action plus role list are front-loaded. The bug-specific clauses at the end are justified but make the block long; a bulleted role mapping would scan faster, though nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-required-param mutation with full schema coverage, no output schema, and annotations covering safety, the description supplies what structured data omits: role directionality and bug-specific semantics. Missing only edge behavior (duplicate relation, permission requirements, effect on existing links), which keeps it short of a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and key/otherKey are self-documenting, but the schema's role description is only a bare enum label ('Role relative to the current task'). The description defines the direction of every role value (dependencies = current waits for other; blockedTasks = other waits for it; etc.) and explains desc as the relation explanation — real meaning beyond the structured data.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Links two tasks or bugs') and immediately scopes it: the role is interpreted relative to the current task key, with the counterpart given as otherKey. This is enough for an agent to distinguish it from task_relation_update/delete without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description effectively tells the agent which role to pick for which real-world situation (waits-for, duplicate, regression, found-during, neutral), which is the main 'when to use' decision here, and clarifies foundDuring/regressionOf are bug-specific with regressionOf sourced from git history. It does not, however, say when to use this tool versus task_relation_update or task_relation_delete for an existing relation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_relation_deleteA
Destructive
Inspect

Deletes a task relation addressed by key, otherKey and role, as in the task's relations; removing dependencies / blockedTasks unblocks the waiting task. Irreversible; only at the user's direct request, after the user confirms what will be affected.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesCurrent task number — the one the role refers to
roleYesCurrent role of the relation
otherKeyYesRelated task or bug number

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, but the description adds meaningful context beyond them: it is 'Irreversible', requires explicit user confirmation, and explains the side effect that removing dependencies/blockedTasks unblocks the waiting task. That unblocking consequence is genuinely non-obvious and high value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, front-loaded with the action, followed by the irreversibility/confirmation caveat. No filler, though the clause about roles is slightly compressed and could be clearer.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no output schema, the description covers what is deleted, the side effect on blocked tasks, irreversibility, and the confirmation requirement. An agent has enough to call it correctly; only error/response behavior is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with each parameter described and role enumerated, so the schema already does the heavy lifting. The description only restates the three identifiers without adding syntax or semantics beyond the schema (e.g., it does not clarify that removal is directional). Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Deletes') and resource ('a task relation'), and identifies it precisely via key/otherKey/role. Clearly distinguishable from siblings task_relation_create and task_relation_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use constraint: 'only at the user's direct request, after the user confirms what will be affected.' This is strong gating guidance. It does not, however, name the alternative (e.g., task_relation_update) for changing rather than removing a relation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_relation_updateAInspect

Changes the role or explanation of a task relation instead of deleting and recreating it; the relation is addressed by key, otherKey and the current role, as in the task's relations. Moving it away from dependencies / blockedTasks unblocks the waiting task with notifications, as in the web app.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesCurrent task number — the one the role refers to
descNoNew explanation; null or empty string — remove; omitted — unchanged
roleYesCurrent role of the relation
newRoleNoNew role relative to key; omitted — unchanged
otherKeyYesRelated task or bug number

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false; the description goes further by disclosing the side effect that moving a relation away from dependencies/blockedTasks unblocks the waiting task and triggers notifications, matching web-app behavior. Auth requirements and rate limits remain unstated, so it is not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with a semicolon, front-loaded with the core purpose before the addressing detail and the unblock side effect. Efficient, though the middle clause about addressing is slightly compressed for a first-time reader.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema and no destructive flag, the description covers the mutation's purpose, how the target relation is addressed, and the notable unblocking/notification side effect. Minor gaps (permission needs, what the response returns, behavior when neither role nor desc actually changes) remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds genuine meaning beyond the schema by explaining that the relation is identified by the composite of key, otherKey and the *current* role (as seen in the task's relations), which clarifies why role is required even when newRole is supplied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (changes) and resource (role or explanation of a task relation), and explicitly contrasts with the delete-and-recreate path used by sibling tools task_relation_delete/task_relation_create. An agent can distinguish it from those siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the intended usage pattern ('instead of deleting and recreating it') and the identifying triple (key, otherKey, current role), which is exactly what an agent needs to invoke it rather than delete+create. It stops short of naming alternative sibling tools or stating when the tool should NOT be used.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_spec_syncA
Read-only
Inspect

Returns the context for updating a task's or bug's spec: spec, feature, recent chat, a code summary (basis=actual, with git read access) and the workspace's spec-update guidelines (instructions); writes nothing. mode=analyze - for "check, compare, propose" requests (differences and a full new version); mode=apply - when the user asks to update or sync. basis=requirements - "from the chat"; basis=actual - "from the code, as built".

ParametersJSON Schema
NameRequiredDescriptionDefault
IDNoTask or bug ID; not with key
keyNoTask or bug number; not with ID
modeYesanalyze - propose without saving; apply - the user asks to apply
basisNorequirements (default) - confirmed in the chat; actual - actual state: code, designs, result
chatLimitNoLatest chat messages to return, 1-100 (default 20); older via chat.before

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint is already declared, and the description corroborates it with 'writes nothing', which resolves the potential ambiguity that mode=apply might imply a write. It also discloses the dependencies of the returned content (code summary requires basis=actual and git read access). It stops short of describing pagination or response shape, but adds meaningful context beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the 'writes nothing' constraint are front-loaded, followed by compact mode/basis breakdowns. It is dense but every clause carries information; the run-on use of semicolons is slightly harder to parse than ideal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description enumerates the returned context (spec, feature, chat, code summary, guidelines), covers both enum parameters and their interactions, and confirms the read-only behavior. Nothing needed to invoke this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five params, setting the baseline at 3. The description adds user-intent phrasing for the enum values ('check, compare, propose', 'the user asks to update or sync', 'from the chat', 'as built') that maps natural-language requests onto parameter values, giving it an edge over the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns the context for updating a task's or bug's spec') and then enumerates exactly what is returned: spec, feature, recent chat, code summary, and spec-update guidelines. The task/bug scoping implicitly distinguishes it from the sibling feature_spec_sync, and 'writes nothing' pins down the nature of the operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent by mode: 'mode=analyze - for "check, compare, propose" requests' versus 'mode=apply - when the user asks to update or sync'. The basis param is similarly disambiguated ('from the chat' vs 'from the code, as built'), which is exactly the when-to-use guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_updateAInspect

Edits a task or bug (shared team data); only the passed fields change: title, spec, gist, result, priority, labels, placement (featureID or featureKey, featureID null removes it from the feature; versionID), bug fields (build, bugType, stage, browser), non-bug type, modules and components (full new lists), stage assignees and ord. Status is not changed here; returns link, status, transitions, specRevision and changed.

ParametersJSON Schema
NameRequiredDescriptionDefault
IDNoTask ID; preferred over key.
keyNoTask number, if the ID is unknown.
ordNoPosition in lists, as drag-and-drop in the web UI. Tasks are sorted by ord DESCENDING (a feature's tasks follow its plan top down): a value between two neighbors' ord goes between them, above the largest - on top, 0 - last (the server assigns it). When neighbors are under 100 apart or the user asks for a re-sort, all of the feature's tasks including closed ones are renumbered: Date.now() for the top one, then steps of -10000.
gistNoGist: one or two plain-text sentences "problem → solution", up to 300 characters, without Markdown, links or task keys, in the language of the workspace data; an empty string clears it.
specNoNew spec (requirements): Markdown in the language of the workspace data, HTML is rejected. Replaces the whole field - text, `::: doc` sections and `attach:`, `design:`, `color:` tokens left out are removed.
typeNoNew work type (non-bugs only).
buildNoBuild where the defect was reproduced (bugs only).
stageNoStand where the bug was reproduced: dev - development and autotests, alpha - manual QA, prod - production data.
titleNoNew name
labelsNoLabels; an empty array removes all.
resultNoReport on the work done: Markdown in the language of the workspace data, HTML is rejected; replaces the whole field, like spec.
browserNoBrowser where the bug was reproduced (bugs only).
bugTypeNoDefect nature (bugs only): ui - interface, fn - functionality, ux - usability, spec - documentation; not the severity, which is type.
priorityNoPriority; a human decision, set only at the user's explicit request.
featureIDNoFeature to move the task to; null removes it from its feature.
moduleIDsNoFull new list of module IDs, replacing the current one; [] removes all.
versionIDNoVersion to move the task to when no feature fits.
featureKeyNoKey of the feature to move the task to (#F12), instead of featureID.
componentIDsNoFull new list of the actually affected component IDs, replacing the current one; [] removes all, adding means the current ones plus new.
designUnlinkNotrue lets a spec / result write drop design markers (designs left without links are deleted); otherwise losing a marker is rejected.
testAssigneeNoTester: employee uuid, "any" (any employee) or null (no testing stage).
workAssigneeNoAssignee: employee uuid or "any" (any employee); null is not allowed - the work stage always exists.
workTimePlanNoPlanned work time in minutes, per the assignee's role level
reviewAssigneeNoReviewer: employee uuid, "any" (any employee) or null (no review stage).
approvalAssigneeNoApprover: employee uuid, "any" (any employee) or null (no approval stage).
expectedSpecRevisionNoThe task's current specRevision; required when spec changes.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only carry readOnlyHint=false and destructiveHint=false, so the description must add the rest — and it does: it discloses merge-vs-replace behavior, that spec/result writes remove omitted doc sections and tokens, that featureID null detaches the task, and it even names the returned payload (link, status, transitions, specRevision, changed) despite no output schema. It stops short of noting permission/prerequisite requirements, keeping it from a full 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first clause and the sentences are dense but information-bearing, with no filler. It is one long semicolon-heavy sentence plus a compact second sentence, which is efficient but slightly difficult to parse for a 26-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutation with no output schema, the description covers the core edit semantics, the return payload, and the status exception. It leaves some parameters (designUnlink, expectedSpecRevision requirement, work/review/approval assignees) to the schema, but the schema's 100% coverage makes that acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all 26 parameters in detail; the description largely restates and groups them (placement, bug fields, full new lists) rather than adding new syntax or constraints. Baseline 3 is appropriate, with minor credit for the grouping that aids navigation of the large param set.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Edits a task or bug') and immediately scopes it ('shared team data'), then enumerates the editable field families. The clause 'Status is not changed here' implicitly distinguishes it from a status/action tool, so an agent can separate it from task_create, task_delete and task_action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description signals partial-update semantics ('only the passed fields change') and that status changes belong elsewhere, which is useful routing context. However, no alternative tool is named explicitly and there is no statement of when to prefer this over task_create or task_action, so guidance remains implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

user_contextB
Read-only
Inspect

Returns an employee's profile, access and roles by userID. Useful when the user asks about a specific colleague.

ParametersJSON Schema
NameRequiredDescriptionDefault
userIDYesEmployee ID.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already establishes this as a safe read, so the description doesn't need to restate safety. It does add content-scope information (profile, access, roles), which tells the agent this returns permission data rather than just a name. It discloses nothing about sensitivity, authorization requirements, or what happens for an unknown ID.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler, and the capability is front-loaded ahead of the usage hint. Nothing is padded, though the second sentence is thin enough that it could have carried the missing alternative-tool guidance instead.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with no output schema, the agent knows what it gets (profile, access, roles) and the entry condition (a userID). The unresolved gap is how to obtain that userID and why this beats user_search, which the description leaves entirely to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single userID parameter (uuid format, 'Employee ID'), so the schema carries the load. The description's 'by userID' merely confirms lookup-by-identifier behaviour and adds no format or validation detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: returns an employee's profile, access and roles, keyed by userID, which is materially more informative than the tautological title 'User context'. It does not, however, distinguish itself from the sibling user_search, which an agent could easily confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Useful when the user asks about a specific colleague' gives implied usage context, but the crucial precondition — that you must already hold a userID and should use user_search to obtain one — is never stated or excluded. There is no explicit when-not or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

version_createAInspect

Creates a product version numbered major.minor with an optional build; the server composes its name. Changes shared product structure; only at the user's request or with their consent. Returns the version.

ParametersJSON Schema
NameRequiredDescriptionDefault
buildNoBuild number, from 1; omitted - a "major.minor" version (there is no build 0).
majorYesMajor part of the number.
minorYesMinor part of the number.
notesNoVersion notes in Markdown (no HTML), in the language of the workspace data.
productIDYesProduct ID
designFileUrlNoLink to the version's Figma file; null or empty - the product's main design file.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false; the description adds real context beyond them: that the change is to shared product structure, that user consent/request is required, that the server composes the version name, and that it returns the version. It does not disclose side effects on existing versions or permissions beyond consent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short clauses, zero waste, and the core action plus numbering format are front-loaded before the consent constraint and return statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by stating it returns the version, and it covers the mutation's governing constraint (user consent). The schema fully documents the six inputs, so little is missing, though side effects on existing version lineage go unmentioned.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, and the description adds meaning the schema cannot: the component parts (major.minor, optional build) are combined and the name is server-composed, so the agent knows it need not supply a name. It does not add anything about notes or designFileUrl beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Creates a product version') and adds the numbering scheme (major.minor with optional build) and that the server composes the name. It does not explicitly differentiate the tool from the sibling version_update or version_release, so an agent must infer the distinction from names alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a genuine usage precondition: 'only at the user's request or with their consent', plus the warning that this changes shared product structure. However, it never names alternatives (version_update, version_release) or states when to prefer those over creating a new version.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

version_releaseA
Destructive
Inspect

Releases a version: marks it completed with a release date, after which it is not offered for new work; isPublished also publishes it. Without releaseDT the version keeps its existing (possibly planned) date or gets the current date if it has none; the response has the resulting releaseDT. Irreversible; only at the user's explicit request.

ParametersJSON Schema
NameRequiredDescriptionDefault
releaseDTNoRelease date, ISO.
versionIDYesVersion ID.
isPublishedNotrue - also publish the version.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as destructive and non-read-only, and the description adds real context on top: irreversibility, the completed/no-new-work consequence, the publishing side effect of isPublished, and the date-defaulting rule (keeps existing planned date or gets current date). That is well beyond what the hints convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense multi-clause sentence with the primary effect front-loaded, followed by the date default and the irreversibility warning. Every clause carries information, though the semicolon-chained structure makes it slightly run-on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by noting the response returns the resulting releaseDT, plus it covers side effects, defaulting, and irreversibility for a destructive 3-param mutation. Remaining gaps are minor (permissions, downstream impact on tasks) but not agent-blocking.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds behavior the schema descriptions don't: what happens when releaseDT is omitted (existing date retained, or current date applied if none) and that isPublished triggers publication. versionID semantics remain unelaborated, keeping it from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Releases a version') and explains the resulting state change: marked completed, no longer offered for new work, and the isPublished side effect. It semantically separates release from sibling mutators like version_update, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear usage gate: 'Irreversible; only at the user's explicit request,' which tells the agent when not to call it unprompted. It does not point to version_update or any alternative for the reversible cases, so it stops short of full when/when-not/alternatives coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

version_updateAInspect

Changes a version's number, notes or design file (only the passed fields); does not release it. Changes shared product structure; only at the user's request or with their consent. Returns the version.

ParametersJSON Schema
NameRequiredDescriptionDefault
buildNoNew build number, from 1 (there is no build 0).
majorNoNew major part of the number.
minorNoNew minor part of the number.
notesNoNew version notes in Markdown (no HTML), in the language of the workspace data. Replaces the whole field: existing text, `::: doc` sections and `attach:`, `design:`, `color:` tokens stay only if included.
versionIDYesVersion ID.
designFileUrlNoLink to the version's Figma file; null or empty - the product's main design file.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false; the description adds genuinely useful behavior beyond that: it is a partial update (only passed fields change), it does NOT release the version, it mutates shared product structure (a consent/caution signal), and it returns the version. It does not explain permissions or error behavior, but the added context is substantial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with what changes and what does not, then the caution and return value. Every sentence carries information; the middle caution clause is slightly dense but earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-param mutation with one required field, full schema coverage, and annotations covering safety, the description supplies the remaining essentials: partial-update scope, non-release behavior, shared-structure side effect, and the return value (compensating for the absent output schema). Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so build/major/minor/notes/designFileUrl semantics (including the full-replacement behavior of notes and the null meaning of designFileUrl) are already documented in the schema. The description only restates the affected field set, adding no syntax or constraint beyond what the schema supplies — the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (changes) and resource (a version's number, notes, design file) and immediately scopes it with 'only the passed fields' and 'does not release it', which cleanly separates it from version_release and version_create. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a real usage condition — only perform this at the user's request or with their consent because it changes shared product structure — and implicitly routes release operations elsewhere via 'does not release it'. It stops short of explicitly naming alternative tools for adjacent cases (e.g. version_release), so it is clear context rather than exhaustive when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

worklog_listA
Read-only
Inspect

Lists worklog (time-tracking) records of tasks and bugs for time and cost reports; filters by task, feature, employee, period and task status combine (none - all accessible records), meetings excluded. Amounts the user may not see come as 0, meaning hidden, not unpaid; isActive marks ongoing work whose time and cost still grow. totals covers the whole selection, employees and tasks are in the users and tasks dictionaries, and the next page comes with nextCursor and the same filters.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNoPeriod end, YYYY-MM-DD inclusive, in the company's time zone.
fromNoPeriod start, YYYY-MM-DD inclusive, in the company's time zone.
limitNoRecords per page, up to 500 (default 100).
cursorNonextCursor of the previous page of the same request.
statusNoOnly work tracked while the task had this status, e.g. review - review time, testing - testing time.
taskIDNoTask or bug ID (preferred).
userIDNoEmployee ID: that employee's records.
taskKeyNoTask or bug number, if the ID is unknown.
featureIDNoFeature ID: all its tasks and bugs.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well past the readOnlyHint annotation by disclosing non-obvious semantics: hidden amounts are returned as 0 and mean hidden rather than unpaid, isActive marks work whose time/cost still grow, meetings are excluded, and totals span the whole selection. These are exactly the behavioral facts an agent cannot infer from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, front-loaded with what the tool returns before moving to filter and response semantics. Every clause carries information, though the second sentence strings together several distinct return-value facts and could be split for easier scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must carry the return contract, and it does: totals scope, where employees/tasks are resolved, pagination via nextCursor with the same filters, hidden-amount encoding, and isActive meaning. An agent has everything needed to call and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter in detail, establishing a baseline of 3. The description adds only marginal value, restating the filter dimensions and noting that filters combine (none = all accessible records), without new format or interaction rules beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Lists worklog (time-tracking) records of tasks and bugs for time and cost reports', with concrete scope notes ('meetings excluded'). It is clearly a distinct capability from any sibling, but it never explicitly names or contrasts a sibling tool, so it stops short of the top tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the stated purpose ('for time and cost reports') and the enumerable filters, but there is no explicit when-to-use guidance, no when-not-to-use, and no named alternative among the many task/report siblings. Adequate but with a clear gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 73 tool updates
    • First observedbranch_list
    • First observedbranch_name
    • First observedbug_create
    • First observedchat_message_delete
    • First observedchat_message_update
    • First observedchat_question_inbox
    • First observedchat_reaction
    • First observedchat_read
    • First observedchat_send
    • First observedcomponent_create
    • First observedcomponent_update
    • First observedcreate_rules
    • First observedcurrency_convert
    • First observeddesign_comment_reply
    • First observeddesign_comments
    • First observeddesign_import
    • First observedevaluation_rules
    • First observedfeature_code_context
    • First observedfeature_context
    • First observedfeature_create
    • First observedfeature_delete
    • First observedfeature_list
    • First observedfeature_search
    • First observedfeature_spec_sync
    • First observedfeature_update
    • First observedfeed_work_queue
    • First observedfile_get
    • First observedfile_info
    • First observedfile_upload
    • First observedfile_upload_complete
    • First observedfile_url
    • First observediteration_key
    • First observedmemory_create
    • First observedmemory_delete
    • First observedmemory_list
    • First observedmemory_read
    • First observedmemory_update
    • First observedmodule_create
    • First observedmodule_update
    • First observedproduct_context
    • First observedproduct_list
    • First observedproduct_modules
    • First observedproduct_update
    • First observedrepo_blame
    • First observedrepo_branches
    • First observedrepo_compare
    • First observedrepo_context
    • First observedrepo_file
    • First observedrepo_history
    • First observedrepo_line_history
    • First observedrepo_search
    • First observedrepo_tree
    • First observedsession_context
    • First observedtask_action
    • First observedtask_code_context
    • First observedtask_commits
    • First observedtask_context
    • First observedtask_create
    • First observedtask_delete
    • First observedtask_evaluate
    • First observedtask_evaluation_context
    • First observedtask_relation_create
    • First observedtask_relation_delete
    • First observedtask_relation_update
    • First observedtask_search
    • First observedtask_spec_sync
    • First observedtask_update
    • First observeduser_context
    • First observeduser_search
    • First observedversion_create
    • First observedversion_release
    • First observedversion_update
    • First observedworklog_list

Publisher details

Operator
Soft-Artel · Publisher source
Vendor relationship
Not applicable
Documentation
Unknown
Trust center
Unknown
Restrictions
Not available

Related MCP Connectors

Related MCP Servers

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources