Cursor Cloud MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Cursor Cloud MCPlist my recent cloud agents and read the latest transcript"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Cursor Cloud MCP
Unofficial MCP server for Cursor Cloud Agents. Same sign-in and API handling as Cursor for Android. Not affiliated with Anysphere, Inc.
List agents, read transcripts, follow runs, and send follow-ups from any MCP client.
Install
Node 20+.
npm install
npm run buildRelated MCP server: scriptit
Auth
Paste an API key from cursor.com/dashboard/api:
export CURSOR_API_KEY=key_...Or sign in in the browser. That stores a key at ~/.config/cursor-cloud-mcp/credentials.json.
npm run login
npm run status
npm run logoutRun
npm start~/.cursor/mcp.json:
{
"mcpServers": {
"cursor-cloud": {
"command": "node",
"args": ["/absolute/path/to/Cursor-Cloud-MCP/build/index.js"],
"env": {
"CURSOR_API_KEY": "${env:CURSOR_API_KEY}"
}
}
}
}npm run http serves POST /mcp on port 8091.
Set CURSOR_EXTENDED_MODE=1 for the same unofficial account endpoints Cursor for Android keeps behind Extended mode. Off by default. Call cursor_status first. The server describes its own tools.
License
MIT © 2026 Bennett Buhner
Available Tools
29 toolsagent_lifecycleC
Extended: pause | resume | wake | cancel_tool_call.
| Name | Required | Description | Default |
|---|---|---|---|
| bcId | Yes | Agent bc id | |
| runId | No | Run id (pause) | |
| action | Yes | pause | resume | wake | cancel_tool_call | |
| toolCallId | No | Tool call id (cancel_tool_call) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It never states whether pause is durable or reversible, whether cancel_tool_call aborts only the current tool call or the whole run, or what permissions/state changes are involved for these mutation actions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single short fragment with no waste, but the 'Extended:' prefix is cryptic noise and the terseness here reflects under-specification rather than efficient information density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a multi-action mutation tool with no annotations and no output schema, the description omits everything an agent needs: action-specific behavior, side effects, and expected results. Only the schema's parameter names carry the semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents bcId, runId, action, and toolCallId including which action each applies to. The description merely repeats the action enum values already present in the schema and adds no format or constraint detail, matching the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description enumerates the four lifecycle actions but never names the resource or states the overall purpose in a verb+resource form — the agent must infer 'control agent execution state' from the tool name. It also does not distinguish this from siblings like cancel_run or steer_agent, which overlap with pause/cancel semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no routing to alternatives such as cancel_run for termination or steer_agent for redirection. An agent is left to guess which of these overlapping sibling tools applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
archive_agentB
Archive (soft-delete). Idempotent.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Agent id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose two genuine traits: the deletion is soft (implying recoverability) and the call is idempotent. It omits permissions required, effect on the agent's runs/data, and whether the operation can be undone, which is a meaningful gap for a destructive-leaning tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two terse fragments, front-loaded with the core action and free of filler. It is efficient, though so minimal that it may under-serve rather than being genuinely well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with full schema coverage and no output schema, the essentials are nominally present. However, the absence of any guidance on reversibility, side effects, or routing against agent_lifecycle leaves gaps an agent would need filled before invoking it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single 'id' parameter ('Agent id'), so the schema already documents the input fully. The description adds no syntax, format, or constraint information beyond that, making 3 the correct baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb 'Archive' plus the tool name 'archive_agent' makes the operation on the agent resource unambiguous, and the parenthetical '(soft-delete)' sharpens what archiving means. It does not, however, distinguish itself from the nearby 'agent_lifecycle' sibling, so an agent must infer which of the two handles archiving.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives such as agent_lifecycle or cancel_run, and no prerequisites or preconditions are given. The agent is left to infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_runB
Cancel the active run (terminal, cannot resume).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Agent id | |
| runId | Yes | Run id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the key trait that cancellation is terminal and irreversible, but says nothing about permission requirements, idempotency, or what happens when the run is already finished.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence, front-loaded with the action and immediately followed by the most important caveat. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter cancel with no output schema, the description covers the essential irreversibility note. However, with no annotations at all, it leaves gaps around authorization, behavior on non-active runs, and idempotency that an agent would benefit from knowing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both required parameters (agent id, run id), so the schema already documents them fully. The description adds no parameter-level meaning beyond the schema, which matches the baseline 3 for full coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('cancel the active run') and clarifies scope with 'active'. It does not name or contrast with any sibling, but the action is unambiguous and distinct from the read-oriented siblings like get_run or list_runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to cancel versus alternatives such as steer_agent or letting a run finish, nor any prerequisites or error conditions. The only context is the consequence, not the selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_followup_runB
Send a follow-up prompt to an existing agent (new run).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Agent id | |
| mode | No | agent | plan | |
| text | Yes | Follow-up instruction |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. The parenthetical '(new run)' usefully discloses that this spawns a distinct run rather than mutating an existing one, which is real behavioral context. But it says nothing about prerequisites (agent must exist and not be archived/terminal), permissions, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the resource and effect front-loaded and zero filler. It is efficient, though arguably too terse for a mutation-style operation, which is a completeness issue rather than a structure one.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Parameters are fully covered by the schema and no output schema exists to explain, so the remaining burden is guidance and behavior. For a tool that creates a run, the description omits prerequisites, the meaning of 'mode', and how the result relates to list_runs/get_run, leaving it only minimally viable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (id, mode, text all documented), so the baseline is 3. The description adds no format detail, no explanation of the 'agent | plan' mode values, and no indication of whether text is required to be non-empty — it merely restates the text parameter as 'follow-up prompt'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('send a follow-up prompt to an existing agent') and clarifies the effect of producing a new run, which separates it from a plain message send. However, it does not name or distinguish itself from closely related siblings such as steer_agent, start_side_chat, or list_pending_followups, so an agent still has to infer which is right.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no when-not-to-use, and no named alternative. With siblings like steer_agent (modify in-flight work) and start_side_chat (branch a conversation) present, the description leaves the routing decision entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cursor_statusA
Check auth (v1/me) and report documented vs Extended-mode availability. Start here.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, yet it only discloses that an auth endpoint is hit. It does not say whether the tool itself requires credentials, whether it is read-only, what failure modes look like, or what 'Extended-mode' refers to — notably ironic for a tool whose purpose is auth checking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two terse sentences with zero filler, the core action front-loaded and the routing hint ('Start here') placed last for emphasis. Nothing could be cut without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool this is close to sufficient, but with no output schema and no annotations the description is the only source of information, and it never clarifies the shape or meaning of the 'documented vs Extended-mode' result it promises. Minimum viable rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which sets a baseline of 4 per the rubric. The description adds no parameter detail, but none is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete action (check auth via v1/me) and a concrete output (documented vs Extended-mode availability), which is more specific than the bare name cursor_status. It does not need to distinguish from siblings since none of them are auth/status probes, but the phrase 'documented vs Extended-mode availability' remains partly opaque.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Start here' is an explicit, unambiguous directive about when to call this tool relative to the rest of the toolset. There is no exclusion or when-not guidance, so it falls short of the top mark.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_agentB
Durable metadata for one agent (repos, autoCreatePR, latestRunId).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Agent id, e.g. bc-... |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden but only names a few returned fields. It does not state that the operation is read-only, whether it requires special permissions, how errors (e.g., missing ID) are handled, or any rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the key scoping phrase ('one agent') and lists the returned metadata fields without any filler. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with no output schema, the description is minimally adequate: it names key return fields to set expectations. However, it omits usage context and behavioral details (read-only nature, auth, error cases) that would help an agent call it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single 'id' parameter is fully documented in the schema. The description adds no extra meaning beyond what the schema already provides, making the baseline score of 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the resource (durable metadata for one agent) and lists key fields, which clearly implies a retrieval operation. It distinguishes itself from the sibling list_agents by scoping to a single agent, though it omits an explicit verb like 'retrieve' or 'get'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, prerequisites, or alternatives are provided. It does not say when an agent should prefer get_agent over list_agents or other tools, leaving the caller to infer context from the name and schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_artifact_urlB
15-minute presigned URL for one artifact path.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Agent id | |
| path | Yes | Relative path, e.g. artifacts/screenshot.png |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the single most important trait: the URL expires after 15 minutes. However, it says nothing about permissions, whether the URL is safe to share, or whether the artifact must exist first.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the expiry caveat is placed prominently. It is arguably under-specified rather than verbose, but as a sentence it is tight and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with full schema coverage and no output schema, stating that the result is a presigned URL covers the essential return contract. What is missing is workflow context: when to obtain this URL rather than fetching the blob directly, and whether the artifact must be listed first.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (both id and path are documented in-schema), so the baseline is 3. The phrase 'one artifact path' loosely maps to the path parameter but adds no format or relative-path nuance beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (artifact path) and the return value (a URL), and the 'presigned' qualifier plus 15-minute window makes it specific. It does not differentiate from the sibling get_blob, which plausibly returns artifact content rather than a link, so an agent must infer the split.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no mention of prerequisites (e.g. the artifact must already be listed via list_artifacts), and no routing versus alternatives like get_blob or list_artifacts. The agent is left to guess the workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_blobC
Extended: GetBlobForAgentKV one record blob (base64 blob_data).
| Name | Required | Description | Default |
|---|---|---|---|
| bcId | Yes | Agent bc id | |
| blobId | Yes | Blob id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full behavioral burden. It does disclose that the returned data is base64-encoded, which is useful, but it says nothing about required permissions, size limits, error behavior, or whether missing blobs return null versus an error.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single sentence, but it is under-specified rather than concise. The 'Extended:' prefix is unexplained and the phrase 'GetBlobForAgentKV one record blob' is grammatically awkward, so the one sentence does not earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read tool with no annotations and no output schema, the description should explain what a blob is, how to obtain bcId/blobId, and what exactly comes back. It only hints at base64 data, leaving the agent without enough context to use the tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there are only two parameters, so the baseline is 3 even without description-level parameter detail. The description adds no meaning beyond the schema for bcId or blobId, but the schema already documents both.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a verb and resource ('GetBlobForAgentKV one record blob') and notes the return payload is base64 blob_data. However, 'one record blob' and 'GetBlobForAgentKV' are opaque, and it does not differentiate this tool from near-siblings like read_store_file or get_artifact_url. Adequate but vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives among the many sibling tools. An agent cannot infer from the text whether this is the right blob-fetching tool or when to prefer it over read_store_file.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_conversationC
Full verbatim user/assistant transcript (v0/conversation). v1 has no equivalent.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Agent id | |
| timeoutMs | No | Wait budget in ms (default 120000). Some very large transcripts never respond; use get_run result instead. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It states the return content and version availability but omits read-only semantics, timeout behavior, error handling, and the fact that very large transcripts may never return (that note lives only in a parameter schema).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no wasted words, and the main output description is front-loaded. The second sentence is terse but relevant to version navigation, though it could be clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter retrieval tool with no annotations and no output schema, the description identifies the content and a version constraint. It still leaves gaps around usage context and behavioral traits such as timeout limitations, so it is adequate but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds no parameter-level detail beyond what the schema already provides for id and timeoutMs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States it returns the full verbatim user/assistant transcript, a specific resource that distinguishes it from run-oriented siblings. It lacks an explicit action verb and the v0/v1 parenthetical is cryptic, but the output nature is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Offers no when-to-use guidance or named alternatives. The clause 'v1 has no equivalent' hints at API-version applicability but does not tell an agent when to choose this over get_run, get_run_stream, or prewarm_transcript.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_diff_detailsC
Extended: branch diff (GetBackgroundComposerDiffDetails).
| Name | Required | Description | Default |
|---|---|---|---|
| bcId | Yes | Agent bc id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing about read-only status, authentication, scope, return format, or side effects. It only restates the tool's apparent domain with an internal method name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short but under-specified rather than usefully concise. It is a fragment with an internal identifier, and the key routing information is absent rather than front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 1-parameter tool with no annotations and no output schema, the description should clarify what details are returned and how the branch diff relates to the agent/bcId context. Instead it gives only a terse label, leaving significant gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the only parameter bcId is documented as 'Agent bc id' in the schema. The description adds no parameter meaning beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description says it returns a branch diff, which gives a rough resource, but the leading 'Extended:' is unexplained and the parenthetical is an internal function name rather than an agent-facing purpose. It does not distinguish this from sibling tools like get_pr_diff, leaving the agent to infer the difference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool, when not to use it, or which sibling alternative to prefer. The description only names a diff type and does not relate it to get_pr_diff, get_pr_status, or any workflow context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_prC
Extended: account view of any PR (SCMService/GetPullRequest).
| Name | Required | Description | Default |
|---|---|---|---|
| prUrl | Yes | PR URL |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing beyond the name. It does not state whether this is a read-only fetch, what auth scope is required, what the response contains, or how it differs in behavior from the other PR tools.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short, but the two tokens present are low-value: the unexplained "Extended:" prefix and the internal service method reference occupy space that could have stated the actual purpose. Brevity here reflects under-specification rather than efficient structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description is the only behavioral signal available, and it supplies nothing about return contents or scope. For a simple single-parameter fetch this is minimally adequate at best, and the sibling overlap is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with a single required parameter (prUrl, "PR URL"), so the schema already documents input fully. The description adds no format guidance (e.g., full URL vs. host/path), so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Account view of any PR" restates the tool name without stating a clear verb+resource purpose; it is a near-tautology rather than a definition. The internal reference (SCMService/GetPullRequest) conveys nothing to an agent, and no distinction from siblings like get_pr_status, get_pr_diff, or get_diff_details is offered.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives despite three closely related PR/diff siblings. The agent is left to infer entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pr_diffC
Extended: PR file diffs (SCMService/GetPullRequestDiff).
| Name | Required | Description | Default |
|---|---|---|---|
| prUrl | Yes | PR URL |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not state that this is a read-only operation, whether any permissions or scopes are required, what happens with large diffs, or whether pagination applies; only the resource name weakly implies retrieval.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short, but it is a sentence fragment rather than a well-structured summary. The leading 'Extended:' label is undefined, and the SCMService method reference consumes space without helping an agent select or invoke the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema or annotations, the description is too thin. It says what resource is involved but leaves out usage context, return shape, and behavioral expectations that an agent needs when deciding between this and related PR/diff tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single prUrl parameter is documented in the schema as 'PR URL'. The description adds no further meaning beyond that, so the baseline of 3 applies when the schema already does the parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies PR file diffs, which is a specific resource, but it never states the operation as a clear verb and adds no sibling differentiation. Siblings like get_diff_details and get_pr are equally plausible for various diff/PR needs, and 'Extended:' plus the SCMService parenthetical do not help an agent distinguish this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no when-to-use guidance, no prerequisites, and no alternatives. An agent cannot tell from the text when to call get_pr_diff instead of get_diff_details, get_pr, or get_pr_status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pr_statusC
Extended: checks + review decision + counts.
| Name | Required | Description | Default |
|---|---|---|---|
| prUrl | Yes | PR URL |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It only lists output fields and says nothing about read-only status, authentication needs, rate limits, error behavior, or side effects; the verb 'get' in the name implies safe reading, but the description itself adds no behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short with no wasted words, but its structure is cryptic. Leading with 'Extended:' does not front-load a clear action or purpose, and the fragment omits essential context that a selection decision would need.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with one required parameter and no output schema, the description should at least clarify what is returned and when it is preferable to siblings. It hints at returned components but leaves the usage context and precise return shape incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single prUrl parameter is already documented as 'PR URL' in the schema. The description adds no syntax, format, or meaning beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives no clear verb or resource, only fragmentary output hints ('checks + review decision + counts'). 'Extended' is ambiguous and does not explicitly distinguish it from the sibling get_pr tool, though the word suggests a richer status payload. An agent can guess the purpose from the name but not from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus get_pr, get_pr_diff, or any other sibling. No prerequisites, alternatives, or exclusions are stated. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_runB
One run: status, result text, durationMs, pushed git branches/PRs.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Agent id | |
| runId | Yes | Run id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden; it helpfully discloses returned data fields but does not state that the operation is read-only, whether it requires authentication, or how missing/invalid run IDs behave. The 'get' naming implies a read, but behavior is not explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact fragment lists the return fields without filler; it is front-loaded and every word earns its place. No redundant restatement of the tool name or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-param read tool with no output schema, the description supplies the key return fields, which is useful. However, with no annotations and no usage guidance, it leaves gaps about sibling selection and operational preconditions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both id and runId, and the description adds no parameter-level detail such as where to obtain runId or whether id must match the run's agent. Baseline 3 applies because the schema already documents both parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a single-run retrieval and enumerates top-level return fields (status, result text, durationMs, branches/PRs), so the resource is clear. It stops short of a specific verb and does not distinguish from siblings like list_runs or get_run_stream.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or alternatives are given; the agent must infer that this is for fetching exactly one run rather than listing runs or streaming one. No prerequisites or exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_run_streamC
Fetch SSE run stream and summarize: status/assistant/thinking/tool_call/result events, tail text.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Agent id | |
| runId | Yes | Run id | |
| maxChars | No | Max assistant text chars (default 8000) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It hints that the output is a summary of event types rather than raw SSE bytes, which is useful, but it omits key traits: whether the call blocks until the run finishes, whether it works on in-flight runs, auth requirements, or truncation behavior. Significant gaps for a streaming tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence, front-loaded with the core action and the event summary that follows. No filler, though it is terse to the point of under-specifying behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Three parameters and no output schema, so the description's enumeration of summarized event types partially compensates. However, for a streaming tool with zero annotations and no output schema, it should clarify blocking/liveness behavior and result shape more than it does.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (id, runId, maxChars with default 8000) are already documented. The description's mention of 'tail text' loosely relates to maxChars' truncation role but adds no syntax or format detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Fetch SSE run stream') and adds what the result contains (status/assistant/thinking/tool_call/result events, tail text). This distinguishes it from get_run/list_runs reasonably well, though the description never explicitly names a sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No indication of when to choose this over get_run, list_runs, or get_conversation. The agent must infer that this is the streaming/summary view rather than a metadata fetch, and there is no guidance on prerequisites (e.g. whether the run must be complete first).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_usageB
Token usage for an agent, total + per run.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Agent id | |
| runId | No | Scope to one run |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the shape of the returned data (total plus per-run breakdown), which is genuinely useful. However, it never states that this is a read-only, side-effect-free lookup, nor does it mention auth requirements or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short phrase with no wasted words and the key scoping information front-loaded. It is arguably terse to the point of under-specification, which keeps it off a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with full schema coverage and no output schema, the description covers the essentials but is minimal. It hints at the return shape ('total + per run') without confirming what an omitted runId yields, leaving a small gap an agent must infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents both 'id' (agent id) and 'runId' (scope to one run). The description's 'per run' phrasing only loosely reinforces the runId semantics and adds no syntax or format detail, so it sits at the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase names a specific resource (token usage) scoped to an agent and clarifies the granularity (total plus per-run). The verb is only implicit in the name 'get_usage', and no sibling is named for differentiation, but the intent is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives, no prerequisites, and no mention of how it relates to siblings like get_agent, get_run, or list_runs. The only hint is that a runId scopes the result, which is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_agentsC
List Cloud Agents (v1, newest first). Filter triage queue.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max agents (default 20, max 100) | |
| cursor | No | Pagination cursor (nextCursor) | |
| includeArchived | No | Include archived (default true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses sort order (newest first) and implies pagination, but says nothing about permissions, rate limits, or what happens to the archived/active split. The word "List" weakly implies read-only, which is the only safety signal available.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact fragments, front-loaded with the core action and scope. Nothing is padded, though the dangling "Filter triage queue" clause is under-specified rather than economical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, fully-schema-documented list tool with no output schema, the description is minimally adequate: ordering and version are covered. It is incomplete in that the triage-queue filtering claim is never operationalized and no return-shape or pagination behavior is described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so limit, cursor, and includeArchived are already fully documented in the schema — the baseline of 3 applies. The description adds no parameter-level meaning and never connects "triage queue" to any of the three fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("List Cloud Agents") and adds scope detail (v1, newest first) that distinguishes it from the sibling list_agents_v0. The trailing "Filter triage queue" fragment is too terse to fully explain what the tool does, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance, and no alternative is named despite list_agents_v0, get_agent, and archive_agent being plausible siblings. "Filter triage queue" hints at a triage use case but never explains the condition or which parameter triggers it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_agent_storesD
Extended: Project Context stores (ListAgentStores).
| Name | Required | Description | Default |
|---|---|---|---|
| pageToken | No | Page token |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full behavioral burden. It provides no information about pagination, authentication, side effects, or return behavior beyond the implicit listing implied by the name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short but consists of an under-specified fragment; 'Extended' does not add useful meaning. Brevity here reflects missing content rather than efficient front-loading.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list operation over project context stores, the description omits what a store is, how results relate to agents, and pagination behavior. No annotations or output schema compensate for these gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single pageToken parameter is documented in the schema. The description adds no syntax, format, or additional meaning beyond what the schema already provides, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description merely restates the tool name in parentheses and labels it 'Extended', without a clear verb or scope. It does not distinguish this operation from siblings like list_agents or list_workspace_files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to use this tool rather than list_agents, list_workspace_files, or read_store_file. No prerequisites, conditions, or exclusions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_agents_v0A
Legacy v0 list with repo/branch/PR/summary in one round-trip. Best for PR triage.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max agents (default 20) | |
| cursor | No | Pagination cursor |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It usefully adds the 'one round-trip' efficiency trait and flags the tool as legacy, but says nothing about read-only safety, auth/permission needs, or pagination behavior beyond what the schema shows.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler; the legacy/scope qualifier and the recommended use case each earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, no-annotation list tool with a fully documented schema, the description covers the value proposition and hints at return fields (repo/branch/PR/summary), but with no output schema the return shape and cursor semantics remain partially unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters (limit default 20, cursor), so the schema already documents them fully. The description adds no parameter-level detail, which is the expected baseline when the schema does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies this as the 'Legacy v0 list' and names the payload it returns (repo/branch/PR/summary), which together with the tool name makes the resource clear and distinguishes it from the sibling list_agents. It stops short of an explicit verb+resource statement like 'lists agents', relying on the name to supply that.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Best for PR triage' provides an implied use case, and 'Legacy v0' hints that list_agents is the preferred alternative. However, it never states when to choose this over list_agents or when not to use it, leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_artifactsC
List screenshots/videos/logs under artifacts/.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Agent id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it discloses almost nothing: no statement of read-only safety, no pagination/limit behavior, no indication of whether results are scoped to the supplied agent id or span the workspace. Only the artifact location hint is added beyond what the name conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler, front-loaded with the verb and resource. It is appropriately sized, though its terseness borders on under-specification rather than genuine economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description must explain the return shape; it says nothing about the fields, ordering, or pagination of the returned artifact list. For a listing tool whose output is the entire point, this leaves a significant gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter exists and the schema documents it fully ('Agent id'), so the schema does the heavy lifting. The description adds no meaning about how the id scopes the listing (e.g. only that agent's artifacts), so baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (artifacts), and enumerates what counts as an artifact (screenshots/videos/logs) as well as the storage location (artifacts/). It is clear what the tool does, though it never names or distinguishes itself from the natural sibling get_artifact_url, which fetches a single artifact.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus get_artifact_url or get_blob, nor any prerequisite (e.g. that the agent id must exist). The agent is left to infer that listing precedes fetching a URL.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsA
Recommended model ids for create_agent.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It signals the payload is a set of recommended ids (so a read-only, non-mutating call), but says nothing about return format, cardinality, or whether recommendations are workspace-specific. For a zero-parameter listing tool the exposure is small, so this is adequate but thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single fragment with no filler and the purpose front-loaded. It is efficient, though it borders on under-specification rather than tight phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument, no-output-schema lookup whose result type is stated ('model ids'), the agent has what it needs to call it. Only the return shape remains unspecified, which is a minor gap for a tool this simple.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema declares zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies; no parameter-level guidance is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (model ids) and the concrete purpose (feeding create_agent), which distinguishes it from the many agent/run-oriented siblings. It is clear but telegraphically terse, and the 'list' verb itself only comes from the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'for create_agent' implies the usage context, so an agent can infer this is a lookup to run before creating an agent. There is no explicit when-not guidance or named alternative, so usage is only implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_pending_followupsD
Extended: account queue (ListPendingFollowups).
| Name | Required | Description | Default |
|---|---|---|---|
| bcId | Yes | Agent bc id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations and the description discloses nothing about behavior: not whether it is read-only, whether it requires authentication, what the return looks like, pagination, or rate limits. It carries the full burden and fails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is short but uninformative; it is under-specified rather than concise. The cryptic 'Extended: account queue' prefix is not front-loaded with useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list operation with one required parameter and no output schema, the description should at least explain the purpose, scope, and return type. It provides none of this, leaving the agent unable to confidently invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the only parameter bcId is described in the schema as 'Agent bc id'. The description adds nothing beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Extended: account queue (ListPendingFollowups)' does not state a clear verb+resource; it restates the internal function name and adds an opaque 'Extended' label. It gives no sense of what the tool does beyond the tool name, and does not distinguish it from siblings like create_followup_run or list_runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. No context about what 'pending followups' are or when the account queue should be queried.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_runsB
List runs (turns) for an agent, newest first.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Agent id | |
| limit | No | Max runs (default 20) | |
| cursor | No | Pagination cursor |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the ordering ('newest first') but says nothing about permissions, pagination behavior beyond the cursor param, rate limits, or whether runs can be empty. For a list tool with no annotations, more context is expected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence, front-loaded with the verb and resource, and no wasted words. The ordering detail is included efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with a fully documented schema and no output schema, the description covers the essential operation. However, with no annotations and no other context, it leaves behavioral aspects like permissions and pagination edge cases unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents 'id', 'limit', and 'cursor'. The description adds no parameter-level meaning beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('runs (turns) for an agent'), and clarifies that runs are turns. It does not explicitly distinguish itself from sibling 'get_run' or 'list_agents', though the 'for an agent' scope makes the target clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the name and description (list runs for a given agent), but there is no explicit when-to-use guidance or mention of the alternative 'get_run' for fetching a single run. With 26 sibling tools, explicit routing would help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_workspace_filesD
Extended: live VM workspace (ListWorkspaceFiles).
| Name | Required | Description | Default |
|---|---|---|---|
| bcId | Yes | Agent bc id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it discloses nothing behavioral: not whether it is read-only, whether it requires a running/prewarmed VM, what it returns, pagination, or failure modes. 'live VM workspace' hints at environment but not behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is short but not informative—it is under-specified rather than concise. The fragmentary 'Extended:' prefix and redundant parenthetical consume the whole budget without conveying usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and a crowded sibling namespace, the description omits everything an agent needs to invoke it correctly: triggering conditions, return shape, and environment assumptions. It is not adequate to the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (bcId) exists and schema description coverage is 100%, so the schema already documents it adequately. The description adds no meaning beyond the schema, which is the expected baseline of 3 when the structured data does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Extended: live VM workspace (ListWorkspaceFiles)' essentially restates the tool name and adds a cryptic parenthetical. It never states a verb+resource plainly (e.g., 'lists files in the agent's VM workspace'), so an agent must infer the operation from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to use this tool versus the many sibling listing tools (list_artifacts, list_agent_stores, read_store_file, get_blob). No context, prerequisites, or exclusions are given, though nothing is actively misleading.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prewarm_transcriptD
Extended: StreamConversation PREWARM initial_state + blob ids. Raw record behind the transcript viewer.
| Name | Required | Description | Default |
|---|---|---|---|
| bcId | Yes | Agent bc id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden and largely fails it. It only vaguely gestures at return content ('initial_state + blob ids'), with no statement of whether the call is read-only, whether it mutates or warms server state (the name suggests PREWARM), what auth scope is needed, or how large/failure-prone the response is.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short but telegraphic fragments ('Extended: ...', 'Raw record behind the transcript viewer') rather than front-loaded explanation. Brevity here is under-specification, not economy; the reader gains no actionable framing from either fragment.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and a single undocumented-in-prose parameter, the description should define what is returned and under what conditions it is called. It leaves the core question — what does an agent get back and when should it use this versus get_blob or get_conversation — entirely unanswered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
High schema coverage (100%) sets a baseline of 3, but the description adds nothing about bcId — it does not clarify whether this is an agent ID, a blob container ID, or the same 'bc id' used elsewhere. The isolated mention of 'blob ids' muddies rather than clarifies the parameter's role.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is telegraphic jargon: 'StreamConversation PREWARM initial_state + blob ids' names an internal artifact rather than a concrete verb+resource. An agent cannot confidently say whether this reads a cached prewarm record, warms a cache, or fetches transcript data, and nothing distinguishes it from siblings like get_blob or get_conversation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, prerequisites, or alternative is stated. The sibling set (get_blob, get_conversation, get_artifact_url) overlaps plausibly with this tool's territory, yet the description offers no routing guidance at all.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_lineageC
Extended: Project lineage = ListWorkersForManager + ListBackgroundComposerChildren. Workers, side chats, subagents.
| Name | Required | Description | Default |
|---|---|---|---|
| managerBcId | Yes | Project/root bc id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full behavioral burden. It implies a read-only aggregation of two lists but does not state permissions, side effects, rate limits, or return format, leaving key safety and operational traits undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact line that front-loads the key composition and returned entities. The 'Extended:' prefix is slightly cryptic but does not detract significantly from the otherwise efficient structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple aggregate-list tool with one required parameter and no output schema, the description states what is returned but omits return structure, pagination, and read-only status. It is minimally adequate but leaves clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the single parameter (managerBcId) already documented as 'Project/root bc id.' The description adds no further parameter semantics, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description defines project lineage as the union of two specific list operations and enumerates the returned entity types (workers, side chats, subagents). While it relies on internal function names rather than a clean verb+resource phrase, the scope is clear enough to distinguish it from sibling listing tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives such as list_agents or start_side_chat. The implied usage is to retrieve a comprehensive project lineage, but no conditions or exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_store_fileC
Extended: ReadAgentStoreFile (Project Context / shared notes).
| Name | Required | Description | Default |
|---|---|---|---|
| storeId | Yes | Store id | |
| relativePath | Yes | e.g. notes/plan.md |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure, and it discloses nothing: no permission requirements, no indication of what happens on a missing path, no return format, no size limits. Only the word 'Read' hints at a non-destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short, but brevity here reflects under-specification rather than conciseness — the single cryptic fragment omits everything an agent needs and wastes space on the meaningless 'Extended:' prefix.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must compensate and does not. For a two-required-parameter file read tool, an agent cannot tell the expected path root, error behavior, or output shape from this text alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the two parameters (storeId, relativePath) are documented in the schema, including an example path. The description adds no parameter meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The string 'ReadAgentStoreFile' conveys a read operation against a store file, and the parenthetical '(Project Context / shared notes)' hints at the content domain. However, the 'Extended:' prefix is internal jargon that adds no meaning, and the description never states in plain language what the tool returns or how it relates to siblings like get_blob or list_workspace_files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the many sibling listing/fetching tools (list_agent_stores, list_workspace_files, get_blob). The only hint is the implicit pairing of storeId + relativePath, which the agent must infer on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_side_chatD
Extended: StartSideChatBackgroundComposer.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Side chat name | |
| parentBcId | Yes | Parent bc id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden and delivers nothing: no statement that this creates/mutates state, no permission requirements, no side effects, no return behavior. The 'Extended:' prefix adds no usable information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short, but brevity here comes from under-specification rather than conciseness — a single fragment consisting of a prefix and an internal class name. Not a sentence an agent can act on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and a near-empty description, an agent has essentially no context for a tool that apparently creates a background composer session. The definition is inadequate for the operation's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (name, parentBcId) are documented in the schema itself. The description adds no meaning beyond that, which meets the baseline of 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description restates the tool name as an internal class identifier ('StartSideChatBackgroundComposer') rather than explaining what the tool does. An agent can guess it starts a side chat, but nothing distinguishes it from siblings or clarifies what a 'side chat' or 'background composer' is.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance at all on when to use this tool versus alternatives, nor any prerequisite or exclusion information. The one required parameter (parentBcId) implies a parent context, but the description never explains it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
steer_agentC
Extended: InjectBackgroundComposerContext (steer / queue message).
| Name | Required | Description | Default |
|---|---|---|---|
| bcId | Yes | Agent bc id | |
| text | Yes | Steering text | |
| expectedRunId | No | Expected run id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It indicates an action ("steer / queue message"), but does not disclose permissions, mutating effects, expectedRunId handling, reversibility, or response behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short, but much of that brevity comes from cryptic internal terminology rather than useful front-loaded information. "Extended: InjectBackgroundComposerContext" does not help an agent understand the operation efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a likely mutating agent-control tool with three parameters and no annotations or output schema, the description is largely incomplete. The schema covers parameters, but purpose, usage, and behavioral impact remain unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so bcId, text, and expectedRunId are documented in the schema itself. The description adds no additional parameter meaning, making the baseline of 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is mostly internal jargon: "InjectBackgroundComposerContext" plus "steer / queue message." It hints at sending steering text to an agent, but does not clearly define the resource, scope, or how it differs from siblings such as create_followup_run or list_pending_followups.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. The phrase "steer / queue message" implies a context, but no prerequisites, exclusions, or sibling comparisons are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
29 tool updates
v0.1.0- First observed
agent_lifecycle - First observed
archive_agent - First observed
cancel_run - First observed
create_followup_run - First observed
cursor_status - First observed
get_agent - First observed
get_artifact_url - First observed
get_blob - First observed
get_conversation - First observed
get_diff_details - First observed
get_pr - First observed
get_pr_diff - First observed
get_pr_status - First observed
get_run - First observed
get_run_stream - First observed
get_usage - First observed
list_agent_stores - First observed
list_agents - First observed
list_agents_v0 - First observed
list_artifacts - First observed
list_models - First observed
list_pending_followups - First observed
list_runs - First observed
list_workspace_files - First observed
prewarm_transcript - First observed
project_lineage - First observed
read_store_file - First observed
start_side_chat - First observed
steer_agent
TDQS
Scored across 29 tools
Most tools target distinct resources and actions, and descriptions clarify differences. The main overlap is list_agents vs list_agents_v0 (legacy version), with minor potential confusion between get_run/get_run_stream and PR-related tools, but these are mostly resolvable by descriptions.
Most names follow snake_case verb_noun conventions (list_*, get_*, create_*, cancel_*), but several are noun-first or state-oriented (cursor_status, project_lineage, agent_lifecycle, prewarm_transcript), creating a mixed but still readable pattern.
29 tools is heavy for the server's apparent scope. The set includes legacy variants (list_agents_v0) and many extended endpoints, pushing well beyond the 3-15 range that tends to be well-scoped.
A critical gap is the absence of a create_agent tool, despite list_models explicitly referencing create_agent. There is also no update_agent or initial run creation, leaving core lifecycle operations incomplete despite extensive read and extended coverage.
Maintenance
Related MCP Connectors
Governed MCP gateway: one endpoint for your tools, with credential custody and audit log.
One MCP endpoint for Claude, GPT & Gemini: 100+ tools + no-code connectors + agent workers.
- QuallaaOAuthcom.quallaa
Talk to your public-facing AI from any MCP client — Claude, ChatGPT, Cursor, Cline, Windsurf.
Your org's AI agents, tasks, runs, search, and brain files as MCP tools and resources.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables management of Brainbase agents, components, evals, tasks, and more via a remote MCP connection with OAuth authentication.2MIT
- FlicenseNot gradedqualityCmaintenanceEnables agents to connect to remote MCP servers once, access their tools through a compact MCP endpoint, pair a CLI inside sandboxes, and create watches that turn command or tool output into pollable structured events.2-
- AlicenseNot gradedqualityBmaintenanceEnables agent clients to safely connect to tools and execution resources through MCP with authorization, approvals, audit, chat-context isolation, SSH/Docker access, and long-running command session tracking.MIT

meshive-mcpofficial
AlicenseNot gradedqualityBmaintenanceEnables MCP-capable agents to manage Meshive GPU Cloud resources, including account, workspaces, pods, storage, GPUs, templates, serverless deployments, tasks, assets, machines, and billing history.Apache 2.0