dialogbrain
Server Details
Unified inbox MCP for WhatsApp, Telegram, Email, voice — read/send messages, search, AI agents.
- Status
- Healthy
- Uptime
- 20.2% over 54 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- saloprj/dialogbrain-mcp
- GitHub Stars
- 0
TDQS
Scored across 277 tools
Despite very detailed descriptions, several clusters overlap heavily: web_search / web_research / web__local_search all answer queries from the web, ai_filters_* (semantic trigger filters) vs ai_tags_* (auto-classification tags) are easily conflated, and agent_handoff / agents_ask / agents_simulate_inbound all route work to agents. With 277 tools, boundaries between search, query, note, and prompt operations blur, so misselection is likely.
Names largely follow a predictable snake_case namespace_action pattern (messages_send, calls_make, group_add, browser_click). Minor deviations exist — singular agent_handoff vs plural agents_*, a double underscore in web__local_search, and a few verbless names like agents_activity — but the convention is readable and mostly systematic.
277 tools is an extreme mismatch for any single server surface, far beyond the 50-tool threshold. Even for a broad multi-channel AI platform this dilutes discoverability and forces agents to guess among many near-neighbors.
Coverage is very broad: CRUD for agents, prompts, collections, threads, files, contacts, tasks, workspaces, plus calls, campaigns, social publishing (IG/TikTok/YouTube/X/Threads), browser automation, and analytics. Minor gaps exist (e.g. no unified cross-entity search, no LinkedIn list/analytics beyond raw requests), but core lifecycles are well covered.
Available Tools
277 toolsagent_handoffARead-onlyIdempotentInspect
Delegate a multi-step task (research, composing messages, booking, scheduling) to the full agentic planner. Use when a user ask needs more than a direct answer. The specialist runs synchronously — its response is already shown to the user in real-time. Summarize the OUTCOME in past tense (e.g. 'The Media Creator generated your video' or 'The Document Composer failed because...'). Do NOT say 'I will delegate' — the delegation already happened. If status is timeout or error, explain what went wrong and offer to retry.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Execution mode: 'sync' (wait for result, default) or 'async' (fire and forget, child runs in background). Async is only available in background/trigger context. | sync |
| payload | No | Optional structured data for the target agent. For a rule_based (script) target this becomes the script's inputs['raw_data'] verbatim — pass the exact fields its script reads (same contract as the trigger event that script normally handles). For LLM targets it is appended to the task text as a [PAYLOAD] JSON block. VOICE callers: the voice pipeline's strict schemas seal free-form objects, so from a live call put the data in task_description instead — script targets receive it as inputs['message_text']. | |
| agent_id | No | Optional ID of another agent in the same workspace to delegate the task to. When set, this becomes cross-agent delegation; the target agent runs with ITS OWN prompt, tools, and model. Use this for specialty tasks (see agents.list to discover specialists). Prefer the in-loop variant (no `agent_id`) for one-off escalations. Spawns a new trace linked back to this trace via parent_trace_id (visible in the admin lineage card). | |
| target_slug | No | Optional stable slug of a system-template specialist to delegate to (e.g. 'doc-composer' for the Document Composer). Env-portable alternative to agent_id — resolves the workspace's fork of that template (auto-forking on first use). Used by async handoffs that target a specialist without knowing its per-workspace id. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| task_description | Yes | Plain-language description of what the planner should accomplish. Include everything the planner needs: the user's goal, constraints, and any context already gathered in this voice call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing that the specialist 'runs synchronously' and that its response is 'already shown to the user in real-time,' plus the existence of `timeout`/`error` statuses and the expected post-call behavior. It adds no information on auth requirements, rate limits, recursion/depth limits, or what a failed delegation returns, which keeps it out of the top score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and trigger are front-loaded in the first two sentences, and the remaining sentences each address a distinct failure mode (past-tense summary, not saying 'I will delegate', error handling). Slightly dense with imperative post-call instructions, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter, side-effect-bearing delegation tool with no output schema, the description covers the odd sync semantics, the fact that output is already user-visible, and error states — the highest-risk gaps. It omits guidance on choosing between agent_id and target_slug in description text, though the schema covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (mode, payload, agent_id, target_slug, in_workspace, task_description) is already documented in the schema with more detail than the description offers. The description adds no parameter-level meaning beyond that, which is the expected baseline when the schema carries the full burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Delegate a multi-step task ... to the full agentic planner' — with concrete examples (research, composing messages, booking, scheduling) and a clear scope ('more than a direct answer'). It does not, however, distinguish itself from closely related siblings such as calls_dispatch_agent, agents_ask, background_run, or job_escalate, leaving the agent to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger condition ('Use when a user ask needs more than a direct answer') and rich guidance on what to do after the call, including how to phrase the result and how to handle `timeout`/`error` statuses. It stops short of naming when NOT to use it or which sibling (e.g. calls_dispatch_agent, agents_ask) to prefer instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_activityARead-onlyIdempotentInspect
See what you — or another agent in your workspace — actually did over a time window: messages sent, documents created, calls made, plus a summary (run counts, per-day, top tools). Use this to answer 'what did I do today / yesterday / last week / in the last hour?' or 'what did do?' with real data instead of guessing.
Omit agent for your own activity, or pass another workspace agent's name, slug, or id. Pass since/until as ISO datetimes (e.g. '2026-06-03T09:00:00') for sub-day windows like the last hour, or plain dates ('2026-06-03') for whole days — compute them from the current date/time you were given. Defaults to the last 24h. Traces are retained 30 days.
Times are interpreted as UTC — if the current time you were given is in another timezone, convert to UTC before passing since/until.
| Name | Required | Description | Default |
|---|---|---|---|
| agent | No | Target agent: name, slug, or numeric id. OMIT for yourself. | |
| limit | No | Max actions / recent runs to return. | |
| since | No | Window start — ISO datetime or date. OMIT for last 24h. | |
| until | No | Window end — ISO datetime or date (exclusive day-end for a bare date). OMIT for now. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive behavior, so the description focuses on useful extras: 30-day trace retention, a default 24h window, and the UTC interpretation of since/until — all genuinely beyond the structured fields. It omits explicit auth/error behavior, but retention and timezone semantics are real added context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what it does, then parameters, then the timezone caveat — a sensible order. It is slightly padded (the 'compute them from the current date/time you were given' instruction recurs near the timezone note), but every sentence carries useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with no output schema, it covers defaults, retention, timezone handling, and a sketch of the return shape (run counts, per-day, top tools), which is enough to call it correctly. The only uncovered item is in_workspace, and its session-scoping caveat is left to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning beyond the schema: ISO vs bare-date formats for since/until, how to compute them from the supplied current time, and that omitting `agent` means 'yourself'. It does not touch in_workspace or limit, so not a full 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource — 'see what you or another agent actually did over a time window' — and enumerates the returned content (messages, documents, calls, plus a run-count/per-day/top-tools summary). It is unambiguous what the tool produces, but it never names or contrasts a sibling (e.g. agents_traces_list / agents_traces_stats), so an agent must infer the boundary from the returned payload.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is explicit and framed around concrete questions ('what did I do today / yesterday / last week / in the last hour?' or 'what did <agent> do?'), which tells the agent exactly when to reach for it. There is no exclusion guidance or routing to overlapping trace tools, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_add_fileAInspect
Attach a file to this agent's private knowledge (agent-specific files, not shared with other agents).
Workflow:
Upload the file with files_upload (pass source_url for remote files)
Index it with files_ingest (pass the file_id)
Call this tool with agent_id + file_id
Returns chunk_count — shows 0 while still processing. Call agents.list_files later to see the final chunk count once indexing completes.
| Name | Required | Description | Default |
|---|---|---|---|
| file_id | Yes | file_id returned by files_upload or files_ingest | |
| agent_id | Yes | ID of the agent to attach the file to | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it is a write (readOnlyHint=false), not destructive, and not idempotent, so the safety profile is covered. The description adds genuinely new behavioral context: indexing is asynchronous, the returned chunk_count reads 0 while processing, and final state must be checked later via agents.list_files. It omits failure modes such as re-attaching a duplicate file or a file that was never ingested.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in one sentence, followed by a tight three-step numbered workflow and a single return-value caveat. Every sentence carries information; nothing is restated or padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by explaining the chunk_count return and its transient zero value, and the pipeline prerequisites are fully spelled out. It leaves minor gaps around duplicate attachments and permission requirements, but is otherwise sufficient to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so file_id, agent_id, and in_workspace are already documented in the schema, including that file_id comes from files_upload or files_ingest. The description reinforces the file_id provenance through its workflow but adds no new syntactic or semantic detail beyond the structured fields, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb and resource ("Attach a file to this agent's private knowledge") and immediately scopes it against the sibling concept by clarifying these are agent-specific files "not shared with other agents." An agent can distinguish this from collections_add_file or a shared-knowledge tool without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The numbered workflow explicitly sequences prerequisites (files_upload, then files_ingest, then this tool) and names the exact parameters to pass at each step, which is strong when-to-use guidance. It stops short of stating when NOT to use this tool (e.g. shared collections vs. agent-private knowledge), so it falls just short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_approve_draftADestructiveInspect
Approve a pending agent draft and send the message.
The draft will be sent to the conversation it was generated for. You can optionally edit the text before sending.
Use this when user says:
'Approve this draft'
'Send this reply'
'Approve and send'
'Looks good, send it'
IMPORTANT: This will send a message to a real person.
| Name | Required | Description | Default |
|---|---|---|---|
| draft_id | Yes | ID of the draft to approve | |
| edited_text | No | Optional edited response text (if user wants to modify before sending) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true and idempotentHint=false, so the safety profile is covered structurally. The description adds genuinely new context: the message goes to the draft's originating conversation, edited_text can substitute, and it reaches a real person. It doesn't state irreversibility or any undo path, keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and consequence in the first sentence, followed by the effect, then routing phrases, then the warning. The four trigger-phrase bullets are slightly repetitive but earn their place as routing cues. No filler beyond that.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers what is affected (the draft's conversation), the optional modification path, and the real-world consequence. With annotations supplying the destructive/non-idempotent profile, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so draft_id, edited_text and in_workspace are already documented in the schema. The description only alludes to the optional edit ('You can optionally edit the text before sending'), adding no syntax, format, or constraint detail beyond the schema. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('Approve a pending agent draft') plus the consequential effect ('send the message'), so the agent knows exactly what the call does. It does not name the obvious alternative (agents_reject_draft) or any other sibling, so it falls short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete trigger phrases ('Approve this draft', 'Send this reply') that map user intent to this tool, which is strong positive routing guidance. It never states when NOT to use it (e.g., use agents_reject_draft if the user wants changes), so the exclusion side is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_askAInspect
Send a message to an AI agent and get its response.
The agent runs with its configured prompt, tools, and knowledge. Use this to test agents or have them process a task.
Returns: {status: 'replied'|'silent', response_text, messages[], full_reply, model_used, tokens_*, send_mode, execution_mode, tool_calls[]}. tool_calls[] is the per-tool trace in call order — each {tool, success, error, duration_ms} — so you can see which tool the agent ran and why it failed (e.g. a workbench script error) directly from this response, no trace lookup needed. messages[] carries each messages.send invocation the agent made (text, subject, reply_to_message_id, timestamp, message_id, attachments=[{file_id,name,mime}]). full_reply concatenates text only — attachment-only sends show up in messages but not full_reply. status='silent' iff both response_text is empty AND messages is empty.
Execution may take 10-60s depending on agent complexity. For runs that may exceed ~2 minutes (heavy multi-step agents), pass background=true: the call returns immediately with status='started' and the run continues server-side, detached from this connection — poll agents.traces_list / agents.trace_get for the outcome and agents.list_drafts for produced drafts.
| Name | Required | Description | Default |
|---|---|---|---|
| message | Yes | Message/goal to send to the agent | |
| agent_id | Yes | ID of the AI agent to ask | |
| send_mode | No | Send mode for the agent run: 'draft' = create drafts, 'auto' = send directly. Defaults to the agent's configured default_send_mode. Does NOT change execution_mode — that is fixed by the agent's config. | |
| background | No | Run detached from this MCP connection. Returns immediately with status='started'; the run survives client timeouts and disconnects (up to 15 min). Poll agents.traces_list for the outcome. Use for runs expected to exceed ~2 minutes — a synchronous call is cancelled when the MCP request dies. OMIT to run synchronously and get the answer in this call (the default). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| attachment_file_ids | No | File ids to attach to the message, as a customer would have sent them. Use this to test how the agent handles a photo or a document: a vision-capable model receives the image itself, a text-only one receives a one-line description of it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only give the coarse safety profile (non-read-only, non-destructive, non-idempotent, closed-world). The description adds real behavioral context: execution takes 10-60s, background runs detach from the connection and survive up to 15 min, status='silent' is defined precisely, and it explains that a synchronous call is cancelled if the MCP request dies. This is well beyond what the annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first line, then return shape and execution semantics. It is dense and the return-value enumeration is long, but nearly every clause carries usable information (status semantics, full_reply vs messages distinction, attachment handling). Slightly verbose for a definition but well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of explaining returns and does so thoroughly (status values, messages[], full_reply, tool_calls[] trace, model/tokens). Combined with the background/synchronous tradeoff and attachment behavior, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters, and much of the description's parameter text (send_mode vs execution_mode, background, attachments) mirrors the schema. It adds marginal framing but does not materially extend meaning beyond the structured fields, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: send a message to an AI agent and get its response, plus the scope that the agent runs with its own prompt/tools/knowledge. An agent can tell this apart from pure-read siblings like agents_get or agents_traces_list. It stops short of explicitly naming the sibling alternatives (e.g. agents_simulate_inbound, calls_dispatch_agent), so 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: 'Use this to test agents or have them process a task.' It also gives conditional guidance for background=true (runs expected to exceed ~2 minutes) vs the synchronous default. It does not name when NOT to use it or point at specific alternative siblings, so it lands at 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_createAInspect
Create a new AI agent in the workspace.
Execution modes:
ai_assisted (default): Two-phase AI — fast pre-classifier (Haiku) for keyword filtering and simple replies, then full AI with tools for complex messages.
agentic: Autonomous multi-step agent with planning and tool execution.
rule_based: Simple pattern matching without AI.
Keyword filtering is available in ai_assisted mode via keywords in trigger conditions (free, deterministic) and/or auto_reply_rules (LLM-based) set through agents.update.
Pass prompt_text for the agent's instructions (stored inline on the agent) or prompt_id to link an existing prompt row, not both.
Tools: a new agent starts with the standard tool set (knowledge base lookup, replying, lead capture, reading files, handoff, plus the default voice-call tools). Pass allowed_tools to replace that set.
From a template: pass template (e.g. 'dm-auto-reply') to create the agent AND its built-in trigger in one call — deterministic, no need to add a trigger separately. When template is set, name/text_engine/send_mode come from the template.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Name of the AI agent (1-100 characters) | |
| model | No | LLM the agent runs on. OMIT to use the platform default (deepseek-v4-flash-nothink). Applied right after creation so you don't need a follow-up agents_update. | |
| template | No | Optional template slug to instantiate the agent + its built-in trigger from (deterministic). Use 'dm-auto-reply' for the customer DM auto-reply agent (incoming DM trigger, draft mode). When set, the trigger comes from the template — you don't need agents.trigger_create. OMIT to create a plain agent (no template). | |
| prompt_id | No | ID of the prompt to assign to this agent | |
| send_mode | No | Default send mode: 'auto' or 'draft'. OMIT to use 'draft' (the default). | |
| description | No | Optional description of what this agent does | |
| prompt_text | No | The agent's instructions, stored inline on the agent (this is what drives how it replies). Works with or without `template`; no separate prompts tool needed. Write it from the business to its customers. | |
| remote_tool | No | Only for text_engine='external_agent': the workspace integration tool that starts the remote agent, e.g. 'ext42_run_routine' (list them with integrations.search_tools). On each trigger DialogBrain calls it once with the event, the agent's instructions and the ids to answer with; the remote agent replies through the DialogBrain MCP tools (messages.send, tasks.comment, agents.task_complete). The endpoint and its secret belong to the integration, not to the agent. | |
| text_engine | No | Text-execution engine: 'rule_based', 'ai_assisted', 'agentic' (default), 'claude_channels', or 'external_agent' (the run is handed to an agent outside DialogBrain; needs remote_tool). Voice is derived from triggers, not engine. OMIT to use the default ('agentic'). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| allowed_tools | No | Explicit allow-list of tool IDs the agent may call on triggered runs (e.g. ['knowledge.query', 'messages.send']). OMIT to give the agent the standard set: knowledge.query, messages.send, contacts.capture_lead, files.read, agent.handoff (plus the default voice-call tools). When passed, the list REPLACES the standard set entirely — include 'knowledge.query' for an agent that must answer from its knowledge base, or the model has no such tool and tends to imitate the call in its reply text. An empty list leaves the agent with no tools. | |
| max_iterations | No | Hard cap on agentic-loop turns per run (1-50). OMIT for the default (10). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint=false, idempotentHint=false, destructiveHint=false) already establish that this is a non-idempotent write, and the description usefully adds that `model` is applied post-creation (no follow-up update), that `allowed_tools` REPLACES the standard set, and that an empty list leaves no tools. Against that, it introduces a real ambiguity: the body calls `ai_assisted` the default mode while the schema declares `agentic` the default and omits `ai_assisted` from the enum entirely, muddying the core behavior of an omitted text_engine.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then uses labeled bullets for modes and short paragraphs for tools/templates, which is well structured for a 12-parameter tool. It repeats a fair amount of schema text (model, template, allowed_tools) and the execution-mode list carries the misleading default claim, costing the top score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should ideally say what creation returns (agent id/handle), which it omits. Otherwise it covers the complex boolean-like interactions (template suppressing name/text_engine/send_mode, tool-list semantics, engine modes) well enough for an agent to call it correctly, modulo the default-mode inconsistency.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description earns above that by adding semantics the schema lacks, notably the mutual exclusivity of `prompt_text`/`prompt_id` and the replacement (not additive) behavior of `allowed_tools`. It does not add much to `model`, `max_iterations`, or `in_workspace` beyond what the schema text already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new AI agent in the workspace') and immediately distinguishes itself from siblings by covering execution modes, template instantiation, and tool assignment. An agent can tell it apart from agents_update, agents_update_from_template, and agents_trigger_create without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear routing conditions: use `template` to also get the built-in trigger ('no need to add a trigger separately'), use `prompt_text` or `prompt_id` but not both, and reference to agents.update for keyword/reply-rule config. It does not, however, state when to prefer this over agents_update or agents_update_from_template for an existing agent, leaving one sibling choice to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_deleteADestructiveIdempotentInspect
Permanently delete an AI agent.
WARNING: This cannot be undone. The agent and all its triggers will be removed.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | ID of the agent to delete | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered structurally. The description adds genuinely new context beyond them: the operation is irreversible and cascades to delete all triggers, which the annotations do not convey. It stops short of mentioning auth requirements or effects on references.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler. The permanence warning is front-loaded and the cascade consequence follows immediately, so the most decision-relevant information hits first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no output schema and full annotation coverage, the description supplies the two facts that matter most: irreversibility and cascading trigger removal. It omits secondary concerns like workspace scoping behavior and permission requirements, keeping it just short of complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters (agent_id, in_workspace) are fully documented in the schema itself. The description adds no parameter-level detail, which is acceptable at full coverage but earns only the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Permanently delete an AI agent') and implicitly distinguishes itself from siblings like agents_trigger_delete and agents_update by noting the cascade to triggers. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to prefer this over alternatives such as agents_update or agents_trigger_delete, and no prerequisites or required permissions stated. The warning implies caution but not selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_getARead-onlyIdempotentInspect
Get detailed information about a specific AI agent.
Returns full agent config including:
Execution configuration
Tool configuration
Knowledge configuration
Escalation configuration
Triggers list
Knowledge collections
Custom AI instructions (prompt_text)
Auto-reply rules override (auto_reply_rules)
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | ID of the AI agent to fetch | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description goes beyond that by disclosing the shape of the payload — execution, tool, knowledge, escalation config, triggers, collections, prompt_text and auto_reply_rules — which materially helps an agent decide whether this call has the data it needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The one-line purpose is front-loaded and the returned-config bullets are scannable. The list is somewhat long but each item names a distinct config area an agent would otherwise have to discover by calling the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by enumerating the returned config sections, which is the main thing an agent needs to know before calling. Combined with complete annotation coverage and full parameter documentation, the definition is nearly self-sufficient; only error/not-found behavior is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters, including the workspace-scoping semantics of in_workspace, so the schema carries the burden. The description adds no parameter-level detail (no format, no constraints on agent_id), making this a baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get detailed information about a specific AI agent') and the bullet list pins down exactly which config sections are returned, so the agent knows this is a full-config fetch rather than a summary. It does not explicitly name siblings such as agents_list or agents_trace_get, so differentiation is inferred from 'a specific agent' rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance and no named alternative. The description never tells the agent to use agents_list for enumeration or agents_update for modification, nor does it state the prerequisite of having a valid agent_id.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_silenceARead-onlyIdempotentInspect
End this turn without sending any message. Use when the thread is owned by a human operator after job.escalate, when the guest is self-resolving, when the message is a duplicate, or for observation-only turns. Calling this tool is the ONLY correct way to stay silent — narrated silence text (e.g. '(Staying silent…)', 'Internal:…') would be delivered to the guest verbatim.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | Free-form explanation for admin audit. Stored in trace_tool_executions.tool_params (ClickHouse String; reason filters are scan-only). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description goes beyond them by disclosing the critical behavioral trap: narrated silence text ('*(Staying silent…)*', 'Internal:…') is delivered to the guest verbatim, so this tool is the only safe way to stay silent. It also notes the reason field is retained for admin audit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, tightly front-loaded: the action first, then use-cases, then the failure mode being warned against. Every sentence carries distinct, non-redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-op control tool with full schema coverage and a rich annotations block, the description supplies everything needed to call it correctly: what it does, when to invoke it, and why substituting text is harmful. No output schema is needed for a tool that by definition sends nothing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already explains both parameters in detail (reason stored in trace_tool_executions.tool_params; in_workspace semantics). The description adds no parameter-level detail beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and effect ('End this turn without sending any message'), which is unambiguous and clearly distinct from siblings like messages_send or job_escalate. An agent can immediately tell this is the silence/no-op path, not an outbound-message tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It enumerates explicit when-to-use conditions (thread owned by a human after job.escalate, guest self-resolving, duplicate message, observation-only turns) and names the wrong alternative (narrated silence text). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_listARead-onlyIdempotentInspect
List all AI agents configured in the workspace.
Returns agents with their basic info, trigger count, and knowledge collection count.
Each agent's description field tells you when that agent is useful. If you're a router-style agent deciding whether to delegate via agent.handoff, read descriptions and pick the best fit.
Use this to:
See all configured AI agents
Filter by status (active/paused/archived)
Get agent IDs for further operations
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Filter by status ('active' / 'paused' / 'archived'). Omit for all. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, closed-world, and non-destructive, so the safety bar is satisfied. The description adds genuine behavioral value by disclosing the return shape (basic info, trigger count, knowledge collection count) and the semantic meaning of the per-agent description field, which the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and return summary, then uses a short bullet list for the enumerated uses. Slightly padded by the router/handoff paragraph, but every part is relevant and no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-shape burden and does so (basic info plus trigger and knowledge counts). It is complete enough to call correctly; only minor gaps remain, such as pagination or result size behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters—including the status enum and the workspace override—are already fully documented in the schema. The description restates the status filter ('active/paused/archived') but adds no syntax or semantics beyond it, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List all AI agents configured in the workspace'), immediately distinguishing it from single-agent siblings like agents_get. An agent can tell what this returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear when-to-use context: see all agents, filter by status, and retrieve agent IDs for further operations. It also frames the router/delegation scenario, but names no explicit exclusion or alternative sibling (e.g., 'use agents_get for a single agent'), so it stops short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_list_draftsARead-onlyIdempotentInspect
List pending agent drafts awaiting approval.
Shows drafts that have been generated by AI agents but not yet sent. Each draft includes:
Thread/conversation info
Trigger message (what prompted the reply)
Generated response text
Creation time and expiration
Use this when user asks:
'Show pending agent drafts'
'What messages are waiting for approval?'
'List drafts to approve'
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of drafts to return | |
| thread_id | No | Filter by specific thread ID (optional) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that: it discloses the contents of each draft returned (thread info, trigger message, response text, creation/expiration).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then a compact bulleted field list, then usage triggers. The bullet list is slightly longer than needed but each line earns its place by describing returned content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates well by enumerating the returned fields and covering the read-only safety case via annotations. It is essentially complete for a zero-required-param list tool, with only minor gaps around pagination/result ordering.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the read-only parameters (limit, thread_id, in_workspace) are already fully documented in the schema. The description adds no syntax or formatting detail for them, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List pending agent drafts awaiting approval') and adds a clarifying scope ('generated by AI agents but not yet sent'). This clearly separates it from sibling list tools like agents_list by specifying the pending-draft domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit trigger phrasing ('Show pending agent drafts', 'What messages are waiting for approval?', 'List drafts to approve') that maps user intent to this tool. It does not name alternatives or exclusions (e.g., that approval itself belongs to agents_approve_draft), so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_list_filesARead-onlyIdempotentInspect
List files directly attached to this agent (agent-specific files, not shared collections).
Returns file_id, title, status, and chunk_count for each file. chunk_count shows how many indexed chunks were created — 0 means the file is still processing.
Use agents.add_file to attach a new file, or agents.remove_file to detach one.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | ID of the agent whose files to list | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds real value beyond that: it enumerates the returned fields and explains that chunk_count=0 means the file is still processing, which is genuinely useful operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core purpose, followed by return semantics and a routing hint. No filler; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description compensates by listing the returned fields and their meaning. Combined with full schema coverage and annotations covering the safety profile, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both agent_id and in_workspace are fully documented in the schema. The description adds no syntax, format, or constraint detail beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (files attached to this agent) with an explicit scope qualifier distinguishing it from shared collections. An agent can tell it apart from collections_list_files without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the complementary tools agents.add_file and agents.remove_file for attaching/detaching, giving clear context. It does not, however, explicitly state when to prefer this over collections_list_files or search_files, so it falls short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_list_integrationsARead-onlyIdempotentInspect
List the workspace integrations enabled (or available) for an AI agent, with each one's workspace_integration_id, provider, enabled flag, and count of denied tools. Use this to see what an agent can call before toggling with agents_set_integration.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | ID of the agent to inspect | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, non-open-world, so safety needs no restating. The description adds the return shape (which fields come back per integration) plus the "enabled (or available)" nuance, which is useful since there is no output schema; it stops short of noting pagination or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the resource and its returned fields front-loaded and the routing hint second. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully compensates by naming the fields returned, and the read-only nature is fully covered by annotations. It leaves the in_workspace parameter entirely to the schema and says nothing about result limits, so it is strong but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both agent_id and in_workspace are already documented, and in_workspace's non-persistent behavior is spelled out in the schema itself. The description adds no parameter syntax or format detail beyond the implicit scoping to a single agent, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ("List the workspace integrations enabled (or available) for an AI agent") and enumerates exactly what each entry contains (workspace_integration_id, provider, enabled flag, denied-tool count). It also names the sibling it complements (agents_set_integration), so an agent can distinguish it without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the use case explicitly ("Use this to see what an agent can call") and points at the alternative action ("before toggling with agents_set_integration"), making the read-then-write pairing unambiguous. Nothing about when to reach for this vs. the setter is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_prompt_historyARead-onlyIdempotentInspect
List past versions of an agent's prompt_text. Every edit to the agent's prompt is snapshotted to an append-only table — use this tool to browse history, find a prior known-good version, and copy it into agents.prompt_restore.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max versions to return (1-200, default 50) | |
| agent_id | Yes | ID of the agent | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| before_version | No | Cursor: return versions strictly below this version_number |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real behavioral context beyond them: every prompt edit is snapshotted into an append-only table, implying complete and immutable history, which is useful for judging result reliability.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first establishes what is listed, the second establishes why and what to do next. No filler, and the purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only history listing with no output schema, the description is sufficient: it tells the agent what the records are, why the list is complete, and how the results are meant to be consumed (via agents.prompt_restore). Annotations cover the safety profile, and pagination parameters are fully described in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (agent_id, limit, before_version cursor, in_workspace) are already documented in the schema. The description adds no extra syntax or semantics for them, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List past versions of an agent's prompt_text') and adds the domain detail that these are append-only snapshots of every edit. It also distinguishes itself from the related restore path by naming `agents.prompt_restore` as the destination for a chosen version.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the use case: browse history to find a prior known-good version and then copy it into `agents.prompt_restore`, which effectively routes the agent between the two tools. It stops short of stating when *not* to use it (e.g., to fetch current prompt text, use agents_get), so it is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_prompt_restoreAInspect
Restore a past version of an agent's prompt_text by version_number. Creates a new version pointing at the restored content — history is preserved. Use agents.prompt_history first to find the version_number you want.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | Optional: why this restore is happening (shows up in history UI) | |
| agent_id | Yes | ID of the agent | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| version_number | Yes | The version_number to restore (get it from agents.prompt_history) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=false; the description adds the important non-obvious behavior that the restore appends a NEW version rather than overwriting, so history is preserved. This explains why a mutation is non-destructive, which the bare annotations cannot convey. It does not cover idempotency nuances (e.g. that repeating the call creates another version) despite idempotentHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and the append-only consequence, then the prerequisite. No filler and no repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and full annotation coverage, the description supplies the remaining essentials: what is being restored, the versioning semantics, and the prerequisite lookup tool. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents each parameter including the note that version_number comes from agents.prompt_history. The description largely restates that, so it meets the baseline without adding syntax or constraint detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (restore) plus the exact resource and field targeted (`prompt_text`), keyed by version_number. It is clearly distinguishable from siblings such as agents_prompt_history, agents_update, and prompts_prompt_restore.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly prescribes the prerequisite workflow: call `agents.prompt_history` first to obtain the version_number. That sequencing guidance is exactly what an agent needs to invoke this tool correctly rather than guessing a version number.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_reject_draftAInspect
Reject a pending agent draft without sending.
The draft will be marked as rejected and won't be sent. Use this when the generated response isn't appropriate.
Use this when user says:
'Reject this draft'
'Don't send this'
'Cancel this reply'
'Delete this draft'
'This response is wrong'
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | Optional reason for rejection (for logging/feedback) | |
| draft_id | Yes | ID of the draft to reject | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=false, readOnlyHint=false, and idempotentHint=false. The description adds useful context that the draft 'will be marked as rejected and won't be sent,' clarifying the state transition. However, it does not disclose whether rejection is reversible, whether a reason is logged (though schema hints at this), or any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action in the first sentence, followed by behavioral note and usage examples. The list of user phrasings is somewhat verbose but serves as helpful trigger recognition; overall efficient with no wasted sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the essential action, effect, and when-to-use examples. Minor gaps exist: no mention of irreversibility or alternative tools, but for a simple state-change tool with full schema coverage and annotations, this is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three parameters are already documented in the schema with descriptions. The description adds no new semantic details about parameters beyond what the schema provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (reject) and resource (pending agent draft) with the key scope constraint 'without sending.' An agent can immediately distinguish this from siblings like agents_approve_draft.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear when-to-use guidance with example user utterances ('Reject this draft', 'Don't send this'). However, it does not explicitly name the alternative (agents_approve_draft) or state when-not-to-use conditions such as rejecting an already-sent message.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_remove_fileAInspect
Remove a file from this agent's private knowledge.
The file itself is not deleted — it's just detached from this agent. Use agents.list_files to find the file_id to remove.
| Name | Required | Description | Default |
|---|---|---|---|
| file_id | Yes | ID of the file to detach (from agents.list_files) | |
| agent_id | Yes | ID of the agent to remove the file from | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=false and readOnlyHint=false, which on its own is ambiguous for a removal operation. The description resolves that ambiguity by explaining the file is not deleted, only detached — genuinely additive context. It doesn't address idempotency (annotations say idempotentHint=false) or what happens when the file isn't attached, leaving a small gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and the single most important caveat (non-destructive detach), then the prerequisite. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple detach operation with full schema coverage, annotations, and no output schema, the description covers what is needed to call it correctly. Only minor behavior (error on a non-attached file, idempotency) is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented in the schema, including file_id's origin. The description reinforces the file_id source but adds no semantics beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Remove a file from this agent's private knowledge') and immediately disambiguates scope by clarifying it is a detach rather than a delete. This distinguishes it from files_delete and agents.delete, which an agent could otherwise confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It points to agents.list_files as the way to obtain the file_id, which is a useful workflow hint, but never states when to prefer this over siblings like files_delete, collections_remove_file, or agents.add_file's inverse. Usage is implied rather than contrasted with alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_set_integrationAInspect
Enable or disable a connected workspace integration on an AI agent — this controls which ext_ integration tools the agent (and its sandboxed workbench runs) may call. Use integrations_list to get the workspace_integration_id.
When enabling with no explicit denied_tools, WRITE-class tools are auto-disabled by default (read tools stay on); pass denied_tools=[] to force-allow everything, or a list of ext slugs / bare tool names to block specific ones. Idempotent upsert — safe to call repeatedly.
| Name | Required | Description | Default |
|---|---|---|---|
| enabled | No | True to enable the integration on this agent, False to disable | |
| agent_id | Yes | ID of the agent to configure | |
| denied_tools | No | Optional explicit block-list of tool names to deny (ext slugs or bare names). Omit to auto-deny WRITE-class tools on first enable; pass [] to allow all. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| workspace_integration_id | Yes | ID of the connected integration (from integrations_list) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description claims "Idempotent upsert — safe to call repeatedly," but the annotations declare idempotentHint=false, a direct conflict on the tool's most safety-relevant property. The claim is also internally strained by the schema's own "on first enable" caveat, so an agent relying on it could misjudge whether re-calling changes write-tool exposure. This is an annotation contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short paragraphs, front-loaded with the purpose and effect, then the default-block behavior and the idempotency note. The only mildly redundant sentence is the integrations_list pointer, which also duplicates the schema, but nothing else is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
It covers purpose, the sibling needed to obtain the required ID, the default WRITE-tool blocking, and the ephemerality of in_workspace (also in schema), which is good coverage for a config-mutating tool with no output schema. However, it omits any permission/authorization requirements for changing agent configuration and its idempotency claim is misleading, leaving an agent with a false model of repeat-call behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter including denied_tools and in_workspace is already documented in the schema. The description reinforces the non-obvious denied_tools default (auto-deny WRITE-class tools on enable) but adds no new syntax, format, or interaction detail beyond what the schema states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ("Enable or disable a connected workspace integration on an AI agent") and immediately scopes the effect to which ext<id>_<name> tools the agent and its sandboxed workbench runs may call. This clearly separates it from siblings like agents_list_integrations (which reads) and integrations_list (which resolves IDs), so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It routes the agent to a specific sibling for prerequisite data ("Use integrations_list to get the workspace_integration_id") and spells out the when-to-use branches for denied_tools (omit for auto-deny of WRITE tools, [] to force-allow, list to block specific ones). What is missing is any explicit when-not-to-use guidance or a note about which sibling to call instead when the goal is merely inspecting the current integration state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_simulate_inboundARead-onlyIdempotentInspect
Replay an inbound message on a thread through the real trigger pipeline and return what would have happened. The router auto-picks the winning enabled agent + trigger by priority/specificity (same logic as production). By default send_mode='draft' so no real message is sent; pass send_mode='auto' on a test account to let the matched agent actually deliver (drafts get overwritten by the next draft, so 'auto' is the only way to verify Telegram/email delivery end-to-end).
Use to verify routing for a thread: which agent answers, which trigger wins, or — when nothing matches — the structured skip reason. Pass blockchain_tx_data instead of message_text to simulate a blockchain:transfer event on the thread.
Returns: {matched: true, matched_agent: {id, name, execution_mode}, matched_trigger: {id, trigger_type, conditions, specificity_score}, routing_reason, response_text, messages[], execution_mode, send_mode, model_used, tokens_input, tokens_output, latency_ms, } on a hit, or {matched: false, skip_reason, simulator_warnings} on a miss.
| Name | Required | Description | Default |
|---|---|---|---|
| send_mode | No | How the matched agent should deliver its reply. 'draft' (default, safe) creates a draft only — no real send, no idempotency key. 'auto' lets the agent deliver through the channel adapter exactly as it would in production — use this on a test account to verify Telegram/email delivery end-to-end. Drafts get overwritten by the next draft on the thread, so 'auto' is required when you want to see the message persisted. | draft |
| thread_id | No | Thread ID to route the simulated event from. Must belong to the API key's workspace. Omit and set create_new_livechat=true to test on a FRESH thread with no history. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| message_text | No | Inbound message body to simulate. Defaults to '[MCP simulation test]' when omitted. | |
| system_message | No | Tag the simulated inbound as a system/service-message row (missed call, group join, pinned message, etc.) so the `excluded_system_message_kinds` trigger filter can be exercised end-to-end. Shape: {"category": <one of call_event | membership_change | contact_signup | pinned_message | chat_metadata_change | voice_chat_event | other_service>, "native_kind": <free-form upstream event class name, e.g. 'MessageActionPhoneCall'>}. The category is written into `message.meta.system_message` (mirroring the real Telegram ingest path) AND surfaced on the synthetic IncomingEvent so the trigger evaluator honors the block-list. Omit for a normal text-message simulation. | |
| blockchain_tx_data | No | When set, simulate a blockchain:transfer event instead of a channel:message:new event. Expected keys: chain, to_address / from_address, tx_hash. | |
| channel_account_id | No | Optional. The livechat widget's channel_account_id to host the fresh chat when create_new_livechat=true. Omit to auto-pick the workspace's active livechat widget. | |
| attachment_file_ids | No | Optional list of workspace file IDs to attach to the simulated inbound message — same shape as a real Telegram message with image/document attachments. Use this to test agent behavior on incoming messages that carry images (e.g. logos for invoices) or documents the agent must reference. File IDs must belong to the API key's workspace. | |
| create_new_livechat | No | Start a FRESH, history-free livechat chat instead of using an existing thread_id — creates a new visitor + thread on the workspace's livechat widget and routes your message onto it. The response's thread_id is the new thread; pass it back (with send_mode='auto') to continue the conversation. Ideal for clean multi-turn text tests. OMIT to reuse an existing thread_id. Ignored if thread_id is given. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description adds real value on top: the draft-vs-auto distinction, the warning that drafts get overwritten, and that 'auto' is the only way to verify end-to-end delivery. It does not explicitly reconcile readOnlyHint=true with the fact that 'auto' mode actually sends a real message, leaving a minor tension an agent must infer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and routing mechanics before the return shape. The inline enumeration of the full response object is long but earns its place given there is no output schema; slightly verbose for a prose description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the routing model, the safe/unsafe delivery modes, the alternate blockchain event path, and both hit and miss response shapes — filling the gap left by the missing output schema. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all nine parameters in depth; the description only re-touches send_mode and blockchain_tx_data. Baseline 3 is appropriate since the schema, not the description, carries parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('replay an inbound message on a thread through the real trigger pipeline') plus the outcome ('return what would have happened'). An agent can immediately tell this apart from observers like agents_trace_get or agents_activity, which report on past runs rather than simulating a new routing decision.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says what to use it for (verify which agent/trigger wins, or capture the structured skip reason), gives the safe default (send_mode='draft'), names the exact condition for the risky variant (send_mode='auto' on a test account to verify delivery), and points to blockchain_tx_data for a different event type. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_task_completeAInspect
Report that an agent run handed to you by DialogBrain is over: a Claude Code agent task, or a run started in an external agent. Call this when you finish processing it.
| Name | Required | Description | Default |
|---|---|---|---|
| success | Yes | Whether the task completed successfully | |
| summary | No | Brief summary of what was done | |
| trace_id | Yes | Trace ID from the agent task event | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only write (readOnlyHint=false), non-idempotent (idempotentHint=false), and non-destructive, so the safety profile is covered. The description adds the scoping detail that the run was 'handed to you by DialogBrain,' but does not explain side effects of reporting, whether it can be called twice safely despite idempotentHint=false, or what happens afterward.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences that front-load the core purpose before the call trigger. Nothing is redundant, though the second sentence's 'Call this when you finish processing it' is slightly generic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter reporting tool with annotations covering the safety profile and no output schema needed, the description covers what the tool does and when to invoke it. It could be stronger on post-call behavior and non-idempotency implications, but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (success, summary, trace_id, in_workspace) are already documented in the schema. The description adds no parameter-level meaning beyond that, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Report') and the resource/event being reported ('an agent run ... is over'), enumerating the two origin flavors (Claude Code agent task, external agent run). It is clear enough to distinguish from most siblings, though it doesn't explicitly contrast with tools like agent_handoff or agents_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Call this when you finish processing it' gives an explicit trigger condition tied to task completion. There is no when-not guidance or named alternative (e.g., what to use if the task failed or was abandoned), but the activation context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_trace_getARead-onlyIdempotentInspect
Fetch the full execution detail for a single trace — tool executions, events timeline, LLM call spans (with error_message on failures), and what the run cost.
cost_usd is the run's billed cost in USD, recorded even when the workspace pays with its own vendor key; 0 means nothing billable was recorded, null means the lookup could not run. Per-token detail is on each LLM span (input_tokens, output_tokens, cache_read_tokens, cache_creation_tokens).
Use after agents.traces_list identifies a specific trace of interest (failed run, slow run, unexpected outcome).
By default LLM system_prompt and prompt_messages are stripped — set include_llm_bodies=true to fetch them when diagnosing prompt engineering issues (emits a WARNING audit log). Set full=true to disable all field truncation. completion_text on failed LLM calls is always returned (capped at 8 KB).
| Name | Required | Description | Default |
|---|---|---|---|
| full | No | Disable all field truncation. Escape hatch for a human operator. OMIT for the standard truncated view. | |
| agent_id | Yes | Expected agent_id — used for scope validation. Mismatch returns not_found. | |
| trace_id | Yes | Trace identifier returned by agents.traces_list. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| include_llm_bodies | No | Include system_prompt and prompt_messages in LLM spans. Audited at WARNING level. OMIT to keep them stripped (the default). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. Beyond that, the description discloses non-obvious behavior: LLM system_prompt and prompt_messages are stripped by default, include_llm_bodies=true triggers a WARNING-level audit log, full=true disables truncation, and completion_text is capped at 8 KB. This is substantial disclosure the annotations do not carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and payload are front-loaded in the first sentence, followed by cost semantics, usage routing, then option defaults. Dense and nearly waste-free, though three paragraphs is longer than strictly needed for a single-trace getter and some field detail (per-token names) could have been left to the payload itself.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the return-value burden and does so: it names the span fields, explains cost_usd's 0-vs-null distinction, and covers token-level fields. For a 5-param read tool with full schema coverage, an agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents each parameter; the baseline would be 3. The description adds marginal-but-real semantics: include_llm_bodies is for 'diagnosing prompt engineering issues' and its audit consequence, and full is framed as a 'human operator' escape hatch rather than a routine option.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Fetch the full execution detail for a single trace') and enumerates the payload (tool executions, events timeline, LLM spans with error_message, cost). It is immediately distinguishable from agents_traces_list (enumeration) and agents_traces_stats (aggregates), which appear as siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says 'Use after agents.traces_list identifies a specific trace of interest (failed run, slow run, unexpected outcome)', naming the upstream sibling and the triggering conditions. The agent knows both the prerequisite call and the scenarios that select this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_traces_listARead-onlyIdempotentInspect
List recent execution traces for an agent — the same data as /admin/requests, scoped to one agent and readable by an LLM.
Use this when an agent call timed out, drafted the wrong response, or you want to know which tool/LLM call burned the latency. Pair with agents.trace_get for full detail on a specific trace.
Filters: status, success, source (single value or comma-separated: agent,voice), date_from/date_to (ISO-8601), pagination via limit/offset.
Returns returned_count, dropped_on_page (should be 0 — positive means the backend agent_id predicate let something through), and has_more. Edge case: a raw page of all-dedup-dropped rows yields returned_count=0, has_more=true; re-call with offset += limit.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max rows per page (1–100). | |
| offset | No | Rows to skip for pagination. OMIT to start at row 0 (default). | |
| source | No | Filter by trace source. Single value or comma-separated, e.g. 'agent,voice'. Values: agent / auto_reply / agentic / outreach / voice. Note: source='agent' also matches voice traces today (known upstream bug). | |
| status | No | Filter by status. OMIT to include all statuses. | |
| date_to | No | ISO-8601 upper bound on created_at. | |
| success | No | Filter to succeeded (true) or failed (false) runs only. OMIT to include both. | |
| agent_id | Yes | Agent ID to pull traces for (must belong to your workspace). | |
| date_from | No | ISO-8601 lower bound on created_at, e.g. '2026-04-10T00:00:00Z'. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), yet the description still adds real behavioral context: the response fields (returned_count, dropped_on_page, has_more), what a non-zero dropped_on_page implies, and the tricky edge case where a fully deduped page returns returned_count=0 with has_more=true and requires re-calling with offset += limit. That is exactly the kind of operational nuance annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then usage, then filters, then return-shape caveats — a logical progression with no filler sentences. It is on the longer side, but nearly every clause carries actionable information, so the size is defensible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes on the burden of explaining the return shape and does so (returned_count, dropped_on_page, has_more) plus the pagination edge case and recovery step. Combined with the filter inventory and the pointer to agents.trace_get, an agent has everything needed to call and interpret this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema, and the description's filter paragraph largely restates it (status, success, source, date range, limit/offset). It adds modest consolidation value by grouping the filters in one place, but no new semantics such as default behavior or interaction rules. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List recent execution traces for an agent') and immediately scopes it ('scoped to one agent'), plus maps it to a known surface (/admin/requests). It also names the sibling that provides the complementary capability (agents.trace_get), so an agent can distinguish it from agents_traces_stats and agents_activity without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggering scenarios — an agent call timed out, a wrong draft, or latency investigation — and explicitly routes to agents.trace_get for per-trace detail. That is a clear when-to-use plus a named alternative, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_traces_statsARead-onlyIdempotentInspect
Aggregated trace statistics for one agent over the last N days — total runs, success rate, avg duration, error breakdown, top tools used, runs-per-day histogram, and what the agent spent.
spend totals the agent's LLM usage over the same window: llm_calls, tokens_input, tokens_output, cost_usd, and own_key_cost_usd (the slice paid with the workspace's own vendor key, included in cost_usd rather than added to it). It covers usage recorded since per-agent attribution shipped, so it reads 0 for older runs; null means the lookup could not run.
Use this when you want a bird's-eye view of an agent's health before diving into individual traces with agents.traces_list / agents.trace_get. Scoped to the target agent (exact match, no substring bleed). days is capped at 30 — matches the ClickHouse request_traces TTL.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Rolling window in days (1–30). | |
| agent_id | Yes | Agent ID to compute stats for (must belong to your workspace). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the readOnly/idempotent annotations by disclosing data caveats: spend reads 0 for runs predating per-agent attribution, null means the lookup could not run, the window is capped at 30 days to match the ClickHouse request_traces TTL, and matching is exact with 'no substring bleed'. These are non-obvious behaviors an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and output inventory, then the spend semantics, then usage guidance. Dense and mostly waste-free, though the spend paragraph is on the longer side for a single returned field.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing returns and does so by enumerating the aggregate metrics and spend fields, including the subtle own_key_cost_usd relationship. An agent has everything needed to call it and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: the rationale for the `days` cap (ClickHouse TTL) and the exact-match scoping of `agent_id` ('no substring bleed'). It stops short of explaining `in_workspace`, but the schema already covers it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Aggregated trace statistics for one agent') and enumerates the exact fields returned (runs, success rate, avg duration, error breakdown, top tools, histogram, spend). This clearly distinguishes it from the sibling trace tools that return individual records.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('bird's-eye view of an agent's health') and names the alternatives it complements ('before diving into individual traces with agents.traces_list / agents.trace_get'), with the routing condition stated rather than implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_trigger_createBInspect
Create a new trigger for an AI agent.
Triggers determine when the agent activates.
Trigger types:
incoming_message: Activates on new incoming messages
schedule: Activates on a schedule
webhook: Activates on webhook events
event: Activates on system events
| Name | Required | Description | Default |
|---|---|---|---|
| enabled | No | Whether the trigger is enabled. OMIT to use the default (true). | |
| agent_id | Yes | ID of the agent to create a trigger for | |
| priority | No | Trigger priority — lower numbers run first (default: 100) | |
| send_mode | No | Send mode override for this trigger. OMIT to inherit from the agent. | |
| conditions | No | Trigger conditions (JSON). Supported fields for incoming_message: - keywords: ["pricing","demo"] — message must contain keyword(s) (free, no LLM cost) - keyword_match: "any" (default, OR) or "all" (AND) - channel_types: ["telegram","whatsapp","livechat_voice","twilio_voice","telegram_voice","voice",...] — filter by channel. For voice, use EITHER the three per-channel keys (scoped) OR "voice" alone (wildcard matching all three) — mixing them is redundant. Per-channel keys: "livechat_voice" (web widget), "twilio_voice" (PSTN inbound), "telegram_voice" (Telegram p2p calls) - context_types: ["dm","group","channel","livechat"] — filter by chat type - group_mode: "mentions_only" or "questions" — for group chats - channel_account_ids: ["123"] — restrict to specific accounts - folder_ids: [5,10] — restrict to threads in folders - ai_tag_ids: [1,2] — restrict to threads with AI tags - ai_filter_ids: [1,2] — semantic intent filters (message matched via embedding similarity, works in noisy groups) - ai_filter_mode: "any" (default, OR) or "all" (AND) — how multiple AI filters combine - ai_filters: [{id: 1}, {name: "...", description: "..."}] — shorthand: reference existing by id or create inline (calls Voyage embedding API). If a filter with the same name already exists, it is reused by id. Prefer referencing existing filters by id when available. Use ai_filters.create + ai_filters.test for fine-tuning before assigning. - template: the reply a rule_based agent sends, with no model run. Either a string ("Hi {from_name}, we open at 9") or a card object {text, attachments, buttons} — text is required, attachments is a list of file ids (files.upload), buttons is a list of {label, value|url|web_app} carrying exactly ONE target each: value = a reply the bot receives back, url = opens a link, web_app = opens a Mini App over https. The string, and a card's text, substitute {from_name}, {message_text}, {channel_type}, {thread_id}, {raw_data}; button labels and targets are sent verbatim. A card with buttons needs send_mode auto — a draft row cannot carry a keyboard. Buttons need a channel that renders them (telegram_bot). A card needs a chat to land in, so it refuses on webhook, schedule and other thread-less triggers where a string still drafts. - handoff: {agent_id, task} — only on a rule_based agent's trigger (refused elsewhere). After the rule's deterministic part has actually happened — the template's message delivered, or the agent's script completed — the named agent runs on the same thread, in its own send mode, exactly as its own trigger would run it: the customer's message reaches it as the incoming message, and `task` reaches it as the handoff brief from this rule. The rule stays instant; the thinking — a reminder, a referral code, a personal follow-up — happens behind it. Script plus handoff is how a workspace does something deterministic the platform has no verb for (file the thread, stamp a tag, push a row) and still lets an agent answer, with the deterministic part living in that agent's own script rather than in platform code. `task` substitutes {channel_type} and {thread_id} only; {message_text}, {from_name} and {raw_data} are refused (the sender would be writing the instruction — the message and its sender already reach the target as the incoming message). The target must be an active agentic or claude_channels agent in this workspace (the engines that read the brief). The target's run is a normal trigger run: it holds the thread lock, its own safety gates apply, and its final answer is delivered like any DM reply — a follow-up that should stay silent must call agent.silence. Nothing runs when the message was drafted, refused or failed, when the script failed, when a safety gate has the rule in draft mode, or when the event carries no thread for the target to answer on. A target paused or deleted after the trigger was written does not stop the rule's message: it goes out, and the reason the follow-up did not start is recorded in the rule's activity row (error field) and in the target's activity as an execution_error. - contact_states: ["active"] — filter by contact state - cooldown_seconds: 30 — min gap between runs per thread - max_runs_per_thread_per_hour: 5 — rate limit - once_per_thread: true — after this trigger completes on a conversation it steps aside there and the next matching trigger answers later messages. For a rule that files / tags / hands off once, not for the agent that holds the conversation. - answer_delay_s: 15 — voice pickup only (incoming_call, or incoming_message scoped to a voice channel). Ring the human's own devices this long before the agent answers; if they pick up, the agent stands down. 0/absent = answer immediately. Honoured on WhatsApp and Telegram 1:1 — other voice channels have nothing ringing to wait for and still answer at once. - auto_join: true|false — voice pickup, GROUP calls only. Whether the bot enters a matching group call on its own (true) or the call becomes a pending invite an operator accepts (false). OMIT THE KEY to defer to the channel setting (channel_account.state.voice_auto_join_policy, default 'approval'); resolution is trigger-when-present, then channel, then approval, so an explicit false beats a channel set to auto. Absent and false are different answers. - chat_ids: ["-5172634473"] — voice pickup, Telegram GROUP calls only. Scopes the trigger to specific groups; omit for every group call. Any id spelling works (5172634473, -5172634473, -1005172634473 are the same group). A call whose chat is unknown never satisfies it, so do NOT set this together with 1:1 / Twilio / Telnyx / WhatsApp / Android / LiveChat voice channels or the "voice" wildcard: those calls have no chat and the trigger would stop answering them. Supported fields for job_completed (proactive callback when a delegated job finishes): - source_agent_id: <int> — fire only when this agent's job completed - source_agent_slug: <str> — alternate to source_agent_id - job_type: "agentic_session" — match a specific job type (default: any) - outcome: ["completed"] | ["escalated"] | ["completed","escalated"] — default ["completed"] - min_duration_seconds: <int> — skip very-short jobs (noise filter) - thread_filter: {thread_ids: [<int>...]} — restrict to specific threads Supported fields for calendar_event (fires N minutes before a Google Calendar event starts): - window_minutes_before: <int 1-1440> — REQUIRED, fire when an event starts within this window - channel_account_ids: [<int>...] — restrict to specific calendar accounts (default: all) - keywords: ["standup"] — word-boundary match on event title - prepare_meet_join: true — pre-invite pool bots to the event (enables unattended Meet join) incoming_message action fields: - action: "reply_text" (default, normal agent run) or "join_voice" (deterministically join the voice channel resolved from the message — requires send_mode=auto) - message_source: "real" (default) | "transcript" | "both" — real messages vs turns during a live call. Transcripts are turns that address the agent. calendar_event run-mode field (incoming_message uses `action` instead): - run_mode: "text" (default) or "voice" (join the meeting — requires send_mode=auto) - voice: {speak_first: <bool — greet immediately vs stay silent until addressed>, vision_mode: "off"|"on_demand"|"continuous_0_3fps"} — pairs with action=join_voice (incoming_message) or run_mode=voice (calendar_event) | |
| thread_ids | No | Restrict this trigger to specific threads (chats) by their numeric thread IDs. When set, the trigger only fires for messages in these threads. Only for incoming_message and job_completed triggers. Maps to conditions.thread_filter.thread_ids. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| trigger_type | Yes | Type of trigger: 'incoming_message', 'incoming_call', 'schedule', 'webhook', 'event', 'blockchain_event', 'job_completed', 'calendar_event', or 'lead_captured' | |
| target_session | No | Which Claude Code desktop session this trigger's runs go to, by the name / client id / session id that workspace.desktops lists. Only meaningful on a claude_channels agent: that engine hands the run to a connected desktop, and without an address it can only be delivered when EXACTLY ONE is connected (a second open window turns every run into channel_target_missing). A task assigned to the agent still wins — its own target_session is per-request and more specific than the trigger's default. Stored as conditions.target_session. | |
| blocked_sender_ids | No | Never react to these senders/callers, by their channel-side id — a phone number, a WhatsApp JID, a Telegram user id, an email address. Explicit deny: it beats every allow-list, and it applies to ALL trigger types, including incoming_call, so it is how you stop an agent answering one nuisance caller without silencing it for everybody. Phone-shaped ids are matched by their digits, so '+998901234567', '998901234567' and '998901234567@s.whatsapp.net' are the same person. Maps to conditions.sender_filter.excluded_external_ids. | |
| excluded_thread_ids | No | Exclude specific threads (chats) by their numeric thread IDs — the opposite of thread_ids. When set, the trigger NEVER fires for messages in these threads, even if thread_ids would otherwise allow them (explicit deny wins). Only for incoming_message and job_completed triggers. Maps to conditions.thread_filter.excluded_thread_ids. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnly=false, destructive=false, idempotent=false, openWorld=false), so the description's remaining burden is light. It contributes the trigger-type activation semantics, but says nothing about persistence, default enabled state, draft/send_mode interactions, or what happens on repeat calls — all of which matter for a mutating tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Six short lines, purpose front-loaded, zero filler prose. Minor waste: the trigger-type bullet list duplicates (and under-represents) the schema enum, so it earns less than a full 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter mutating tool with a nested conditions object, the description is thin, but the schema carries essentially all parameter detail, so the gap is mostly the unstated return value (no output schema exists) and the incomplete type enumeration (4 of 9 enum values, missing incoming_call, job_completed, calendar_event, blockchain_event, lead_captured).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema fully documents all 11 parameters (including the very detailed `conditions` object); baseline 3 applies. The description adds no parameter meaning of its own, and its four-item type list is narrower than the nine-value enum, so it does not compensate further.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (create) and resource (trigger) plus the owning entity (AI agent), and adds that triggers determine when an agent activates. It is clearly distinguishable from agents_create and agents_trigger_update, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The type list implies when each trigger fires, which gives some usage context, but there is no explicit when-to-use guidance, no mention of prerequisites (the agent must already exist), and no routing to agents_trigger_update / agents_trigger_delete for modifying or removing triggers.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_trigger_deleteADestructiveIdempotentInspect
Delete a trigger from an AI agent.
WARNING: This cannot be undone.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | ID of the agent that owns this trigger | |
| trigger_id | Yes | ID of the trigger to delete | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true and idempotentHint=true, so the agent knows this is destructive and idempotent. The description adds the irreversible nature ('cannot be undone'), which is useful but partly redundant with destructiveHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two very short sentences, zero waste, and the warning is front-loaded after the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple delete tool, the description covers purpose and irreversibility, which is adequate. However, it doesn't mention required permissions or side effects beyond deletion, and there's no output schema to rely on.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents parameters. The description adds no parameter syntax or constraints, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Delete a trigger from an AI agent.' It clearly distinguishes this tool from siblings like agents_trigger_create and agents_trigger_update, which perform different operations on the same resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (to delete a trigger) but doesn't provide explicit when-to-use or when-not-to-use guidance, nor does it reference alternatives such as agents_trigger_update for modification.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_trigger_updateBInspect
Update an existing AI agent trigger.
All parameters are optional — only provided fields will be updated.
| Name | Required | Description | Default |
|---|---|---|---|
| enabled | No | Enable or disable this trigger. OMIT to leave the enabled flag unchanged. | |
| agent_id | Yes | ID of the agent that owns this trigger | |
| priority | No | Trigger priority — lower numbers run first | |
| send_mode | No | New send mode override. OMIT to leave the send-mode unchanged. | |
| conditions | No | New trigger conditions (replaces existing). Same fields as trigger_create: keywords, keyword_match, channel_types, context_types, group_mode, channel_account_ids, folder_ids, ai_tag_ids, ai_filter_ids, ai_filter_mode, ai_filters: [{id: 1}, {name: "...", description: "..."}] — shorthand: reference existing by id or create inline (calls Voyage embedding API). If a filter with the same name already exists, it is reused by id. contact_states, cooldown_seconds, max_runs_per_thread_per_hour, once_per_thread (true: completes once per conversation, then the next matching trigger answers there). template: the rule_based reply — a string, or a card object {text, attachments, buttons} whose buttons each carry exactly one of value|url|web_app; a card with buttons needs send_mode auto. handoff: {agent_id, task} — rule_based owner only. After the rule's message has gone out, that agent runs on the same thread as its own trigger would, with task as the handoff brief; task may use {channel_type} and {thread_id} only. The target must be an active agentic or claude_channels agent in this workspace. incoming_call (voice pickup): answer_delay_s, plus auto_join and chat_ids for GROUP calls — auto_join true|false decides whether the bot enters on its own or the call becomes a pending invite, and OMITTING the key defers to the channel setting (trigger-when-present, then channel, then approval; absent and false are different answers). chat_ids scopes to specific Telegram groups (any id spelling) and must not be combined with channels whose calls have no chat, which never satisfy it. Conditions are REPLACED, so dropping a key clears it. calendar_event: window_minutes_before (1-1440, required), channel_account_ids, keywords, prepare_meet_join. incoming_message: action "reply_text"|"join_voice" (join requires send_mode=auto), message_source "real"|"transcript"|"both". calendar_event: run_mode "text"|"voice" (voice requires send_mode=auto). Both: voice: {speak_first, vision_mode} | |
| thread_ids | No | Restrict this trigger to specific threads (chats) by their numeric thread IDs. When set, merged into conditions.thread_filter.thread_ids. If conditions is also provided, thread_ids is merged into it. Only for incoming_message and job_completed triggers. | |
| trigger_id | Yes | ID of the trigger to update | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| trigger_type | No | New trigger type. OMIT to keep the existing type unchanged. | |
| target_session | No | Which Claude Code desktop session this trigger's runs go to (name / client id / session id from workspace.desktops). Only meaningful on a claude_channels agent. Pass "" to clear it and go back to deliver-only-when-one-desktop-is-connected. Stored as conditions.target_session. | |
| blocked_sender_ids | No | Never react to these senders/callers, by their channel-side id — a phone number, a WhatsApp JID, a Telegram user id, an email address. Explicit deny: it beats every allow-list, and it applies to ALL trigger types, including incoming_call, so it is how you stop an agent answering one nuisance caller without silencing it for everybody. Phone-shaped ids are matched by their digits, so '+998901234567', '998901234567' and '998901234567@s.whatsapp.net' are the same person. Empty list [] = unblock everyone (the allow-list, if any, is left alone). Maps to conditions.sender_filter.excluded_external_ids. | |
| excluded_thread_ids | No | Exclude specific threads (chats) by their numeric thread IDs — the opposite of thread_ids. When set, the trigger NEVER fires for messages in these threads (explicit deny wins). Merged into conditions.thread_filter.excluded_thread_ids. Only for incoming_message and job_completed triggers. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, so the safety profile is covered. The description adds patch semantics (unlisted fields are preserved), which is genuine behavioral value. But it omits a key hazard disclosed only in the schema — that conditions are fully REPLACED, and that thread_ids/excluded_thread_ids merge into conditions — which matters given destructiveHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, well front-loaded: the action first, then the update semantics. Efficient, though the second sentence is partly inaccurate relative to the schema's required fields, so it does not fully earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter mutation tool with nested objects and no output schema, the description is minimal — the massive conditions schema does nearly all the work. It neither routes the agent away from trigger_create/delete nor warns about the destructive conditions/replacement behavior, leaving the definition only barely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the parameter descriptions are unusually rich (merging behavior, deny-list precedence, phone-number matching), so the schema carries the burden and baseline 3 applies. The description adds no parameter-level meaning beyond that, and its "all parameters are optional" statement is inaccurate against the required trigger_id/agent_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource ("Update an existing AI agent trigger") that plainly tells the agent what the tool does. However, it does no sibling differentiation — agents_trigger_create, agents_trigger_delete, and agents_update all sit nearby, and the description never distinguishes between them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"only provided fields will be updated" implies partial-update usage, but there is no explicit when-to-use vs agents_trigger_create or agents_trigger_delete, and no prerequisites or exclusions. Notably, the claim that "all parameters are optional" conflicts with the schema's two required fields (trigger_id, agent_id), which makes the usage hint misleading.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_updateAInspect
Update an existing AI agent's configuration.
All parameters are optional — only provided fields will be updated.
Use this to:
Enable or disable an agent
Change agent name or description
Assign or detach a prompt
Change default send mode
Replace knowledge collections
Update agent status
Change agent priority for trigger matching (lower number = higher priority)
Override which tools the agent can/can't call on triggered runs
Override which context sections (situation, communication style, job state, conversation history, thread summary) the agent receives
Opt into boilerplate prompt sections (safety guidelines, data confidentiality, factual accuracy) — all default OFF
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | New name for the agent | |
| model | No | Canonical source for which LLM the agent runs on. To switch models pass JUST this — do NOT also rewrite prompt_text (any 'duty model' section in the prompt is stale doc, not the config). OMIT to leave the model unchanged. | |
| script | No | rule_based deterministic action (no LLM): Python run in the workbench sandbox on each matched event. Reads `inputs` (raw_data, message_id, from_name, …) and calls the agent's integrations via call_tool('ext<id>_<name>', {..}). Dedupe writes on inputs['message_id'] (retries re-run). Pass null to clear (falls back to per-trigger template). | |
| status | No | Agent status: 'active', 'paused', or 'archived'. OMIT to leave the status unchanged. | |
| agent_id | Yes | ID of the agent to update | |
| priority | No | Agent priority for trigger matching. LOWER number = HIGHER priority (wins tiebreaks). Typical range 1-100. Fallback auto-reply agents use 10; specialised/topical agents use 100. When two agents match the same incoming message, the one with the lower priority number fires. | |
| prompt_id | No | Prompt ID to assign (null to detach) | |
| send_mode | No | Default send mode: 'auto' or 'draft'. OMIT to leave the send-mode unchanged. | |
| fast_model | No | Model for the fast-path responder (voice, text auto-reply, agent executor). Defaults to deepseek-v4-flash-nothink when unset. Non-Anthropic models (deepseek-v4-flash-nothink, gpt-4.1-nano, kimi-k2.6) do NOT use BYOK today — they use the system API key + credits. Pass null to revert to default. | |
| api_surface | No | OpenAI HTTPS endpoint for this agent's LLM calls (Phase 3a). 'chat_completions' (default, also when null) routes to /v1/chat/completions. 'responses' routes to /v1/responses — required for OpenAI native server tools (web_search, code_interpreter, image_generation, input_file PDFs). Capability still wins: agents whose tool list triggers the server_tool_responses_api substitution always route to Responses regardless of this setting. Ignored on non-OpenAI models (Anthropic, DeepSeek, Moonshot). OMIT to leave the api_surface unchanged. | |
| description | No | New description for the agent | |
| prompt_text | No | DESTRUCTIVE — REPLACES the entire system prompt. Pass ONLY when the user explicitly asks to edit/rewrite the prompt. To READ the prompt use prompts.get. When updating other fields (model, name, …) OMIT this. To append, prompts.get first then concatenate. Pass null to revert to the linked template. | |
| remote_tool | No | Only for text_engine='external_agent': the workspace integration tool that starts the remote agent, e.g. 'ext42_run_routine' (list them with integrations.search_tools). On each trigger DialogBrain calls it once with the event, the agent's instructions and the ids to answer with; the remote agent replies through the DialogBrain MCP tools (messages.send, tasks.comment, agents.task_complete). The endpoint and its secret belong to the integration, not to the agent. | |
| text_engine | No | Text-execution engine: 'agentic', 'ai_assisted', 'rule_based', or 'claude_channels'. Replaces the legacy execution_mode field (20260523_002). Voice is now derived from triggers, not engine. OMIT to leave unchanged. | |
| voice_tools | No | The EXACT set of tool IDs exposed to the LIVE VOICE runner (dotted IDs, e.g. ['knowledge.query','messages.send','messages.read_history','agent.handoff','calls.end','calendar.check_availability','calendar.create_event','contacts.capture_lead']). SEPARATE from allowed_tools (which governs TEXT mode): a tool only reaches the voice LLM if listed here. When set this REPLACES the whole voice surface — include EVERY tool the voice agent needs (the small default set is NOT auto-added once any voice tool is set). Tools listed here are also added to the text allow-list if absent. Empty list [] = no voice tools. OMIT to leave the voice tool surface unchanged. | |
| denied_tools | No | Block-list of tool IDs the agent must not call on triggered runs. Applied after allowed_tools and default visibility. Empty list [] = clear the block-list. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| voice_engine | No | Voice execution engine: 'pipeline' (default — Deepgram/Gladia STT + LLM + TTS), 'openai_realtime' (OpenAI Realtime API v2v; requires the workspace to have a BYOK OpenAI key connected — the worker falls back to 'pipeline' and logs why if the key is missing), or 'gemini_realtime' (Gemini Live native-audio v2v; runs on the workspace owner's Gemini key, else on the platform key unless platform keys are switched off). If the chosen realtime engine has no usable key, the result carries `warnings` and calls use the default pipeline. OMIT to leave unchanged. | |
| allowed_tools | No | Explicit allow-list of tool IDs this agent can call on triggered TEXT runs (e.g. ['messages.send', 'agent.handoff']). REPLACES the current text list: only these tools (minus denied_tools) are exposed to text runs. Empty list [] = no tools for text runs (there is no fallback set). Omit to leave the list unchanged. Tools already on the agent keep their current text/voice settings; only NEWLY named tools are switched on for text. So re-sending the current list plus one name adds exactly that one tool, and a tool that is currently call-only stays call-only (switch it on for text in the Tools tab, or remove it with voice_tools and name it again). Call tools are not removed by omitting them here — without voice_tools, the call surface is kept as it is; use voice_tools to change it. Does NOT affect the My AI dropdown path. | |
| resume_policy | No | When a human-paused thread hands control back to the agent: 'manual' (stays paused until an operator resumes it from the thread header), 'next_incoming_from_guest' (default — resumes on the contact's very next message), or 'after_<N>_hours' (e.g. 'after_2_hours'). With 'next_incoming_from_guest' a manual takeover lasts exactly one message; pick 'manual' when a person is expected to finish the conversation. OMIT to leave unchanged. | |
| voice_denoise | No | Server-side noise suppression (DeepFilterNet) on the INBOUND caller audio before STT — cleans background noise so the agent hears callers in loud/public places and gets fewer false barge-ins. Channel-agnostic: applies to every voice channel the agent answers on. Default false (unset). OMIT to leave unchanged. | |
| max_iterations | No | Hard cap on agentic-loop turns (LLM round-trips) per run, 1-50 (default 10). Each turn can call tools; the loop stops when the model replies with no tool call OR this cap is hit. Raise it for multi-step tool chains (e.g. browser automation: open → snapshot → fill → confirm → reply) that otherwise exhaust their turns before producing a final answer. OMIT to leave it unchanged. | |
| vision_enabled | No | Per-agent opt-in for vision content. When true, the executor splices recent image attachments from the active thread into the LLM call (Phase 3a continuous vision for Meet bot screen-share, plus any future channel that uploads images). Requires the agent's model to support vision (model_has_vision check). Default false; new calls pay zero token cost until the operator opts in. OMIT to leave the vision flag unchanged. | |
| voice_greeting | No | Opening line the agent speaks when the call connects. Pass an empty string "" to clear. Omit or null leaves unchanged. | |
| voice_stt_model | No | Speech-to-text model (Deepgram provider only): 'flux' (alias for flux-general-en), 'flux-general-en' (English Flux, LLM-powered end-of-turn), 'flux-general-multi' (multilingual Flux), or 'nova-3' (silence-based fallback). Flux variants are more responsive; nova-3 is the fallback when your Deepgram plan lacks Flux. Ignored when voice_stt_provider='gladia' (Gladia has a single model). OMIT to leave the STT model unchanged. | |
| voice_tts_speed | No | TTS playback speed multiplier (0.5-2.0, default 1.0). Yandex/OpenAI/Cartesia only — ignored for Deepgram. | |
| voice_tts_voice | No | TTS voice id — provider-specific (e.g. 'aura-2-thalia-en' for Deepgram, 'alloy' for OpenAI, 'alena' for Yandex, Cartesia voice UUID). Pass null to revert to provider default. | |
| auto_reply_rules | No | Plain-English rules injected into the fast model's system prompt as a `## Rules` block. No reserved keywords — the fast model reads them as guidance and decides per turn whether to reply directly or escalate to the main model for tools. Example: '- If the user greets, reply "Hi! How can I help?"\n- If the user asks what you can do, reply with a 1-sentence summary\n- If the question needs live data (prices, stock, booking), escalate' Engagement filtering (SKIP) belongs in trigger `conditions` (keywords, ai_filters, channel_types, cooldown), NOT here — if a message should be ignored the trigger shouldn't have fired. Pass null to clear. | |
| debounce_seconds | No | Seconds to wait after the last inbound message before generating, so rapid-fire messages coalesce into one reply (0-120, default 5). OMIT to leave unchanged. | |
| voice_max_tokens | No | Max TTS tokens per voice reply (40-200, default 100). Lower = snappier, higher = more detail. Controls speech brevity only: when the agent has voice tools, the runtime floors the underlying LLM completion cap at 400 so tool-call JSON always fits. | |
| include_job_state | No | Include current job state (active job context, tasks, notes) in the agent's prompt. OMIT to leave this flag unchanged. | |
| include_situation | No | Include situation context (channel, sender info, trigger type) in the agent's prompt. OMIT to leave this flag unchanged. | |
| native_web_search | No | Whether this agent may use the model provider's built-in web search (Anthropic, OpenAI) instead of the platform's web.search tool. Default true. Set false when the agent needs to filter results by publication date, use the workspace's own Serper key, or control the query — the built-in search offers none of those. No effect on providers without a built-in search (DeepSeek, Qwen). OMIT to leave unchanged. | |
| voice_mip_opt_out | No | Opt out of Deepgram's model-improvement program (privacy) for Flux STT. Default false. OMIT to leave unchanged. | |
| voice_speak_first | No | Who speaks first when a call connects. true (default): the agent says its greeting right away. false: the agent listens first and answers what the other person says; it greets only if they are silent for ~2 s. Use false for outbound calls where the person answering introduces themselves first. OMIT to leave unchanged. | |
| voice_record_calls | No | Record voice calls handled by this agent (stereo audio: caller and agent on separate channels), stored with the call for later playback. Default false. Ensure callers are informed of recording where consent rules apply. OMIT to leave this flag unchanged. | |
| voice_stt_keyterms | No | Domain-vocab bias for STT — names, product SKUs, etc. Passed verbatim as repeated `&keyterm=<w>` query params. Works on both Nova-3 and Flux. Prefer short phrases over full sentences. Empty list [] = no bias. Omit leaves unchanged. | |
| voice_stt_language | No | STT language code, validated against the SELECTED voice_stt_provider. 'multi' (default) enables autodetect / code-switching on either provider; a singleton like 'en', 'ru', 'uz' gives higher accuracy when the caller language is known. Deepgram provider: 'en' runs on Flux (fastest, eager end-of-turn); 'multi' and every other language run on Nova-3. Some languages — notably Thai ('th'), Vietnamese ('vi'), Indonesian ('id'), Tagalog ('tl') — are NOT in Nova-3's 'multi' auto-detect set, so those callers MUST be given an explicit code. Gladia provider: use its codes (e.g. 'uz' Uzbek, 'kk' Kazakh, 'az' Azerbaijani, 'tg' Tajik) — these are NOT valid under Deepgram. The enum below lists the union of both providers' codes; a code invalid for the chosen provider is rejected on save. OMIT to leave the STT language unchanged. | |
| voice_stt_provider | No | Speech-to-text provider. 'deepgram' (default) runs Nova-3 / Flux and preserves existing behaviour. 'gladia' routes to Gladia streaming STT (solaria-1), which is Deepgram-class latency and covers ~115 languages Deepgram does NOT — including Uzbek ('uz'), Kazakh ('kk'), Azerbaijani ('az'), Tajik ('tg'). Pick 'gladia' when the caller's language is outside Deepgram's coverage, then set voice_stt_language to that language's code. The provider also decides which language codes voice_stt_language accepts. OMIT to leave the STT provider unchanged. | |
| voice_tts_language | No | TTS language code, BCP-47 lite e.g. 'en', 'es', 'pt-BR' (Cartesia only, default 'en'). | |
| voice_tts_provider | No | Text-to-speech provider: 'deepgram' (default, Aura-2 EN-only), 'openai' (multilingual), 'cartesia' (Sonic-3, ultra-low TTFB, multilingual), 'alibaba' (CosyVoice v3-flash, multilingual, ~95ms TTFB), 'yandex' (best Russian), 'qwen' (Qwen3-TTS, self-hosted), or 'xai' (Grok, ~20 langs). OMIT to leave the TTS provider unchanged. | |
| default_calendar_id | No | Default Google calendar_id applied at the TOOL layer whenever this agent calls a calendar tool without an explicit calendar_id (e.g. 'c_...@group.calendar.google.com'). Use for agents whose bookings must always land on one dedicated calendar — a prompt-only rule is advisory and the model occasionally drops the param. An explicit calendar_id in a tool call still wins. Pass an empty string to clear (falls back to 'primary'). OMIT to leave unchanged. | |
| include_specialists | No | Inject a [SPECIALISTS] block (~50–200 tokens) listing the workspace's delegation-capable agents so a router-style agent can pick a handoff target without first calling agents.list. Default OFF for new agents; the Router template ships with this ON. Agentic mode only. OMIT to leave this flag unchanged. | |
| output_text_filters | No | Deterministic post-processing of the user-facing text an AGENTIC run produces (messages.send text + auto-delivered final answer; drafts inherit; ai_assisted fast-path replies and messages.edit are NOT covered). List of {pattern, replacement} regex rules applied in order. Example — strip AI-tell em dashes while keeping '—————' separator runs and never gluing lines: [{"pattern": "[ \\t]*(?<![—–])[—–](?![—–])[ \\t]*", "replacement": ", "}] (use [ \\t], not \\s — \\s matches newlines). Use for style rules the LLM won't reliably follow via prompt. Patterns are validated (invalid regex rejects the update). Max 20 rules. Empty list [] = clear all filters. OMIT to leave unchanged. | |
| voice_call_analysis | No | Run a post-call LLM analysis after each answered call: summary, success verdict with reason, and a 1-10 quality score, stored on the voice session and returned by calls.get_transcript metadata. Default false (costs one LLM call per call). OMIT to leave unchanged. | |
| voice_primary_model | No | Primary LLM for voice turns (e.g. 'gpt-4.1-mini', 'claude-haiku-4-5-20251001'). Pass null to revert to default. | |
| voice_turn_detector | No | Voice end-of-turn detector: 'vad' (default — sharp, low-latency) or 'multilingual' (semantic model, ~1s slower per turn, fewer mid-pause cuts). OMIT to leave unchanged. | |
| fast_prompt_override | No | Full fast-path prompt override. Placeholders substituted via .replace(): {message}, {history}, {rules}, {tools}, {output_contract}. agent.prompt_text is NOT injected into fast_prompt_override — include it yourself if you want it. Pass null to clear. | |
| pause_on_human_reply | No | Pause the agent on a thread the moment a human operator replies there. Default true. OMIT to leave unchanged. | |
| voice_filler_enabled | No | Emit 'thinking' filler audio while tools run so the caller hears life on the line (default true). OMIT to leave this flag unchanged. | |
| voice_max_tool_calls | No | Max tool calls per voice turn (1-10, default 3). OMIT to leave unchanged. | |
| voice_output_gain_db | No | Output attenuation in dB applied to the agent's WhatsApp call audio before the codec (-12.0 to 0.0, default 0 = unchanged). Use a negative value (e.g. -3.5) when a hot voice engine (OpenAI Realtime rides near 0 dBFS) causes blown-speaker distortion on loud words: the 24 kbps call codec needs headroom, and this restores it. Applied on the next call, no deploy needed. | |
| voice_realtime_model | No | Realtime (v2v) model for the agent's voice_engine. openai_realtime: 'gpt-realtime' (~18c/min on short calls) or 'gpt-realtime-mini' (~75% cheaper). gemini_realtime: a Gemini Live model, newest first 'gemini-3.8-live', 'gemini-3.1-flash-live-preview', 'gemini-2.5-flash-native-audio-latest', 'gemini-2.5-flash-native-audio-preview-12-2025' (today's default). 'default' = back to the engine's default. OMIT to leave unchanged. | |
| voice_thinking_texts | No | Pool of phrases spoken while the agent sets up the turn before calling the LLM (e.g. ['Hmm', 'So', 'One sec']). Pre-rendered to PCM at call start; one is picked at random per turn so the agent doesn't repeat the same word. Pass [] to clear. Omit or null leaves unchanged. | |
| include_learned_style | No | Include learned communication style (per-contact tone, dormancy state) in the agent's prompt. OMIT to leave this flag unchanged. | |
| voice_record_announce | No | On recorded calls, prepend a short 'this call may be recorded' notice to the agent's greeting. Only takes effect when voice_record_calls is enabled. Default false. OMIT to leave unchanged. | |
| voice_v2v_transcripts | No | Enable live transcripts for voice-to-voice engines (default true). OMIT to leave unchanged. | |
| include_thread_summary | No | Include condensed summary of older thread messages in the agent's prompt. OMIT to leave this flag unchanged. | |
| voice_endpointing_mode | No | LiveKit endpointing mode. 'fixed' (default) waits voice_endpointing_min_delay every turn; 'dynamic' adapts the wait from the conversation's own rhythm. OMIT to leave unchanged. | |
| voice_hold_ready_reply | No | Keep a finished reply the caller talked over (before any audio played) and speak it at the next pause, then answer what was said since. Without it (the default) stock LiveKit behaviour applies: that reply is discarded and regenerated from scratch. Default false. OMIT to leave unchanged. | |
| voice_transfer_numbers | No | Phone numbers (E.164, e.g. '+15551234567') the agent may cold-transfer a live call to ('let me put you through to a manager'). Any destination not on this list is refused; an empty list [] means the agent cannot transfer at all. Omit leaves unchanged. | |
| include_factual_accuracy | No | Inject the Factual Accuracy block (~100 tokens, generic anti-hallucination rules) into the system prompt. Default OFF — skip if you write domain-specific accuracy rules in Instructions. Agentic mode only. OMIT to leave this flag unchanged. | |
| knowledge_collection_ids | No | Replace all knowledge collections with these IDs (empty list = clear all) | |
| voice_flux_eot_threshold | No | Flux STT end-of-turn confidence threshold (0.1-1.0). Higher = wait for more certainty before finalizing the turn. Flux STT only (ignored on nova-3). OMIT to keep the worker default (0.7). | |
| voice_greeting_prerender | No | Gemini realtime agents only: synthesize the greeting in the agent's own voice while the phone is ringing, so it plays the instant the call is answered — word for word, no model start-up delay. Default false. OMIT to leave unchanged. | |
| voice_greeting_returning | No | Greeting variant for RETURNING callers (thread has prior history or a resolved name). Supports '{name}' — replaced with the caller's name when known, stripped when not. Pass an empty string "" to clear (always use voice_greeting). Omit or null leaves unchanged. | |
| voice_silence_reprompt_s | No | Seconds of silence from BOTH sides after which the agent briefly checks that the other person can still hear it (at most twice per call; never while the phone is still ringing). 0 = off (default). Typical 3-5 for outbound calls. OMIT to leave unchanged. | |
| include_safety_guidelines | No | Inject the generic Safety Guidelines block (~80 tokens) into the system prompt. Default OFF — enable only if you don't already write safety rules in your Instructions. Agentic mode only. OMIT to leave this flag unchanged. | |
| include_tool_call_history | No | Include the agent's own tool calls and results from the last 3 runs on this thread, compacted to IDs + top hits (~200-1000 tokens). Lets the agent recall file IDs, search hits, and decisions it already made across turns. Default ON. Agentic mode only. OMIT to leave this flag unchanged. | |
| voice_filler_audio_preset | No | Which bundled clip plays as the 'thinking' filler while the agent is working (LLM + tool calls), used when voice_thinking_texts is empty. Requires voice_filler_enabled=true. Built-in presets: 'keyboard_typing' / 'keyboard_typing2' (keyboard-typing SFX — sounds like the agent is typing/looking something up), plus any bundled music preset. Pass '' to clear (fall back to spoken filler). An unknown value silently plays no filler. OMIT or null to leave unchanged. | |
| voice_flux_eot_timeout_ms | No | Flux end-of-turn hard timeout in ms (500-15000). Flux STT only. OMIT to keep the worker default (5000). | |
| voice_max_call_duration_s | No | Wall-clock cap for a voice call in seconds, clamped to 60-14400; the worker ends the call at it regardless of what the conversation is doing. OMIT to leave unchanged (default is no cap — a call ends when the conversation ends). To REMOVE an existing cap, clear the field in the agent's Voice settings UI (0 there = no limit). Two agents talking to each other are capped separately by calls.agent_duel's own max_duration_s. | |
| voice_endpointing_max_delay | No | LiveKit endpointing.max_delay (0.5-10.0s, default 3.0). Ceiling on how long the agent waits for a turn to end. voice_endpointing_min_delay is the floor after silence; this is what stops a thinking pause from holding the turn open, and it is the knob to raise when an agent talks over someone who pauses mid-thought. | |
| voice_endpointing_min_delay | No | Silence after end-of-utterance before agent replies (0.1-2.0s, default 0.3). Higher = fewer false interrupts; lower = snappier. | |
| voice_preemptive_generation | No | Speculatively start the LLM on STT partials so the agent begins responding before end-of-utterance. Matches LiveKit stock template. Default true. OMIT to leave this flag unchanged. | |
| include_conversation_history | No | Include recent messages from this thread (up to 20) in the agent's prompt. OMIT to leave this flag unchanged. | |
| include_data_confidentiality | No | Inject the Data Confidentiality block (~250 tokens, cross-contact PII isolation + prompt-injection defense) into the system prompt. Default OFF. Agentic mode only. OMIT to leave this flag unchanged. | |
| voice_greeting_interruptible | No | Allow the caller to barge in during the opener TTS. Default true (trial-friendly — long greetings can be interrupted). Set false on outbound-call agents whose configured opener would otherwise get preempted by the caller's 'Hello?' triggering an off-script auto-turn. OMIT to leave this flag unchanged. | |
| voice_group_interim_barge_in | No | GROUP/Meet barge-in: interrupt the agent AS SOON AS a participant starts speaking (on the interim transcript), instead of waiting for the finished utterance. Default true. Set true to make a presenter/group agent easy to interrupt mid-sentence (word-gated by voice_group_barge_in_min_words so ambient noise can't trip it); false = only a completed utterance interrupts. OMIT to leave unchanged. | |
| voice_flux_eager_eot_threshold | No | Flux eager end-of-turn threshold (0.1-1.0). Setting this ENABLES EagerEndOfTurn for faster turn-taking at the cost of +50-70% LLM calls. Flux STT only. OMIT to leave eager off. | |
| voice_group_barge_in_min_words | No | GROUP/Meet barge-in word gate (1-6, default 1): a participant line must have at least this many words to interrupt the agent. 1 = any word interrupts (most responsive); raise to ignore short cross-talk. OMIT to leave unchanged. | |
| voice_early_finalize_confidence | No | Flux early-finalize confidence (0.1-1.0, worker default 0.65). Lower = finalize sooner on trailing silence (snappier, small mid-word clip risk). Flux STT only. OMIT to leave unchanged. | |
| voice_group_barge_in_stop_words | No | GROUP/Meet barge-in stop-words: any of these words/phrases ALWAYS interrupts the agent, even below voice_group_barge_in_min_words (e.g. ['стоп','подожди','вопрос','stop','wait','question']). Empty list [] clears. OMIT to leave unchanged. | |
| voice_interruption_min_duration | No | Min caller speech duration to interrupt the agent (0.1-1.5s, default 0.25). Higher = ignore short fillers like 'uh-huh'. | |
| voice_group_barge_in_requires_address | No | GROUP/Meet barge-in: when true, only lines that ADDRESS the agent (by name/keyword) interrupt it — two humans talking to each other won't break the walk. Default false. OMIT to leave unchanged. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds partial-update semantics ('only provided fields will be updated'), which annotations do not state, and that is useful for understanding the mutation behavior. However, it omits that several parameters (prompt_text, allowed_tools, voice_tools) destructively replace entire surfaces, and the claim that 'All parameters are optional' is inaccurate given the required agent_id.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and partial-update note are front-loaded, and the bulleted capability list is scannable and earns its place as a high-level map of an 85-parameter tool. It is neither bloated nor redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the extensive schema coverage and the absence of an output schema, the description provides sufficient orientation for an update tool: it states the scope, optionality, and major update categories. A warning about destructive replacements would improve it, but the schema carries that detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all 85 parameters in detail; the description adds no syntax or format guidance beyond grouping use cases. Its assertion that all parameters are optional conflicts with the required agent_id, though the schema corrects this.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Update') and resource ('AI agent's configuration'), so the action is unambiguous. It does not name or distinguish itself from siblings like agents_create or agents_update_from_template, leaving the agent to infer the boundary from the tool name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Use this to:' list gives concrete update scenarios (enable/disable, change name, assign prompt, replace knowledge collections, etc.), providing clear context for when the tool applies. It stops short of naming alternatives or when-not-to-use conditions, which would be needed for a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agents_update_from_templateAInspect
Update a forked agent's instructions (prompt) to the latest version of the system template it was created from.
Use when the platform has improved a template and the user wants their forked agent to pick up the new prompt. This OVERWRITES the agent's prompt_text with the template's current prompt — any customizations to the prompt are replaced (recoverable via prompt history). Tool/model/execution settings are NOT changed. Only works on agents forked from a template (not from-scratch agents or templates themselves).
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | ID of the forked agent to update from its template | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing what is destroyed (prompt_text overwritten, customizations replaced), the recovery path (prompt history), and what is untouched (tool/model/execution settings). The 'recoverable via prompt history' detail is fully consistent with destructiveHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by when-to-use and then the destructive/scope caveats. Every sentence carries distinct information; nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool, the critical behavioral facts (what changes, what is preserved, recoverability, applicability scope) are all covered, and the annotations already carry the safety profile. No output schema is needed, and nothing an agent requires to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented. The description still adds a meaningful constraint on the agent_id parameter — it must reference a forked agent, not a from-scratch agent or a template — which sharpens selection beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (a forked agent's instructions/prompt) and a precise target (the latest version of the system template it was created from). This is clearly distinguishable from siblings like agents_update or agents_prompt_restore.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('when the platform has improved a template and the user wants their forked agent to pick up the new prompt') plus an explicit when-not ('only works on agents forked from a template — not from-scratch agents or templates themselves'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ai_filters_createAInspect
Create a new AI filter for semantic intent-based message matching.
AI filters use vector embeddings (via Voyage AI) to detect whether an incoming message matches a specific intent or topic. The filter's description is embedded as a reference vector at creation time. When a message arrives, its embedding is compared against this reference using cosine similarity.
The description field is the most important part — it becomes the reference embedding that all incoming messages are compared against. Write it as a clear statement of what kind of messages should match:
'Customer asking about pricing, subscription plans, or billing'
'User reporting a bug, crash, or unexpected behavior in the product'
'Inbound sales lead expressing interest in purchasing or trialing'
The threshold controls sensitivity: 0.5 is a balanced default, lower values (0.3) cast a wider net, higher values (0.8) require closer matches.
Note: This tool calls the Voyage AI embedding API to generate the reference vector.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Filter name — a short, human-readable label (max 100 chars) | |
| threshold | No | Cosine similarity threshold for a message to be considered a match. Range 0.1–1.0. Default 0.50. Lower values (e.g. 0.3) are more permissive and catch more messages. Higher values (e.g. 0.8) require closer semantic similarity. | |
| description | Yes | Reference text that defines what messages should match this filter. This text is embedded as a vector and used for cosine similarity comparison against all incoming messages. Be specific and descriptive — the quality of this text directly determines filter accuracy. E.g. 'Customer asking about pricing, subscription costs, or billing issues'. Max 500 chars. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare the write/open-world/non-idempotent profile, and the description adds genuinely new behavioral context: the creation-time call to the Voyage AI embedding API, that the description field is frozen into a reference vector, and the cosine-similarity matching model. It does not disclose failure modes (e.g. embedding API errors) or the returned object, but it goes meaningfully beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then mechanism, then per-field guidance in bullets, which is easy to scan. It is somewhat long and partially repeats the schema's threshold examples and the description example, but each section still carries usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the mental model, the critical input, tuning semantics, and the external API dependency. The main omission is what the call returns (filter id/object) and any workspace-scoping caveat, though the in_workspace parameter is self-documenting in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, but the description adds interpretive meaning: it explains that the description field becomes the reference embedding and is the most important input, with concrete example phrasings, and it translates threshold numbers into sensitivity behavior (0.3 permissive vs 0.8 strict). This exceeds the schema's own wording for these two params.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new AI filter') and immediately scopes it to 'semantic intent-based message matching', which cleanly separates it from ai_filters_update, ai_filters_delete, ai_filters_list, and ai_filters_test. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives rich guidance on how to configure the tool (what to write in the description field, how threshold values behave) but never states when to create a filter versus when to list, test, or update one. Usage context is implied through the embedding workflow rather than stated, and no alternative sibling is named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ai_filters_deleteADestructiveIdempotentInspect
Permanently delete an AI filter.
When to use:
User wants to remove a filter they no longer need
This action cannot be undone. Any triggers that reference this filter by ID will no longer match it — review and update those triggers after deletion.
| Name | Required | Description | Default |
|---|---|---|---|
| filter_id | Yes | ID of the filter to delete | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true and readOnlyHint=false, so the safety profile is covered. The description adds genuinely new consequence information: irreversibility and the cross-entity side effect that triggers referencing the filter by ID will stop matching, prompting follow-up work. It stops short of 5 because it says nothing about required permissions or the response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then a scannable 'When to use' bullet, then the critical warning. Every sentence carries weight and none restates the tool name or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive single-entity delete with no output schema needed, the description covers irreversibility, downstream trigger breakage, and the use case. An agent has everything required to decide and act; the only minor omission is a pointer to ai_filters_list for obtaining IDs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so filter_id and in_workspace are already fully documented in the schema. The description adds no additional meaning about parameter format, ID acquisition, or workspace-scoping behavior, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Starts with a specific verb+resource ('Permanently delete an AI filter') that immediately distinguishes it from the create/list/test/update siblings in the ai_filters_* family. The word 'Permanently' additionally sets it apart from ai_filters_update, which an agent might otherwise confuse with removal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Includes an explicit 'When to use' block ('User wants to remove a filter they no longer need'), giving clear positive context for selection. It does not name alternatives or when-not-to-use conditions (e.g. disabling rather than deleting), so it falls short of the explicit routing a 5 requires.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ai_filters_listARead-onlyIdempotentInspect
List all AI filters for the current workspace.
AI filters are semantic intent-based message filters that use embeddings (vector representations) to detect whether an incoming message matches a specific intent or topic. Unlike keyword filters, they understand meaning: 'I need help with my order' and 'my package hasn't arrived' both match a 'shipping support' filter even without shared keywords.
Each filter stores a reference embedding of its description. When a message arrives, its embedding is compared via cosine similarity against the filter's reference vector. If the similarity exceeds the threshold, the filter matches.
When to use:
Check which semantic filters already exist before creating a new one
Get filter IDs for use in trigger conditions
Review thresholds and active status of existing filters
Returns all filters with id, name, description, threshold, and is_active.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so safety is covered; the description adds real value by listing the returned fields (id, name, description, threshold, is_active) in the absence of an output schema. It does not mention permissions, pagination, or limits, but for a scoped read tool this is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is correctly front-loaded, but the two-paragraph explanation of embeddings and cosine similarity is heavy for a list operation and repeats conceptual material an agent rarely needs to call the tool. The 'When to use' and return-field lines earn their place; the mechanics exposition is mostly filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the explicit enumeration of returned fields is a meaningful completeness win, and the when-to-use guidance covers the main invocation contexts. Only the optional in_workspace override is left to the schema, which is acceptable given full schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one optional parameter (in_workspace) at 100% schema description coverage, so the schema fully carries its semantics. The description never references the parameter, so the baseline 3 for high-coverage schemas applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb and resource ('List all AI filters for the current workspace'), and the following paragraphs define what an AI filter actually is versus keyword filters. It never names a sibling tool (ai_filters_create, ai_filters_test, ai_filters_update) explicitly, so differentiation is implied through the 'before creating a new one' framing rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
A dedicated 'When to use' block gives three concrete scenarios: audit existing filters before creating, fetch filter IDs for trigger conditions, and review thresholds/active status. That is strong context, but there are no stated exclusions or explicit named alternatives (e.g. 'use ai_filters_test to test a filter'), which keeps it below a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ai_filters_testARead-onlyIdempotentInspect
Test a message against an AI filter to check whether it would match.
This tool embeds the provided message using Voyage AI and computes the cosine similarity between the message vector and the filter's stored reference vector. It returns the similarity score, whether the message would match (similarity >= threshold), and the filter's threshold value.
Use this to:
Verify a filter works as intended before using it in a trigger
Tune the threshold by testing borderline messages
Debug why a message did or did not match a filter in production
Returns: {similarity: float, matched: bool, threshold: float}
Note: This tool calls the Voyage AI embedding API to embed the test message.
| Name | Required | Description | Default |
|---|---|---|---|
| message | Yes | The message text to test. This is embedded and compared against the filter's reference vector via cosine similarity. | |
| filter_id | Yes | ID of the filter to test against | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=true. The description goes beyond them by disclosing the external Voyage AI embedding call (the reason openWorldHint is true), the exact computation performed, and the shape of the returned score. It stops short of noting latency, cost, or failure modes of that external call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then layers mechanism, use cases, return contract, and a dependency note. Sections are clearly delineated and no sentence is filler; the external-API note prevents misreading this as a pure local computation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return shape ({similarity, matched, threshold}), the match rule (similarity >= threshold), all required parameters, and the external dependency. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents filter_id, message, and in_workspace fully. The description adds conceptual context around how the message is embedded but no parameter-specific semantics beyond what the schema states, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Test a message against an AI filter') and immediately clarifies the mechanism (embedding + cosine similarity). It is trivially distinguishable from its CRUD siblings (ai_filters_create/update/delete/list) because 'test' is its own action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The explicit 'Use this to:' block enumerates three concrete scenarios — verify before wiring into a trigger, tune threshold on borderline messages, and debug production match/mismatch. This tells the agent both when and why to reach for this tool rather than the filter CRUD siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ai_filters_updateAInspect
Update an existing AI filter's name, description, threshold, or active state.
When to use:
User wants to rename a filter
User wants to refine the filter description to improve match accuracy
User wants to adjust the similarity threshold (higher = stricter matching)
User wants to enable or disable a filter without deleting it
Provide only the fields you want to change. At least one field is required.
Note: If the description is changed, this tool calls the Voyage AI embedding API to re-generate the reference vector with the new description text.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | New filter name (max 100 chars, optional) | |
| filter_id | Yes | ID of the filter to update | |
| is_active | No | Enable (true) or disable (false) the filter. OMIT to leave the active flag unchanged. | |
| threshold | No | New cosine similarity threshold. Range 0.1–1.0. Optional. | |
| description | No | New reference description text. If changed, the Voyage AI embedding API is called to re-generate the reference vector. Max 500 chars. Optional. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations covering safety (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true), the description still adds real behavioral context: partial-update semantics and, crucially, that changing the description triggers an external Voyage AI embedding call to regenerate the reference vector. This explains the non-idempotent hint rather than just repeating it. It does not discuss permissions or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then a scannable bulleted 'When to use' list and a clearly marked Note for the side effect. Efficient overall, though several bullets largely restate the fields already enumerated in the opening sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with no output schema, the description covers the essential gaps: partial-update contract, minimum-field requirement, threshold direction, and the external embedding side effect. Only auth/permission requirements and error conditions are absent, which are minor for this tool class.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine semantic value the schema lacks: 'threshold (higher = stricter matching)' clarifies the direction of the similarity metric, and 'enable or disable a filter without deleting it' frames is_active's purpose. It does not touch in_workspace, which the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (AI filter) plus the exact mutable fields (name, description, threshold, active state), which cleanly separates it from ai_filters_create, ai_filters_delete, ai_filters_list and ai_filters_test. An agent can identify the tool's role without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'When to use' block gives explicit trigger conditions for each field, and it states the key constraint 'Provide only the fields you want to change. At least one field is required.' It stops short of naming alternatives (e.g., use ai_filters_create to make a new filter, ai_filters_delete to remove one), so the routing guidance is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ai_tags_add_to_threadAInspect
Apply one or more AI tags to a thread (manually).
When to use:
User wants to label a conversation with one or more tags
User asks to categorize or tag a thread
Provide the thread_id (integer) and an array of tag_ids to apply. If a tag is already applied it will be updated to is_manual=true.
| Name | Required | Description | Default |
|---|---|---|---|
| tag_ids | Yes | Array of tag IDs to apply (1–20 IDs) | |
| thread_id | Yes | ID of the thread to tag | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the write/safe/idempotency profile (readOnly=false, destructive=false, idempotent=false). The description adds genuinely new behavior: applying an already-present tag flips its is_manual flag, which explains why the operation is not purely idempotent. Return format is undocumented but that is minor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then a compact 'When to use' block, then the required inputs. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool whose safety profile is carried by annotations, the description covers action, triggers, required inputs, and the update-on-reapply behavior. Missing only edge-case handling (e.g., invalid tag_ids or the 1-20 limit noted in schema), which is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters including in_workspace. The description restates thread_id and tag_ids but adds no format, constraint, or syntax detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (apply) and resource (AI tags to a thread), so an agent can tell what it does at a glance. It does not explicitly distinguish itself from siblings like ai_tags_remove_from_thread or ai_tags_list, so it stays just below the top band.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'When to use' section gives two concrete triggering scenarios (labeling a conversation, categorizing/tagging a thread). There are no explicit exclusions or named alternatives, but the use context is clearly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ai_tags_createAInspect
Create a new AI tag (automatic message filter).
AI tags are lightweight classifiers that run on every incoming message. When a message matches the tag's description/criteria, the thread is automatically labelled — so AI agents can cheaply pre-filter threads instead of running full LLM analysis on everything. Good descriptions are the key: they tell the classifier exactly when to apply this tag.
When to use:
User wants to auto-classify incoming messages (e.g. bug reports, sales leads, support requests)
User wants to reduce AI agent costs by pre-filtering threads by topic or intent
Tips for the description field:
Be specific: 'Messages reporting errors, crashes, or unexpected behavior in the product'
Include examples of what qualifies and what doesn't
Limit: 20 active personal tags / 50 active team tags.
| Name | Required | Description | Default |
|---|---|---|---|
| icon | No | Emoji icon for the tag (max 10 chars, optional) | |
| name | Yes | Tag name (max 100 chars) | |
| color | No | Tailwind color key for the tag badge. Allowed: amber, blue, green, red, purple, yellow, slate. OMIT to use the default color. | |
| applies_to | No | Where the tag can be used: 'thread' (conversations, default), 'task' (the task board), or 'both'. | thread |
| description | No | Classifier prompt: describe exactly when this tag should be applied to a thread. The more specific, the better the auto-classification accuracy. E.g. 'Messages reporting software errors, crashes, or unexpected behavior'. Max 500 chars. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-idempotent create. The description adds meaningful behavior beyond them: the classifier runs on every incoming message and auto-labels matching threads, and there is a hard quota (20 personal / 50 team tags). It stops short of saying what happens at the limit or on duplicate names.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence and the thread-classifier mechanism follows immediately. The tips and limit sections are useful rather than filler, though the block is longer than strictly required for a 6-param create call.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description never indicates what creation returns (e.g., the new tag id), which is the one real gap. Everything else an agent needs — semantics, quota constraints, field guidance — is covered, so it is close to complete for this complexity level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description meaningfully extends the 'description' parameter with guidance on writing classifier criteria and examples ('Messages reporting errors, crashes...'). That is real semantic value the schema text alone does not carry.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Create a new AI tag') and goes further by defining what an AI tag actually is ('lightweight classifiers that run on every incoming message'). An agent understands both the operation and the artifact it produces, which is enough to distinguish this from ai_tags_update/delete/list without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is an explicit 'When to use' section with two concrete scenarios (auto-classify incoming messages; reduce agent cost via pre-filtering). However, it offers no when-NOT-to-use guidance and never mentions the sibling ai_filters_create, which an agent could easily confuse with this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ai_tags_deleteADestructiveIdempotentInspect
Delete a personal AI tag. All thread associations are removed automatically.
When to use:
User wants to permanently remove a tag they no longer need
This cannot be undone. Threads are NOT deleted — they just lose this tag.
| Name | Required | Description | Default |
|---|---|---|---|
| tag_id | Yes | ID of the tag to delete | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, but the description goes further with concrete consequences the annotations cannot express: cascading removal of all thread associations, explicit irreversibility ('This cannot be undone'), and the important negative that threads themselves are preserved. This is exactly the 'what gets destroyed' context the annotations lack.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action and its cascade effect are front-loaded in the first two sentences, and the destructive warning is placed last for emphasis. The 'When to use' header with a single bullet is mild structural overhead, but nothing is padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema mutation tool whose annotations cover the safety profile, the description supplies the missing cascade and irreversibility semantics. It stops short of stating failure behavior (e.g. nonexistent tag_id) or permission requirements, which are the only remaining gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with only two parameters, so tag_id and the workspace override are already fully documented in the schema. The description adds no syntax, format, or constraint detail for either parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb + resource ('Delete a personal AI tag') with an explicit scope qualifier ('personal'), which separates it from ai_tags_create, ai_tags_update, and especially ai_tags_remove_from_thread. The closing sentence 'Threads are NOT deleted — they just lose this tag' resolves the exact ambiguity the sibling name could otherwise create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'When to use' block states a clear trigger condition ('User wants to permanently remove a tag they no longer need'). It also implicitly disambiguates from ai_tags_remove_from_thread by clarifying that threads survive the operation, but it never names that alternative tool or states when not to use this one.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ai_tags_listARead-onlyIdempotentInspect
List all personal AI tags.
AI tags are automatic message filters: the system runs a lightweight classifier on every incoming message and applies matching tags to threads. This lets AI agents skip expensive full analysis on most messages — they only act on threads that match relevant tags, dramatically cutting LLM costs.
When to use:
Check which auto-classification filters exist before creating one
Get tag IDs for add_to_thread / remove_from_thread
See how many threads each tag currently matches
Returns all tags with thread counts (non-archived, included threads only).
| Name | Required | Description | Default |
|---|---|---|---|
| applies_to | No | Filter to tags usable on 'thread' (conversations) or 'task' (the task board). OMIT to list all tags. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so the safety profile is covered. The description adds value beyond that by explaining the classification mechanism and, most usefully, the exact scope of the returned counts ('non-archived, included threads only'). It does not state whether counts are real-time or cached, which is the only remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then a scannable bulleted usage block, then return semantics. The middle explanatory paragraph on classifier cost is longer than strictly needed for tool selection but does carry real context, so it mostly earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of return values and does so ('all tags with thread counts, non-archived, included threads only'), plus usage triggers and the tag domain model. Nothing needed to call it correctly is missing, aside from per-parameter detail that the schema already supplies.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are documented in-schema, including the enum meaning for applies_to and the non-persistent semantics of in_workspace. The description adds nothing about either parameter (it never mentions the thread/task filter or cross-workspace execution), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('List all personal AI tags') and then defines what an AI tag actually is (automatic message filters applied by a lightweight classifier). It differentiates itself from siblings by naming the create and add/remove_from_thread workflows it feeds into, so an agent can place it in the tag lifecycle without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
An explicit 'When to use' block lists three concrete triggers: pre-checking before creating a tag, retrieving tag IDs for add_to_thread / remove_from_thread, and inspecting per-tag thread counts. The relationship to sibling tools is stated, so the routing decision is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ai_tags_remove_from_threadADestructiveIdempotentInspect
Remove a specific AI tag from a thread.
When to use:
User wants to un-label or remove a specific tag from a conversation
User wants to correct an incorrectly applied tag
Provide both thread_id and tag_id.
| Name | Required | Description | Default |
|---|---|---|---|
| tag_id | Yes | ID of the tag to remove | |
| thread_id | Yes | ID of the thread to remove the tag from | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true and readOnlyHint=false, so the mutation/idempotency profile is covered. The description adds that this un-associates a tag from a thread ('un-label') rather than deleting it, which is useful, but says nothing about permissions, side effects, or what happens to other threads using the same tag.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short lines, front-loaded with the action, followed by scannable bullets and a one-line input reminder. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a scoped tag-removal tool with a full schema and annotations covering the safety profile, the description gives enough to invoke correctly. The only gap is the un-explained in_workspace parameter, which is covered by the schema anyway.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so tag_id, thread_id and in_workspace are already documented in the schema. The description only restates that thread_id and tag_id are required and adds no format, range, or scoping detail (in_workspace is never mentioned). Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Remove a specific AI tag from a thread.' The wording 'un-label / remove from a conversation' implicitly separates it from sibling ai_tags_delete (which removes the tag itself) and ai_tags_add_to_thread (the inverse). An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a dedicated 'When to use' block with two concrete scenarios (un-labeling, correcting a misapplied tag) and states the required inputs. However, it does not name an alternative tool or state when-not to use it (e.g., ai_tags_delete when the tag should be destroyed).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ai_tags_updateAInspect
Update an existing personal AI tag's name, description, icon, color, or active state.
When to use:
User wants to rename a tag
User wants to change a tag's icon, color, or description
User wants to enable or disable a tag
Provide only the fields you want to change. At least one field is required.
| Name | Required | Description | Default |
|---|---|---|---|
| icon | No | New emoji icon (max 10 chars, optional) | |
| name | No | New tag name (max 100 chars, optional) | |
| color | No | New color key. Allowed: amber, blue, green, red, purple, yellow, slate. OMIT to leave the color unchanged. | |
| tag_id | Yes | ID of the tag to update | |
| is_active | No | Enable (true) or disable (false) the tag. OMIT to leave the active flag unchanged. | |
| applies_to | No | Where the tag can be used: 'thread' (conversations, default), 'task' (the task board), or 'both'. OMIT to leave it unchanged. Narrowing it away from a side detaches the tag from that side's existing conversations or tasks. | |
| description | No | New LLM hint (max 500 chars; empty string clears it, optional) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (not read-only, not destructive, not idempotent, closed-world). The description usefully adds partial-update semantics ("Provide only the fields you want to change. At least one field is required"), which is real behavioral context. It does not mention permission requirements, reversibility, or side effects of disabling, so it stays at a moderate level.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by a scannable "When to use" list and a closing constraint sentence. Every line earns its place with no redundancy, though the bullets restate field names already implied by the opening sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with a title annotation, full schema coverage, and no output schema, the description covers purpose, selection cues, and the partial-update contract adequately. It is only slightly incomplete in not addressing sibling routing (create vs. delete vs. add/remove-from-thread) that an agent choosing among ai_tags_* tools would want.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so per-parameter meaning is already fully carried by the schema (including OMIT-to-leave-unchanged and enum values). The description only adds the cross-parameter constraint that at least one field must be supplied, which is marginal added value over the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ("Update an existing personal AI tag") and enumerates the mutable fields, so an agent immediately knows what it does. It implicitly separates itself from create/delete/list siblings via "existing" and "update", but never names an alternative to disambiguate against them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The explicit "When to use" bullets map user intents (rename, change icon/color/description, enable/disable) to this tool, which is clear context for selection. It stops short of the 5-level bar because it never states when NOT to use it or names the create/delete siblings that a user might also be reaching for.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
analytics_queryARead-onlyIdempotentInspect
Answer analytics questions about this workspace's own data (conversations, messages, voice calls) — e.g. 'how many new conversations this week by channel' or 'inbound vs outbound messages per day this month'. Returns rows plus a chart hint. Read-only and scoped to the current workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | The analytics question in natural language, e.g. 'how many conversations started this week, by channel?' | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so safety is covered structurally. The description adds value beyond that by disclosing the return shape ('Returns rows plus a chart hint') and the current-workspace scoping, which the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, purpose front-loaded, examples inline, and the return/scoping caveats tacked on last. No filler or repetition; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully notes 'Returns rows plus a chart hint,' and the workspace scoping plus read-only nature round out the picture for a 2-parameter NL query tool. Minor gaps remain around row limits or result shape detail, but nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema (including the in_workspace semantics about not storing anything and not affecting other sessions). The description's examples mirror the schema's own example for 'question' and add no new parameter-level syntax or constraints, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Answer analytics questions) and resource (this workspace's own data: conversations, messages, voice calls), with two concrete example questions that pin down the scope. It is readily distinguishable from data-access siblings like db_query, knowledge_query, and threads_channel_account_stats.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'about this workspace's own data' framing and the example questions give clear context for when to reach for it, and the read-only/current-workspace scoping tells the agent it is safe for exploratory analytics. It stops short of naming alternatives (db_query, search_messages) or stating when NOT to use it, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android_close_appADestructiveInspect
Close an app on the Android device (safe method: HOME key + am kill). Never uses am force-stop (corrupts PackageManager). Example: 'com.whatsapp'.
| Name | Required | Description | Default |
|---|---|---|---|
| package | Yes | Android package name to close (e.g., 'com.whatsapp') | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| channel_account_id | No | Android device channel_account ID. Omit when the workspace has a single Android device; REQUIRED when it has more than one (e.g. WhatsApp + LINE), else the call is rejected as ambiguous. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so safety is covered. The description adds genuine non-obvious context by disclosing the implementation strategy (HOME key + am kill, never am force-stop because it corrupts PackageManager), which tells the agent what actually happens to the device state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, followed by a compact parenthetical on the safe method and an example. Every element is short and purposeful, though the example clause is slightly redundant with the schema's own package example.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-action device control tool with no output schema and full parameter coverage, the description supplies purpose, method, and the key safety caveat. It omits edge cases such as behavior when the app isn't installed or running, but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents package, in_workspace, and channel_account_id in detail. The description only repeats the package example ('com.whatsapp'), which the schema already supplies, adding nothing extra. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (close) and resource (an app on the Android device), making it immediately distinguishable from siblings like android_launch_app or android_current_app. The scope and effect are unambiguous without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clarifies the internal mechanism and rules out am force-stop, but never states when to prefer this over android_shell or other alternatives, nor any prerequisites (e.g. app must be running). Usage is implied rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android_current_appBRead-onlyIdempotentInspect
Get the currently focused app and activity on the Android device.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| channel_account_id | No | Android device channel_account ID. Omit when the workspace has a single Android device; REQUIRED when it has more than one (e.g. WhatsApp + LINE), else the call is rejected as ambiguous. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond that (no statement about the return format, staleness, or whether the device must be unlocked), so it clears the lower bar set by annotations but adds little.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the resource is identified immediately and nothing is repeated from the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only query tool with full annotation coverage and no output schema, the description tells the agent what it retrieves (app plus activity), which is enough to call it correctly. It could note that results reflect a single device when multiple are connected, but the schema already handles that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (in_workspace, channel_account_id) are fully documented in the schema, including the multi-device ambiguity rule. The description adds nothing about parameters, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Get') plus a precise resource: the currently focused app and activity. This implicitly separates it from android_list_apps (enumerates all apps) and android_ui_dump (dumps the UI tree), but it never names those siblings explicitly, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states what the tool returns but gives no guidance on when to reach for it versus android_list_apps, android_ui_dump, or android_screenshot, nor any prerequisites or exclusions. Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android_install_apkADestructiveInspect
Install an app on the Android device from an APK or split bundle (XAPK/APKS — how most large modern apps ship). Provide apk_url (an https link the device downloads itself — preferred, and the only workable option for real apps) or, for tiny payloads, as apk_base64. Max 300MB. Large apps can take several minutes; if this call times out the install keeps running on the device — verify with the app-list tool instead of re-installing. Use reinstall=true to uninstall then install (for signature changes).
| Name | Required | Description | Default |
|---|---|---|---|
| apk_url | No | https URL of the APK — the server downloads it (preferred; the only practical option for real-sized apps). | |
| package | No | Expected package name (for verification) | |
| reinstall | No | Uninstall first if package exists (for different signature). OMIT for a normal install (default false). | |
| apk_base64 | No | APK as base64-encoded bytes. Only usable for tiny payloads (the whole blob is an LLM tool argument); prefer apk_url. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| channel_account_id | No | Android device channel_account ID. Omit when the workspace has a single Android device; REQUIRED when it has more than one (e.g. WhatsApp + LINE), else the call is rejected as ambiguous. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive/not-idempotent, but the description adds non-obvious behavior beyond them: the 300MB cap, that large installs take minutes, and critically that a timeout does NOT cancel the install (it keeps running on the device). That async semantics disclosure is exactly the kind of context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then the input options, then the timeout caveat. Dense with parenthetical asides but each sentence carries actionable content; slightly long for the amount of core instruction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 6 optional params, the description covers what an agent needs: inputs, size limits, reinstall semantics, and the failure/verification path. The device-disambiguation requirement lives in the channel_account_id schema description, so nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema carries the parameter burden (baseline 3). The description still adds value beyond it: the 300MB size ceiling and the APK/XAPK/APKS split-bundle framing, plus reinforcing the apk_url-over-base64 preference. Modest but real addition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (install) and resource (app/APK on the Android device), and even clarifies the artifact types supported (APK/XAPK/APKS). An agent can distinguish it from android_launch_app (launch installed app) and android_list_apps (verify install) without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes between apk_url (preferred, only workable option for real apps) and apk_base64 (tiny payloads only), tells when to set reinstall=true (signature changes), and gives the alternative path on timeout: verify with the app-list tool rather than re-installing. No inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android_keyBDestructiveInspect
Send a key event to the Android device. Use keycode (e.g., '3' for HOME, '4' for BACK) or common name (e.g., 'HOME').
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | Key code: numeric (e.g., '3', '4') or common name (HOME, BACK, POWER, VOLUME_UP, VOLUME_DOWN, MENU, SEARCH, ENTER, DEL) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| channel_account_id | No | Android device channel_account ID. Omit when the workspace has a single Android device; REQUIRED when it has more than one (e.g. WhatsApp + LINE), else the call is rejected as ambiguous. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds only the keycode naming convention (a parameter concern) and discloses nothing extra, such as which app/foreground context receives the event or that BACK/HOME may discard in-app state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed immediately by the accepted input forms. Efficient, though the second sentence partly restates what the schema's 'code' description already provides.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-action tool with 100% schema coverage and full annotations, the description covers the action and the input convention adequately. Device-selection ambiguity and failure behavior are handled by the schema, so no critical gap remains for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%: the schema already documents 'code' with the numeric/name list and fully explains in_workspace and channel_account_id (including the ambiguity-rejection rule). The description's examples ('3' for HOME, '4' for BACK, 'HOME') largely duplicate the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Send a key event to the Android device'), which clearly separates it from android_tap, android_type, and browser_press_key. It does not, however, explicitly name which sibling to prefer for navigation-key input versus text/tap input.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the two accepted input forms ('keycode ... or common name'), which implicitly signals it is for hardware/navigation keys rather than UI targeting. It never states when to use this instead of android_tap, android_type, or android_shell, so the selection guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android_launch_appBDestructiveInspect
Launch an app on the Android device by package name. Example: 'com.whatsapp'. Returns success status and process ID if launched.
| Name | Required | Description | Default |
|---|---|---|---|
| package | Yes | Android package name to launch (e.g., 'com.whatsapp') | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| channel_account_id | No | Android device channel_account ID. Omit when the workspace has a single Android device; REQUIRED when it has more than one (e.g. WhatsApp + LINE), else the call is rejected as ambiguous. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=true, openWorldHint=true, and idempotentHint=false. The description adds only that it returns success status and a process ID if launched; it does not explain what makes the operation destructive or mention auth or rate-limit considerations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, front-loaded with the action and followed by a compact example and return note. It avoids filler and is appropriately sized for a simple tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 3-parameter launch tool with full schema coverage and annotations covering the safety profile, the description covers the action and return values adequately. It omits device-selection prerequisites and destructive-behavior detail, but those are already in the schema and annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents each parameter, including the same 'com.whatsapp' example and the channel_account_id ambiguity rule. The description's example adds no meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Launch an app') plus the Android device context and a concrete package-name example. It does not explicitly name sibling alternatives like android_install_apk or android_close_app, but the action verb 'launch' inherently distinguishes it from install/close/list tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no when-to-use or when-not-to-use guidance, no prerequisites, and no named alternatives. The agent must infer usage entirely from the tool name and the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android_list_appsARead-onlyIdempotentInspect
List installed apps on the Android device. By default returns third-party apps; use all=true for all apps.
| Name | Required | Description | Default |
|---|---|---|---|
| all | No | Include system apps. OMIT for third-party only (default false). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| channel_account_id | No | Android device channel_account ID. Omit when the workspace has a single Android device; REQUIRED when it has more than one (e.g. WhatsApp + LINE), else the call is rejected as ambiguous. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so safety is covered. The description adds one genuine behavioral fact beyond the annotations: the default result set is third-party apps only. It says nothing about ordering, volume, or response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with what the tool does and then the default-vs-override behavior. No filler and nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-required-parameter read tool with 100% schema coverage and annotations covering safety, the description is nearly complete. Return fields are undocumented, though there is no output schema and a list of apps is largely self-explanatory.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (all, in_workspace, channel_account_id) are already documented in the schema. The description only restates the `all` semantics, adding no meaning beyond what structured data provides. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List installed apps') scoped to 'the Android device', which naturally separates it from android_launch_app, android_close_app and android_ui_dump. It is clear on its own, but never explicitly differentiates itself from any sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives one useful routing rule ('use all=true for all apps') and states the default scope, but offers no guidance on when to reach for this tool versus alternatives such as android_ui_dump. Usage is implied rather than explained.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android_screenshotBRead-onlyIdempotentInspect
Capture a screenshot from the Android device. Returns JPEG image data encoded as base64.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| channel_account_id | No | Android device channel_account ID. Omit when the workspace has a single Android device; REQUIRED when it has more than one (e.g. WhatsApp + LINE), else the call is rejected as ambiguous. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description usefully adds the output encoding (JPEG as base64), but says nothing about resolution, latency, or whether the capture includes overlays — modest added value over the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the action stated first and the return format second. Nothing could be removed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by stating the return type and encoding. Given the low complexity (2 optional params, fully documented) and the annotation coverage, the definition is nearly complete; only device-selection context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (in_workspace, channel_account_id) are already fully documented in the schema, including the ambiguity-rejection rule. The description adds no parameter-level meaning, which is the expected baseline here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Capture a screenshot from the Android device") and adds the return format, which clearly separates it from browser_take_screenshot and android_ui_dump. It does not name the sibling it replaces, but the Android-device scoping makes the domain unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to prefer this over android_ui_dump (structured view hierarchy) or browser_take_screenshot (web pages), nor any note about the multi-device ambiguity case. The only selection hint lives in the channel_account_id schema text, not in the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android_shellBRead-onlyIdempotentInspect
Run an allowlisted shell command on the Android device. First element of cmd array must be allowlisted: input, pm, dumpsys, am, getprop, screencap, monkey, uiautomator, settings, wm, cmd, content, ls, cat, pidof, ps, echo, service.
| Name | Required | Description | Default |
|---|---|---|---|
| cmd | Yes | Shell command to run (array form preferred, space-separated accepted). Example: 'getprop ro.build.version.release' or ['getprop', 'ro.build.version.release'] | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| channel_account_id | No | Android device channel_account ID. Omit when the workspace has a single Android device; REQUIRED when it has more than one (e.g. WhatsApp + LINE), else the call is rejected as ambiguous. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and idempotentHint=true, but the description explicitly licenses state-changing commands (input, am, pm, settings, content, monkey) which modify the device and are not idempotent. The description therefore contradicts the safety annotations rather than reinforcing them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: purpose first, then the binding constraint. The 15-item allowlist is bulky but every token is required information, and nothing is redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so return-format omissions are acceptable, and the core allowlist constraint is covered. But for a shell-execution tool the description is silent on timeouts, output capture/truncation, error behavior when a command is not allowlisted, and what execution context the command runs in.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds a real constraint the schema omits: the first element of cmd must be on a fixed allowlist, effectively an implicit enum of valid values. It does not, however, clarify the string-vs-array type handling beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run an allowlisted shell command on the Android device') and immediately scopes it with the command allowlist. An agent can distinguish this generic shell escape hatch from the dedicated action tools (android_tap, android_type, android_screenshot) without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states the allowlist but gives no when-to-use guidance and never mentions the sibling alternatives that overlap (android_tap, android_type, android_screenshot, android_ui_dump). Nothing tells the agent when to reach for android_shell versus those dedicated tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android_swipeBDestructiveInspect
Swipe/drag on the Android device screen from one coordinate to another. Use for scrolling, opening app drawer, etc.
| Name | Required | Description | Default |
|---|---|---|---|
| x1 | Yes | Starting X coordinate in device pixels | |
| x2 | Yes | Ending X coordinate in device pixels | |
| y1 | Yes | Starting Y coordinate in device pixels | |
| y2 | Yes | Ending Y coordinate in device pixels | |
| duration_ms | No | Duration of swipe in milliseconds (default: 300) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| channel_account_id | No | Android device channel_account ID. Omit when the workspace has a single Android device; REQUIRED when it has more than one (e.g. WhatsApp + LINE), else the call is rejected as ambiguous. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false and openWorldHint=true, so the safety profile is covered. The description adds no behavioral context beyond what the annotations provide — it doesn't explain why a swipe is flagged destructive or how the swipe interacts with device state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, zero padding. Slightly weakened by the vague "etc." tail but otherwise efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for selection given full schema coverage and annotations, but it omits how to choose between android_swipe and android_tap, and gives no coordinate-system or duration context beyond the schema. Complete enough to invoke, not enough to route confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all 7 parameters including duration_ms default and the channel_account_id requirement. The description adds no syntax, coordinate-system, or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (swipe/drag) and resource (Android device screen) with the coordinate-to-coordinate scope. The verb naturally distinguishes it from siblings like android_tap or android_key, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use for scrolling, opening app drawer, etc." gives implied usage scenarios but no explicit when-to-use-vs-alternatives guidance (e.g., when to swipe instead of android_tap). The trailing "etc." leaves the boundary fuzzy.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android_tapBDestructiveInspect
Tap on the Android device screen at specified coordinates. Coordinates are in device pixels.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | X coordinate in device pixels | |
| y | Yes | Y coordinate in device pixels | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| channel_account_id | No | Android device channel_account ID. Omit when the workspace has a single Android device; REQUIRED when it has more than one (e.g. WhatsApp + LINE), else the call is rejected as ambiguous. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds essentially nothing beyond that: it does not warn that a tap can trigger irreversible UI actions, and its only extra content (device pixel coordinates) merely restates the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the action and the unit front-loaded, zero filler. Nothing to trim without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple four-parameter tap with full annotations and complete schema descriptions, the core is adequate. What is missing is workflow context: how to obtain the coordinates and how this differs from android_swipe, which an agent driving a device would need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including the nuanced channel_account_id ambiguity rule and the in_workspace scoping note, so the schema carries the parameter burden. The description's coordinate-unit remark duplicates the x/y descriptions rather than adding meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (Tap), a target (Android device screen), and the coordinate unit (device pixels), so the agent knows exactly what the tool does. It does not, however, distinguish itself from the closest sibling android_swipe or place itself in the android_* family workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance: nothing says when to prefer android_tap over android_swipe or android_shell, nor that coordinates typically come from a prior android_screenshot/android_ui_dump. The agent must infer the workflow entirely on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android_typeCDestructiveInspect
Type text on the Android device keyboard.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Text to type | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| channel_account_id | No | Android device channel_account ID. Omit when the workspace has a single Android device; REQUIRED when it has more than one (e.g. WhatsApp + LINE), else the call is rejected as ambiguous. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true, and idempotentHint=false, so the safety profile is covered. The description adds nothing beyond the name: it does not say whether existing text is replaced, whether a field must be focused first, or what happens on failure, which matters for a destructive typing action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the action front-loaded and zero waste. It is terse to the point of under-specification, but structurally sound.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple input tool with full schema coverage and annotations carrying the safety profile, the minimal description is close to adequate. It is missing the key operational context that a text field must be focused and that typing may overwrite existing content, which the destructiveHint implies but the text never confirms.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the descriptions are rich (channel_account_id documents the ambiguity rejection, in_workspace documents scoping). The description contributes nothing about parameters, so the baseline of 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (type) and resource (text on the Android device keyboard), which is clearly distinct from android_tap, android_swipe, and android_shell. It does not, however, explicitly distinguish itself from the closely related android_key or browser_type siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance and no named alternatives. An agent gets no signal about prerequisites (e.g. a focused text field) or how this differs from android_key when it needs to enter text.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android_ui_dumpARead-onlyIdempotentInspect
Get UI structure dump from the Android device. Uses dumpsys (uiautomator OOM-kills on this build). Returns window list and focused activity for frame-based clicking.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| channel_account_id | No | Android device channel_account ID. Omit when the workspace has a single Android device; REQUIRED when it has more than one (e.g. WhatsApp + LINE), else the call is rejected as ambiguous. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the description only needs to add context — and it does: it names the underlying mechanism (dumpsys) and explains why (uiautomator OOM-kills on this build), plus what comes back. It stops short of describing output format or any auth/side-effect detail, but that is minor given full annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the purpose, followed by the mechanism note and the return content. The parenthetical implementation aside is slightly tangential but justifies its place by explaining why dumpsys is used.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with two fully-specified optional params and no output schema, the description covers return content ('window list and focused activity') adequately. It is nearly complete; only finer return-value detail is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters (in_workspace, channel_account_id) are already fully documented, including the ambiguity rule for multiple devices. The description adds no parameter-level meaning, which is the baseline 3 expectation when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get UI structure dump from the Android device') and enumerates the payload ('window list and focused activity'), which separates it from android_screenshot. It does not explicitly name a sibling it replaces or complements, so it falls short of the top band.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'for frame-based clicking' implies the workflow (call this first to obtain coordinates for android_tap) but never states when to prefer it over android_screenshot or android_current_app, nor any exclusions. Usage is inferable rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
artifacts_createAInspect
Host a self-contained HTML page at a stable, default-private, shareable URL — the Artifact experience, in-app.
Pass exactly one of:
html— the full page: your<body>plus any<style>/<script>. Unlike documents.create, the page is served live (JavaScript runs), so charts, interactivity, and small tools work.file_id— a workspace file whose contents are already the HTML page.analytics_card_ids— ids of saved analytics dashboard cards; the platform re-runs their queries and composes one designed report page (static charts, snapshot at build time). Best way to give someone a shareable analytics report.
The page runs in a locked-down sandbox: a dedicated origin + a strict CSP. That means it is fully self-contained — it CANNOT call out to the network (fetch/XHR/WebSocket are blocked) or load anything from a CDN. Inline all assets: CSS/JS inline, images/fonts as data: URIs. Draw charts yourself as inline SVG (no external chart library).
access_level defaults to 'private' (viewable only in-app). Set 'shared' to make the unguessable link itself the capability (anyone-with-link). You can flip this later with artifacts.set_access.
Returns {artifact_id, slug, url, app_url, version, access_level}. app_url always opens for workspace members, in the app. url is the public link and is set ONLY for a shared artifact (a private one has no public page: handed out bare, the public link answers "not available"). Republish with artifacts.update — both links stay the same.
| Name | Required | Description | Default |
|---|---|---|---|
| html | No | Full self-contained HTML page. Mutually exclusive with file_id. | |
| title | Yes | Short human-readable title (page <title> + gallery label). | |
| favicon | No | Optional emoji used as the browser-tab icon (e.g. '📊'). | |
| file_id | No | Workspace file whose contents are the HTML page. Mutually exclusive with html. | |
| template | No | Optional data-driven template: HTML with {{placeholder}} tokens. When set, later artifacts.refresh(data={...}) re-renders the page server-side from tiny data payloads (no HTML round-trip) — ideal for a scheduled agent that refreshes live numbers. The initial html you pass should be this template already rendered with today's values. | |
| description | No | Optional one-line summary for the gallery card. | |
| access_level | No | 'private' (default, in-app only) or 'shared' (anyone-with-link). | private |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| analytics_card_ids | No | Compose saved analytics dashboard cards into one report page: each card's SQL re-runs through the guarded analytics engine and renders as a static chart. Data is a snapshot at build time. Mutually exclusive with html/file_id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing the locked-down sandbox, strict CSP, blocked network calls, self-contained asset requirements, default-private access semantics, and how shared versus private URLs behave. It also documents the returned fields and clarifies that update preserves both links.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and URL behavior, then uses compact bullets for the mutually exclusive inputs and a short paragraph for sandbox and access semantics. Despite its length, every sentence conveys operational detail needed to call the tool correctly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter write tool with no output schema, the description is complete: it clarifies mutual exclusivity, access modes, sandbox constraints, template behavior, and return fields. Nothing an agent needs in order to invoke it correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the baseline is 3, but the description adds important cross-parameter semantics: exactly one of html/file_id/analytics_card_ids, live rendering for html, and template behavior for server-side refresh. It does not add meaning for every minor parameter (e.g., favicon, in_workspace), which are left entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: hosting a self-contained HTML page at a stable, default-private, shareable URL. It directly distinguishes itself from documents.create by noting the page is served live with JavaScript running, and names the main input modes (html, file_id, analytics_card_ids).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to pass exactly one of the three input modes and gives the condition for choosing analytics_card_ids ('Best way to give someone a shareable analytics report'). It also routes follow-up actions to named siblings: flip access with artifacts.set_access and republish with artifacts.update.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
artifacts_export_pdfAInspect
Snapshot an artifact's current version to a PDF saved in the workspace. NOTE: the PDF is a STATIC capture rendered with JavaScript disabled — interactive or JS-drawn content (charts on , script-generated SVG) will not appear. Author charts as static SVG if they must show in the PDF. Returns the file_id; deliver it with messages.send(attachments=[file_id]).
| Name | Required | Description | Default |
|---|---|---|---|
| artifact_id | Yes | The artifact to snapshot (its latest version). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and idempotentHint=false but explain nothing about rendering. The description adds genuinely important behavior: the PDF is a STATIC capture with JavaScript disabled, so canvas/script-drawn content is silently dropped, plus it discloses the return value (file_id) and the delivery path via messages.send.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the action, then a NOTE for the critical rendering caveat, then the return/delivery contract. Every sentence carries information an agent needs; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description covers the return value and how to use it downstream. Combined with annotations that cover the safety profile and a fully documented schema, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented, including that artifact_id targets the latest version. The description adds no syntax, format, or meaning beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource+target: 'Snapshot an artifact's current version to a PDF saved in the workspace.' An agent can immediately distinguish this from artifacts_create/artifacts_update/artifacts_refresh siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it (capturing a snapshot of the current version) and a workaround guideline ('Author charts as static SVG if they must show in the PDF'). It does not name an alternative tool or an explicit when-not condition, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
artifacts_getARead-onlyIdempotentInspect
Fetch an artifact's metadata by id or slug — title, access level, version, view count, URL. Set include_html=true to also return the latest HTML body (so you can read-then-edit before calling artifacts.update).
| Name | Required | Description | Default |
|---|---|---|---|
| slug | No | The artifact slug. Pass this or artifact_id. | |
| artifact_id | No | The artifact id. Pass this or slug. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| include_html | No | If true, include the latest version's HTML body. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false and closed-world, so safety is covered. The description goes further by disclosing the returned shape (title, access level, version, view count, URL) and the conditional HTML body payload, which annotations cannot express. It stops short of noting any rate limits or the response envelope, hence a 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact clauses: the core fetch and its key identifier first, then the one optional behavior and its rationale. No filler, no repetition of the tool name, and the actionable guidance sits at the end where the agent reading left-to-right reaches it after learning what the call returns.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly compensates by naming the metadata fields returned and flagging the opt-in HTML body. The in_workspace scope caveat lives only in the schema, but that schema text is thorough ('Nothing is stored; other sessions are not affected'), leaving no real gap for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description earns above baseline by explaining the intent behind include_html (read the current body so you can edit it) rather than just restating the schema's 'If true, include the latest version's HTML body'. It does not add anything further for in_workspace.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Fetch an artifact's metadata') and an explicit lookup key ('by id or slug'), then enumerates the returned fields. This clearly distinguishes it from artifacts_list (enumeration), artifacts_update (mutation) and artifacts_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Routes the agent explicitly: set include_html=true when you intend to 'read-then-edit before calling artifacts.update', naming the downstream sibling. The trigger condition for the optional flag is spelled out rather than left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
artifacts_listARead-onlyIdempotentInspect
List the workspace's artifacts (most-recently-updated first): title, url, access level, version, view count. Soft-deleted artifacts are excluded.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so safety is covered. The description adds real behavioral context beyond that: results are ordered most-recently-updated first, soft-deleted artifacts are hidden, and the returned fields are enumerated. It stops short of stating pagination or result-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the resource and scope front-loaded and zero filler; the sort-order parenthetical earns its place and the exclusion note is the last useful fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the returned fields (title, url, access level, version, view count) and the ordering, which is what an agent needs to use the result. Only pagination/volume limits are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single in_workspace parameter is fully documented in the schema. The description mentions no parameters at all and adds no syntax or scoping nuance, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear specific verb (List) plus resource (workspace's artifacts), and it specifies scope details: sort order (most-recently-updated first) and the fields returned. It does not explicitly distinguish itself from the nearest sibling artifacts_get, but the singular/plural naming and the returned-field list make the distinction inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the tool plainly exists to enumerate artifacts, yet there is no explicit 'use this when you need X instead of artifacts_get' routing and no prerequisites mentioned. No misleading guidance, just no when-to-use framing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
artifacts_refreshAInspect
Re-render a data-driven artifact from a small data payload and publish a new version at the SAME URL. The artifact must have been created/updated with a template (HTML containing {{placeholder}} tokens). Pass data as a flat map of placeholder -> value (e.g. {"leads": "3 200", "date": "15 июля 2026"}); the server substitutes them into the stored template — you do NOT send any HTML. Ideal for scheduled refreshes of live numbers. Every template placeholder must have a value in data, or the call fails.
| Name | Required | Description | Default |
|---|---|---|---|
| data | Yes | Flat map of {{placeholder}} name -> value. Every placeholder in the template must be present. Values are HTML-escaped. | |
| label | No | Optional version label (e.g. 'daily refresh'). | |
| artifact_id | Yes | The data-driven artifact to refresh. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false. The description adds meaningful behavior beyond that: a new version is published at the SAME URL, the server does the substitution (no HTML sent), and the call fails if any placeholder lacks a value. That failure condition and the URL-stability guarantee are useful context the annotations don't convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and kept to a handful of sentences. Slight redundancy: the 'every placeholder must have a value or the call fails' rule is already stated in the schema description, so one sentence partially repeats structured data.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does usefully explain the effect (a new version at the SAME URL) rather than a return payload. Combined with the template precondition, failure mode, and data format, an agent has what it needs to invoke it correctly, though the response shape is only implied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description nevertheless adds value: a concrete example of the `data` map ({"leads": "3 200", "date": "..."}), the rule that every placeholder must be present, and the explicit note that values are substituted server-side so no HTML is passed. This clarifies the semantics of `data` beyond the schema's field list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource pair ('Re-render a data-driven artifact') plus the key mechanism (template + data payload, published to the SAME URL). This is highly specific. It does not explicitly name a sibling (e.g. artifacts_update) to contrast against, so it falls just short of the 5 bar for differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use signal ('Ideal for scheduled refreshes of live numbers') and a precondition (artifact must have been created/updated with a template). It does not state when NOT to use it or name an alternative tool, so no exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
artifacts_set_accessAInspect
Change an artifact's access level. 'shared' makes the unguessable link the capability (anyone-with-link); 'private' revokes that — the link 404s for anyone not signed into the workspace. Existing shared links keep the same URL when re-shared.
| Name | Required | Description | Default |
|---|---|---|---|
| artifact_id | Yes | The artifact to change. | |
| access_level | No | 'shared' (anyone-with-link) or 'private' (in-app only). OMIT to leave the access level unchanged — do that when the call only binds or unbinds a Telegram Mini App. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| telegram_web_app_agent_id | No | The agent woken by events from this page. It must own exactly one enabled webhook trigger — that trigger is what the events target. Not needed when unbinding (telegram_web_app_account_id=0). | |
| telegram_web_app_account_id | No | Bind this artifact to a Telegram bot account (channel_account.id) so it can be opened as a Mini App from a `web_app` button and report verified events back. The artifact must already be shared (or be made shared in this same call). Pass together with telegram_web_app_agent_id. Pass 0 to UNBIND — the page stops being a Mini App and stops accepting events; telegram_web_app_agent_id is not needed to unbind. Setting access_level='private' also unbinds. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations covering the safety profile (readOnlyHint=false, destructiveHint=false), the description still adds real behavior: 'shared' makes the unguessable link the capability, 'private' causes the link to 404 for anyone not signed into the workspace, and re-sharing preserves the URL. It omits that 'private' also unbinds a Telegram Mini App, though the schema carries that detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action, then the two modes, then a re-share edge case. No filler, no restatement of the name or title, and each sentence adds distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is non-trivial (Telegram Mini App binding, workspace scoping, enum + omit semantics), but the 100%-covered input schema carries those dimensions and there is no output schema to explain. The description fully covers the access-level dimension it is named for; only the interaction with Telegram binding is left entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3, but the description genuinely enriches the enum beyond the schema's terse 'shared (anyone-with-link) or private (in-app only)': it explains that the link itself is the capability and that revocation produces a 404. The URL-stability note further clarifies re-invocation semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Change an artifact's access level') that is immediately distinguishable from the sibling artifacts_update (which handles content) and artifacts_get/list. The scope is tight enough that an agent knows exactly which operation this is without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the semantics of 'shared' vs 'private' rather than stated: an agent can infer when to pick each value, but the description never says when to call this tool versus artifacts_update, nor does it name alternatives or prerequisites. Adequate but with a clear gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
artifacts_updateAInspect
Republish an existing artifact with new HTML. The slug and URL stay the SAME; a new version is stored (older versions are retained up to a cap). Use this to refresh a shared dashboard — anyone with the link sees the new snapshot.
| Name | Required | Description | Default |
|---|---|---|---|
| html | Yes | The new full self-contained HTML page (same contract as artifacts.create). | |
| label | No | Optional human name for this version (e.g. 'Q3 final'). | |
| template | No | Optional: attach/replace the data-driven template (HTML with {{placeholder}} tokens) on this existing artifact, so later artifacts.refresh(data={...}) can re-render it server-side. Pass the html rendered from this template. Omit to leave the template unchanged. | |
| artifact_id | Yes | The artifact to republish. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the safety profile (not read-only, not idempotent, not destructive); the description adds meaningful behavior beyond that: slug/URL stability, new-version creation, a retention cap on older versions, and that link-holders see the new snapshot. That is real value the agent cannot get from the structured fields. It stops short of covering permissions or the return payload.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and its most important invariant (same URL), then the retention/visibility behavior. Every clause carries information; nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and full schema coverage, the description covers the behavioral consequences an agent needs (URL stability, versioning, retention, visibility). It omits auth requirements and what the call returns (e.g. new version id), which keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents artifact_id, html, label, template, and in_workspace. The description adds no per-parameter syntax or constraints beyond what the schema states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (republish an existing artifact with new HTML) and immediately discloses the key distinguishing behavior: slug and URL stay the same while a new version is stored. An agent can tell this apart from artifacts_create (new artifact) and artifacts_refresh (server-side re-render) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives one concrete use case ("refresh a shared dashboard"), which is implied guidance, but never names or contrasts with the obvious alternatives artifacts_create or artifacts_refresh. The template parameter hints at a relationship with artifacts_refresh, but the routing condition (replace HTML vs. re-render from data) is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
background_runAInspect
Run ANOTHER tool off the critical path: returns INSTANTLY while the target tool runs in the background, so a slow write/log/notify never delays your reply. Pass tool (the id of a tool you already have) and args (the target tool's arguments as a JSON object STRING). Optional serialize_key serializes background runs that share the key (e.g. the spreadsheet id when appending rows). Use ONLY for tools whose result you don't need this turn (logging, side-effect writes). Cannot run a tool you don't already have, cannot run a tool that requires approval, and cannot run a tool that deletes data — call those directly. Only callable from within an agent turn.
| Name | Required | Description | Default |
|---|---|---|---|
| args | Yes | The target tool's arguments as a JSON object STRING, e.g. '{"spreadsheet_id":"abc","values":[["a","b"]]}'. (A JSON string — not a bare object — so the schema stays valid for strict function-calling on voice.) | |
| tool | Yes | Id of the tool to run in the background. Must be a tool you already have. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| serialize_key | No | Optional. Background runs sharing this key run one-at-a-time (e.g. pass the spreadsheet id when appending rows). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnly=false, destructive=false, idempotent=false), but the description adds substantial context beyond them: instant return semantics, background execution model, the serialize_key serialization guarantee, and the approval/deletion restrictions. These are non-obvious behavioral traits that materially change how an agent invokes it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The most important constraint (instant background return) is front-loaded, and every sentence carries routing or constraint value. It is somewhat dense at five sentences, but there is little dead weight to cut.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a fire-and-forget meta-tool with no output schema, the description fully covers what happens (instant return, background execution), what is prohibited, and the serialization option. An agent has everything needed to call it correctly without guessing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters including the JSON-string rationale and serialize_key example. The description restates 'tool must be a tool you already have' and the serialize_key example rather than adding new meaning, so it sits at the schema-driven baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The lead sentence states a specific verb (run) and resource (another tool) plus the key behavioral trait (returns INSTANTLY). An agent can immediately distinguish this meta-tool from every direct sibling in the list, since it operates on other tools rather than a domain resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('Use ONLY for tools whose result you don't need this turn (logging, side-effect writes)') and when-not-to-use ('cannot run a tool that requires approval, and cannot run a tool that deletes data — call those directly'). This is a near-complete routing rule with the alternative action named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_add_init_scriptARead-onlyIdempotentInspect
Register JavaScript that runs BEFORE the page's own code, at document start, on every page opened afterwards in this identity's context. Use when browser.evaluate is too late — most often to hook window.fetch/XMLHttpRequest and capture a SPA's request bodies, which network_requests cannot show. Does NOT affect already-open pages: add the script first, then browser.open. Cannot be removed once added; it lives until the context is evicted or browser.close is called.
| Name | Required | Description | Default |
|---|---|---|---|
| script | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| identity_name | No | _anon |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=true, destructiveHint=false, idempotentHint=true), yet the description goes well beyond them: injection timing relative to page code, forward-only scope, and the fact that registration is irreversible and lives until context eviction or browser.close. That lifetime/irreversibility detail is exactly the kind of context the structured fields cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the action and timing, then the routing rule, then the two constraints (non-retroactive, non-removable). No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 1-required-param action tool with no output schema, the behavioral surface is well covered: timing, scope, irreversibility, and alternatives. The only shortfall is that no parameter-level semantics (script format, identity_name default) are supplied to offset the thin schema descriptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only in_workspace is documented), so the description is expected to compensate. It implies the meaning of the "script" argument (document-start JS) and the identity scoping via "this identity's context," but never names or explains the parameters, and says nothing about duplicate/`_anon` default behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ("Register JavaScript") plus its exact execution point ("runs BEFORE the page's own code, at document start") and scope ("every page opened afterwards in this identity's context"). It distinguishes itself from browser_evaluate and browser_network_requests by name, so an agent can route without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use with a named alternative and the failure mode that selects it: "Use when browser.evaluate is too late... which network_requests cannot show." It also gives an ordering rule ("add the script first, then browser.open") and an exclusion ("Does NOT affect already-open pages").
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_attach_identityARead-onlyIdempotentInspect
Switch the page's identity by loading saved cookies + storage. Use only when switching identity mid-page; for first navigation, pass identity_name to browser.open instead.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| identity_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description usefully adds the mechanism (cookies + storage are loaded). It does not say what happens to the page's existing session/identity data, whether the switch is reversible, or whether the target identity must already exist — which matters for a stateful page mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the operational constraint front-loaded ahead of the redirect to browser_open. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter tool with no output schema and annotations covering the safety profile, the description gives enough to call it correctly and to not confuse it with browser_open. The one remaining gap is what happens to prior page state when the identity is swapped.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 33% (only in_workspace is documented in the schema). The description implies identity_name refers to a saved identity by referencing saved cookies/storage, but adds no format, discovery, or failure semantics, and page_id is left entirely to the schema. Partially compensates for the coverage gap, no more.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — switching the page's identity — and explains the mechanism (loading saved cookies + storage). It also distinguishes itself from browser_open, which handles first navigation, so an agent can route correctly without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scoped: 'Use only when switching identity mid-page' tells the agent the trigger condition, and 'for first navigation, pass identity_name to browser.open instead' names the alternative and the condition that selects it. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_attach_meetARead-onlyIdempotentInspect
Attach to THIS meet's bot browser and return a page_id for its live Google Meet tab. Then use the normal browser.* tools (browser.snapshot to see the page, browser.click to press a control, browser.take_screenshot, browser.evaluate, browser.console_messages, browser.network_requests) with that page_id to inspect why the bot isn't visibly joining (lobby/blocked/admitted) and to recover a stalled join. Call with no arguments from inside a meet.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that: it attaches to a specific live tab, hands back a page_id, and enables a follow-up inspection/recovery workflow. It does not state whether the page_id expires or whether attaching is exclusive, which keeps it out of the top band.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and return value. The parenthetical tool list is dense but each item signals a concrete capability; slightly more enumeration than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by stating the return value (page_id) and how to use it. For a single-optional-param, read-only attach tool this is close to complete, missing only page_id lifetime/exclusivity details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single optional in_workspace param is fully documented in the schema and the baseline is 3. The description's 'call with no arguments' framing is consistent with that param being optional and adds no syntax beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (attach) and resource (THIS meet's bot browser) and declares the return value (page_id for the live Google Meet tab). It is clearly distinguished from the normal browser.* siblings, which it names as downstream tools rather than alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the triggering condition (to inspect why the bot isn't visibly joining: lobby/blocked/admitted, or to recover a stalled join) and the prerequisite (call with no arguments from inside a meet). It does not name a competing tool it should be preferred over, but the when-to-use context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_clickBRead-onlyIdempotentInspect
Click an element. ref is either an aria-ref token from browser.snapshot ('e7') OR a CSS selector ('button.submit'). Prefer the aria-ref token.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| page_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (readOnlyHint, idempotentHint, destructiveHint), so the description is not the primary burden. But it adds nothing behavioral: it doesn't say whether the click waits for navigation, that it may change page state, or what happens if the ref is stale. Notably, readOnlyHint/idempotentHint are questionable for a click action, yet the description offers no clarifying context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no waste, front-loading the core action before explaining the ref parameter. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a browser interaction tool with no output schema and one fully undocumented required parameter, the description is incomplete. It handles ref sourcing well but omits post-click behavior (navigation/waiting), error handling for stale refs, and any explanation of page_id.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 33%, and the description strongly compensates for `ref` by documenting both accepted formats (aria-ref token vs CSS selector), its origin in browser.snapshot, and the preferred form. However, the required `page_id` parameter is undocumented in both schema and description, leaving a real gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Click an element") that is clearly distinct from sibling verbs like browser_hover, browser_fill, and browser_type. However, it never explicitly differentiates from those siblings, so it lands just below the 5 bar.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use click versus alternatives (hover, type, fill, press_key) and no stated prerequisites. The only usage hint is "Prefer the aria-ref token," which is about parameter format choice, not tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_closeBRead-onlyIdempotentInspect
Close a page opened by browser.open.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds nothing behavioral beyond this – no note on what closing does to the session, auth requirements, or effects on other sessions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler, and the action is front-loaded. It is efficient, though its brevity comes partly at the cost of the details penalized in other dimensions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple close tool with annotations carrying the safety profile and no output schema, this is roughly adequate. It nonetheless omits the source of page_id and any consequence of closing (resource release, effect on later calls).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: page_id has a title but no description. The description hints that page_id corresponds to a page from browser_open but never names the parameter or explains where to obtain it, and gives no additional meaning for in_workspace beyond the schema's own text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (close a page) and ties it to the origin action browser_open. However, it does not distinguish this from sibling browser_tabs, which likely also manages page/tab lifecycle, so sibling differentiation is weak.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by 'a page opened by browser.open' – the agent can infer this is called on pages it previously opened. There is no explicit when/when-not guidance, no prerequisites, and no named alternative such as browser_tabs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_console_messagesARead-onlyIdempotentInspect
Return console.log/warn/error events captured since the last drain. Filter by level ('log'|'info'|'warning'|'error'|'debug') and/or pattern (regex). Buffer caps at 500 entries; oldest are dropped first. Set clear=false to peek without draining.
| Name | Required | Description | Default |
|---|---|---|---|
| clear | No | ||
| level | No | ||
| page_id | Yes | ||
| pattern | No | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is unusually informative about behavior (drains by default, buffer caps at 500 with oldest dropped first, clear=false peeks), which annotations alone would not convey. However, the disclosed default drain makes the call state-modifying and non-repeatable, conflicting with the idempotentHint=true and readOnlyHint=true / destructiveHint=false annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core action and scoping before filters and the drain/buffer caveats. Dense and largely waste-free, though the buffer-cap sentence could sit closer to the drain caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should ideally sketch the returned event shape (level, text, timestamp) or ordering, and note that the buffer is per page/session. It covers buffer behavior well but leaves the return contract and page prerequisite implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 20% (only in_workspace is documented), so the description must compensate. It supplies the level enum values (absent from the schema), the regex nature of pattern, and the default-true/peek semantics of clear, covering the ambiguous parameters well; only page_id is left purely implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Return), resource (console.log/warn/error events) and scope (captured since the last drain), which clearly separates it from browser siblings such as browser_network_requests. An agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage context: how to filter (by level and/or regex pattern) and the clear=false option to peek without consuming the buffer. It does not name alternatives or prerequisites (e.g. that an open page via browser_open is required), so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_dragARead-onlyIdempotentInspect
Drag one element onto another. source_ref is the element to grab; target_ref is where to drop. Both are CSS selectors. Used for slider captchas, kanban, drag-and-drop uploads.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| source_ref | Yes | ||
| target_ref | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered. The description adds nothing behavioral beyond that - it does not say whether elements must be visible/scrolled into view, how the drag is stepped, or whether it can time out. Note the mild tension: a tool used for kanban is not obviously read-only/idempotent, but the description makes no explicit contradicting claim.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the action, then operand clarification, then use cases. Every sentence adds distinct information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists and none is needed for a UI action; the description covers the action, both operands and typical scenarios. It leaves page_id context and any wait/visibility prerequisites unaddressed, so it is strong but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 25%, so the description must carry the load, and it does for the two decisive params: source_ref is 'the element to grab', target_ref is 'where to drop', both CSS selectors. page_id is left unexplained, keeping it short of a 5 given the low schema support.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Drag one element onto another') and immediately identifies the two operands, so the action is unambiguous. It implicitly separates itself from browser_click/browser_hover by being a drag, but never names or distinguishes a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The final clause gives concrete use cases ('slider captchas, kanban, drag-and-drop uploads') that tell the agent when this tool applies rather than a plain click. It offers no exclusions or named alternatives, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_evaluateARead-onlyIdempotentInspect
Run JavaScript in the page context and return the result. Use for state not in the a11y tree, captcha iframe inspection, DOM events. Expression is either a plain JS value ('document.title') or a zero-arg IIFE ('(() => { … })()'). Inline any runtime values into the expression itself. Result is JSON-serialized; non-serializable values become strings. 256KB cap on output.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| expression | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (which only assert read-only/idempotent safety), the description discloses real behavioral traits: the result is JSON-serialized, non-serializable values are coerced to strings, and output is capped at 256KB. That is meaningful invocation-relevant context. It does not address async/promise handling or execution-time failures, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: purpose first, then use cases, then expression format, then result serialization and the size cap. No repetition of schema or annotation content and no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries return-value semantics itself, covering serialization and truncation limits. Combined with the expression-format rules this is nearly complete; only execution semantics (awaited promises, timeouts, errors) are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is low (33%), so the description must carry the load, and it does for the critical 'expression' parameter: it defines accepted shapes (plain JS value vs. zero-arg IIFE), gives literal examples, and warns to inline runtime values. 'page_id' is left undocumented, so the compensation is strong but incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run JavaScript in the page context and return the result') and immediately scopes it with use cases (state not in the a11y tree, captcha iframe inspection, DOM events) that separate it from browser_snapshot and the click/type family. An agent can distinguish it from siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear positive routing ('Use for state not in the a11y tree...'), which implicitly excludes cases the accessibility snapshot already covers. No explicit 'do not use when' clause or named alternative tool is given, so it stops just short of a full when/when-not rule set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_file_uploadARead-onlyIdempotentInspect
Attach files to an . Pass either local_paths (absolute host paths) or data (list of {name, mime, base64} blobs written to /tmp). 25MB cap per file.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| data | No | ||
| page_id | Yes | ||
| local_paths | No | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations declare readOnlyHint=true, but the description states the tool attaches files into a page's file input and writes blobs to /tmp — both environment mutations, so the description's disclosed behavior contradicts the read-only annotation. This is a real inconsistency an agent could be misled by, even though the 25MB cap and /tmp destination are otherwise useful behavioral detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: purpose first, then the parameter-choice rule, then the hard constraint. No filler, and each sentence adds information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter browser tool with no output schema, the description covers mechanism, both input modes, and the size limit. It stops short of explaining how `page_id`/`ref` are obtained or what happens on the page after attachment, but no return-value documentation is required since no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% (just `in_workspace`), so the description must carry the load and largely does: it explains that `local_paths` are absolute host paths and that `data` items are {name, mime, base64} blobs staged in /tmp, plus a size cap absent from the schema. It adds nothing for `ref` or `page_id`, which remain unexplained in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Attach files') and an exact target ('an <input type=file>'), which cleanly separates it from storage-oriented siblings such as files_upload, collections_add_file, and agents_add_file. The mechanism (browser DOM file input) is unambiguous from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the mention of a file input; there is no explicit when-to-use/when-not and no alternative is named. It does explain the two mutually exclusive input modes and the 25MB cap, but omits prerequisites such as needing an open page and a `ref` obtained from browser_snapshot.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_fillARead-onlyIdempotentInspect
Fill an input or textarea with the given value. ref is either an aria-ref token from browser.snapshot ('e7') OR a CSS selector ('input[name=email]'). Prefer the aria-ref token — it's stable and matches exactly what snapshot returned.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| value | Yes | ||
| page_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=true, idempotent=true, destructive=false and openWorld=false, so the safety profile is covered. The description adds useful evidence about ref stability and snapshot matching, but omits key fill behavior: whether an existing value is replaced or appended, and what happens on an invalid/stale ref.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the action front-loaded, followed by parameter disambiguation and a concrete recommendation. No filler, though the final preference clause slightly restates the format point already made.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 4-parameter tool with no output schema and thin schema descriptions, the description covers the ambiguous parameter (ref) and the action clearly. Remaining gaps -- page_id meaning and value-replacement semantics -- are minor but real.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (only in_workspace is documented), so the description must compensate, and it does well for the hardest parameter: ref is explained as either an aria-ref token ('e7') or a CSS selector ('input[name=email]'). value is self-evident and page_id is left unexplained, keeping this below a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: "Fill an input or textarea with the given value." An agent immediately knows this writes text into a form field. However, it never distinguishes itself from the very similar siblings browser_type and browser_fill_form, so the differentiation half of the criterion is unmet.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a real decision rule, but only for parameter format ("Prefer the aria-ref token — it's stable"), not for tool selection. There is no guidance on when to choose browser_fill over browser_type or browser_fill_form, so usage is only implied by the stated purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_fill_formARead-onlyIdempotentInspect
Fill multiple form fields in one call. fields is a list of {ref, value} dicts. ref is a CSS selector; value is a string (text) or boolean (checkbox). Saves N round-trips vs calling browser.fill repeatedly.
| Name | Required | Description | Default |
|---|---|---|---|
| fields | Yes | ||
| page_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the concrete field/ref/value shape and the round-trip rationale, but says nothing about event triggering, clearing existing values, or whether the form is submitted—extra context that would be valuable for a form-filling tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action, followed by the parameter shape, then the usage rationale. No filler; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a browser-automation tool with rich annotations and no output schema, the description covers the non-obvious nested parameter and the efficiency rationale. It omits return behavior and any waiting/event semantics, but the annotations carry the safety profile, so the remaining gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, and the nested `fields` items are opaque ({type: object, additionalProperties: true}), so the description's explanation that fields is a list of {ref, value} dicts with ref as a CSS selector and value as string/boolean is essential and compensates well. It does not address page_id, but that parameter is self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fill multiple form fields in one call') and immediately distinguishes itself from the sibling browser_fill by emphasizing the batch behavior. An agent can tell it apart from browser.fill and browser.type without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Saves N round-trips vs calling browser.fill repeatedly' names the alternative and implies the condition that favors this tool (multiple fields at once). It gives clear context for preferring this tool but stops short of explicit when-not guidance or mentioning siblings like browser_select_option for dropdowns.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_handle_dialogARead-onlyIdempotentInspect
Respond to a pending JS dialog (alert/confirm/prompt). Pass accept=true for OK or false for Cancel. For prompt() dialogs also pass prompt_text. Dialogs are queued at page-open time; returns {pending: false} if none is waiting.
| Name | Required | Description | Default |
|---|---|---|---|
| accept | Yes | ||
| page_id | Yes | ||
| prompt_text | No | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so safety is covered. The description adds genuinely new behavioral context: dialogs are queued at page-open time and the call returns {pending: false} when nothing is waiting, which tells the agent how to interpret a no-op result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, purpose front-loaded, then parameter semantics, then runtime behavior. No filler and every sentence adds actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and low schema description coverage, the description steps in to explain return shape and prompt handling, which is adequate for a 4-param tool. Only page_id and error cases are left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 25% (just in_workspace), so the description carries most of the burden and does so well: accept is mapped to OK/Cancel semantics and prompt_text is tied to the prompt() case. page_id is left implicit but is self-evident from the required field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Respond) and resource (pending JS dialog) plus the dialog types (alert/confirm/prompt). No other browser_* sibling handles dialogs, so it is unmistakably distinct from browser_click, browser_snapshot, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly identifies the triggering condition (a pending dialog) and the response options (accept=true for OK, false for Cancel, prompt_text for prompt()). It does not name alternatives or when-not-to-use, but for a uniquely-scoped tool that gap is minor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_hoverARead-onlyIdempotentInspect
Hover the mouse over an element (reveals tooltips + hover menus). ref is a CSS selector.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| page_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare that the operation is read-only, idempotent, non-destructive, and not open-world. The description adds the visual effect of hovering, but it does not explain whether the hover state persists, what page state is required, or what the call returns. With annotations covering safety, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, front-loads the action and its effect, and adds the key parameter semantics without any filler. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple browser action with rich annotations and no output schema, the description covers purpose, effect, and ref semantics. However, it leaves the required page_id parameter undocumented and does not mention preconditions such as an open browser page, so a meaningful gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 33%: page_id and ref are required but have no schema descriptions. The description explains ref as a CSS selector, but leaves page_id completely unexplained, so it only partially compensates for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Hover) and resource (mouse over an element) and immediately clarifies the observable outcome: revealing tooltips and hover menus. This distinguishes it from sibling browser actions such as browser_click, browser_type, and browser_drag.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical 'reveals tooltips + hover menus' gives clear context for when the tool is useful, namely interacting with hover-dependent UI. It does not name alternatives or exclusions, such as when to use browser_click instead, but the usage context is easy to infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_network_requestsARead-onlyIdempotentInspect
List HTTP requests the page made since open or last drain. Optional filters: method (GET/POST/...), url_pattern (regex), status_min (e.g. 400 for errors). Captures up to 200 most recent requests per page.
| Name | Required | Description | Default |
|---|---|---|---|
| clear | No | ||
| method | No | ||
| page_id | Yes | ||
| status_min | No | ||
| url_pattern | No | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive, so the safety profile is covered. The description adds genuinely new behavior: the collection window ('since open or last drain') and a hard capacity limit ('up to 200 most recent requests per page'), which are not derivable from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core purpose, then the filters, then the capacity caveat. Every sentence carries information and none is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with annotations covering safety and no output schema, the description supplies the window semantics and capacity limit an agent needs. The only notable omission is an explicit explanation of the 'clear' parameter's drain behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 17%, so the description must compensate and partially does: it documents method, url_pattern (regex), and status_min (e.g. 400) with useful format hints. However, the 'clear' parameter is never explained (only obliquely hinted at by 'last drain'), leaving a real gap at low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (HTTP requests the page made), plus the collection window ('since open or last drain'). It clearly distinguishes this from sibling tools like browser_console_messages by resource, though it never names an alternative directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no alternatives named. The filters imply a debugging/inspection context but the agent gets no explicit signal about when this tool is the right choice versus browser_console_messages or browser_snapshot.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_openARead-onlyIdempotentInspect
Open a URL in a remote browser. Saved login cookies are auto-attached when the URL domain matches a claimed browser identity. Pass identity_name to override auto-matching or force a specific identity.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| identity_name | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/closed-world, and the description adds genuinely new behavioral context: login cookies auto-attach on matching domains and identity_name can override or force an identity. It does not contradict the readOnlyHint, so no flag. Minor gap: no mention of failure/timeout behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, followed by two tight sentences that each add distinct information (auto-attach rule, override mechanism). No padding or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool whose safety profile is covered by annotations and which has no output schema, the description supplies the key operational context an agent needs (cookie attachment and identity override). It is close to complete, missing only edge-case behavior such as navigation failures.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must carry weight — and it does for identity_name, explaining that it overrides auto-matching or forces a specific identity, which is beyond the bare string|null type. in_workspace is not addressed in the description, but the schema already documents it, and url is self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Open a URL in a remote browser') so the agent knows exactly what the tool does. It doesn't explicitly differentiate itself from siblings like browser_navigate_back or browser_attach_identity, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the automatic cookie/identity behavior and mentions identity_name as an override, which implies when to reach for it. However, it gives no explicit when-to-use versus alternatives (e.g., browser_navigate_back or browser_attach_identity) and no preconditions for choosing an identity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_press_keyARead-onlyIdempotentInspect
Press a keyboard key (e.g., 'Enter', 'Tab', 'Escape', 'ArrowDown') or a single character. Optional ref focuses an element first — aria-ref token from browser.snapshot ('e7') or a CSS selector.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| ref | No | ||
| page_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description usefully explains that ref focuses an element first, but says nothing about error behavior (e.g., unresolvable ref) or side effects of the key press.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences that front-load the core action and then the optional focusing behavior. No filler; each clause adds actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a moderate-complexity input tool with no output schema, the description covers the action and the ref mechanism adequately, but omits return/effect details and the workflow relationship to browser_snapshot beyond the token reference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 25% schema coverage, the description compensates well by explaining the two core params: key accepts named keys or a single character, and ref accepts an aria-ref token ('e7') or a CSS selector. page_id and in_workspace are left to schema (in_workspace documented there).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (press) and resource (keyboard key) with concrete examples ('Enter', 'Tab', 'Escape', 'ArrowDown') or a single character. It is clear what the tool does, though it does not explicitly distinguish itself from nearby siblings like browser_type or android_key.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through the reference to 'aria-ref token from browser.snapshot', suggesting a snapshot-then-press workflow, but there is no explicit when-to-use guidance and no stated alternatives (e.g., browser_type for text entry).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_resizeARead-onlyIdempotentInspect
Resize the page viewport. Useful when a site serves different HTML based on viewport width (mobile vs desktop) or when an anti-bot scores risk by viewport dimensions.
| Name | Required | Description | Default |
|---|---|---|---|
| width | Yes | ||
| height | Yes | ||
| page_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety and repeatability profile is covered. The description adds useful motivation for the operation but does not disclose traits beyond annotations, such as whether the viewport change persists across navigations or applies to the session's active tab.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and the rationale immediately after. Every clause earns its place and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple no-output-schema tool, the description adequately explains what it does and why, so return values need not be covered. The gap is parameter-level: three of four inputs are undocumented in both schema and description, which leaves an agent guessing at expected values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (only in_workspace is documented). The description mentions no parameters at all, leaving page_id, width and height without units or format guidance (e.g., CSS pixels, which tab is targeted). With low coverage the description should compensate, and it does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Resize the page viewport'), which is unambiguous about the operation. It does not explicitly contrast itself with sibling browser tools such as browser_open or browser_take_screenshot, but the action is distinct enough to be identified without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete conditions for use: sites that serve different HTML by viewport width and anti-bot risk scoring by viewport dimensions. That is genuine when-to-use guidance, though it names no alternative tool or exclusion, so it falls short of the 5 level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_select_optionARead-onlyIdempotentInspect
Pick option(s) in a native dropdown. Pass value (matches the option's value attr) OR label (matches its visible text). Lists allowed for multi-select.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| label | No | ||
| value | No | ||
| page_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint, so the safety profile is covered. The description adds matching semantics and multi-select list support, but says nothing about what happens when an option is absent, nor that selecting fires change events (which is relevant given readOnlyHint=true).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, and the core action is front-loaded ahead of the parameter guidance. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and low schema coverage, the description covers option-matching but omits how the required ref/page_id are obtained (presumably from browser_snapshot) and what a failure looks like. Adequate for the action itself, incomplete for a 5-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 20% (only in_workspace is documented), so the description must compensate. It does define two of the four undocumented params well ('value matches the value attr' vs 'label matches visible text', lists for multi-select), but ref and page_id remain unexplained in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Pick option(s) in a native <select> dropdown') and the qualifier 'native' implicitly separates it from custom-dropdown interactions that would use browser_click. It doesn't explicitly name a sibling, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'native <select>' qualifier implies when this tool applies (not for custom dropdowns), and it explains whether to pass `value` or `label`. But there is no explicit when-to-use / when-not guidance against browser_click or browser_fill, so the routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_snapshotARead-onlyIdempotentInspect
Return a YAML aria_snapshot of the page DOM. Each interactive node is tagged with [ref=eN] (e.g. [ref=e7]). Pass that exact token as the ref arg to browser.click / browser.fill / browser.type / browser.press_key. Do NOT pass the role name ('combobox', 'button') as ref — only the eN token. Truncated at 32KB.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real behavioral context not in the annotations: output is YAML, refs are embedded inline, and output is truncated at 32KB. It does not say what happens after truncation or whether refs go stale after navigation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the return value, then the ref contract, then the caveat. Four short sentences, all relevant. The explicit 'Do NOT pass the role name' warning slightly overlaps the preceding sentence but is defensible given how commonly that mistake is made.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the burden of describing the return value (YAML aria tree, eN ref tags, 32KB truncation) and its intended consumer. Remaining gaps — page_id semantics, ref lifetime across navigations, truncation recovery — are minor for a read-only snapshot tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% — `in_workspace` is documented in the schema while `page_id` (required) is not. The description adds no meaning for either input parameter; its discussion of `ref` concerns sibling tools' arguments, not this tool's schema. With the coverage gap uncompensated, this sits below the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Return a YAML aria_snapshot of the page DOM') and defines the distinguishing output artifact, the `[ref=eN]` tokens. An agent can tell this apart from browser_evaluate or browser_console_messages without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly explains the downstream workflow — the eN token is the value to pass to browser.click/fill/type/press_key — and warns against the common misuse of passing a role name. It gives no explicit 'when not to use this' or alternative (e.g. browser_evaluate for non-interactive inspection), so it stops short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_tabsBRead-onlyIdempotentInspect
Manage tabs within the same BrowserContext as page_id. action ∈ {list, switch, close, new}. For list, returns all open tab metadata; for new, returns the new tab's page_id.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| action | Yes | ||
| tab_id | No | ||
| page_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description advertises `close` (terminates a tab) and `new` (creates a tab), which are state-mutating operations, while the annotations declare readOnlyHint=true and destructiveHint=false. This directly contradicts the structured safety profile, which is the exact scenario the rubric treats as a contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with scope, then the action set and return behavior. No wasted words, though the return-value sentence could be deferred and the action set could be formatted more scanably.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a multi-action tab tool with no output schema, the description covers the key actions and even sketches return values for `list` and `new`, which is helpful. However, it omits the semantics of the optional `url`/`tab_id`/`in_workspace` parameters and never addresses the read-only framing that its own annotations assert.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 20% schema description coverage, the description must carry the load. It valuably enumerates the `action` values (absent from the schema, which types `action` only as a bare string) and explains the role of `page_id`, but leaves `url`, `tab_id`, and `in_workspace` semantics unexplained, so it only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('Manage tabs within the same BrowserContext as page_id') and enumerates the supported actions, so an agent knows exactly what the tool does. It does not explicitly differentiate itself from close siblings like browser_open or browser_close, keeping it just below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Listing `action` ∈ {list, switch, close, new} implies the usage contexts, but there is no explicit guidance on when to choose this tool over browser_open, browser_close, or present_tab, nor on which action applies in which situation. Usage is inferable but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_take_screenshotARead-onlyIdempotentInspect
Capture a PNG screenshot of the page or a specific element. Returns base64-encoded image bytes AND a file_id (persisted in DialogBrain files storage). Pass file_id straight to messages.send(attachment_file_ids=[file_id]) — do NOT call files.upload again. Use sparingly — favor browser.snapshot for structured DOM understanding.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | ||
| page_id | Yes | ||
| full_page | No | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| inline_bytes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), and the description adds substantial context beyond that: the return shape (base64 bytes AND a persisted file_id), where the file lives (DialogBrain files storage), and the anti-pattern to avoid (re-calling files.upload). That is genuinely useful behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then output shape, then workflow, then the usage constraint. Three tight sentences with no filler; every clause carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description rightly explains return values and persistence, and annotations cover safety. The main gap is parameter coverage — the effect of inline_bytes and full_page is left unstated, which matters for an agent deciding how to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20%, so the description should compensate but does not. It never explains ref, page_id, full_page, or inline_bytes — notably inline_bytes, which gates whether the advertised 'base64-encoded image bytes' are even returned. The one parameter with schema help (in_workspace) is untouched by the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Capture a PNG screenshot') and scope ('of the page or a specific element'). It explicitly differentiates from the sibling browser_snapshot, so an agent can pick between them without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-not and alternative: 'Use sparingly — favor browser.snapshot for structured DOM understanding.' It also routes the downstream workflow (pass file_id to messages.send, do not re-upload), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_typeARead-onlyIdempotentInspect
Type text into an element with per-keystroke delay (organic). Each character dispatches keydown/keypress/keyup, unlike browser.fill which replaces .value instantly. Use when the page listens to keystroke events or for typing-speed fingerprint checks. ref is an aria-ref token from browser.snapshot ('e7') or a CSS selector. delay_ms defaults to 50.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| text | Yes | ||
| page_id | Yes | ||
| delay_ms | No | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/destructive/idempotent, so the safety profile is covered; the description adds real behavioral detail beyond them (event-by-event dispatch, organic timing, delay default). It stops short of explaining error behavior (element/ref not found) or whether pre-existing element text is cleared first, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, purpose front-loaded, then differentiator, then when-to-use, then parameter clarifications. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-param tool with no output schema and low schema coverage, the description supplies purpose, mechanism, usage trigger and the key parameter semantics. The one material gap is whether typing appends to or replaces existing element content, which the browser.fill contrast hints at but never states.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 20% (just in_workspace), so the description carries the burden and does add meaning: `ref` is an aria-ref token from browser.snapshot or a CSS selector, and `delay_ms` defaults to 50. The obvious params (page_id, text) are left implicit, but the non-obvious ones are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (type text into an element) plus the distinguishing mechanism (per-keystroke delay, keydown/keypress/keyup dispatch). It explicitly contrasts itself with the sibling browser.fill, so an agent can pick between them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule ('page listens to keystroke events or typing-speed fingerprint checks') and names the alternative behavior (browser.fill replaces .value instantly), which implies the condition under which fill is preferred. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_wait_forARead-onlyIdempotentInspect
Wait for a selector to appear OR a navigation URL to match a glob pattern. Provide ref (selector) OR url_pattern (glob).
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | ||
| page_id | Yes | ||
| timeout_ms | No | ||
| url_pattern | No | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is well covered. The description adds the two wait modes but says nothing about timeout behavior, what happens on timeout expiry, or blocking semantics beyond what the schema/annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences, front-loading the purpose and then the required parameter choice. No filler and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with low schema coverage and no output schema, the description covers the core branch logic but omits the default timeout behavior (10000ms only implicit in schema) and the role of page_id. Adequate but incomplete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% (only `in_workspace` is documented), so the description must compensate. It usefully clarifies that `ref` means a selector and `url_pattern` means a glob, but leaves `page_id` and `timeout_ms` unexplained, so the gap is only partially filled.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Wait') plus the two exact conditions it waits on (selector appearance OR navigation URL matching a glob). This is a distinct capability within the browser_* family and an agent can identify it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the caller to provide `ref` OR `url_pattern`, which is mode selection, but gives no guidance on when to use this tool versus alternatives (e.g., waiting before a click vs calling browser_snapshot) or any preconditions. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_check_availabilityARead-onlyIdempotentInspect
Check when you have free time in Google Calendar. Shows busy periods and free slots in a given time range. Useful for finding meeting times or checking schedule conflicts.
| Name | Required | Description | Default |
|---|---|---|---|
| end_time | No | End date/time to check availability (YYYY-MM-DD or ISO 8601). Defaults to end of start_time day, or 7 days from now. | |
| timezone | No | IANA timezone (e.g. 'Asia/Bishkek') that start_time/end_time WITHOUT an offset are in, and that the free slots are reported in. OMIT for UTC. | |
| start_time | No | Start date/time to check availability (YYYY-MM-DD or ISO 8601). Defaults to start of today. | |
| calendar_id | No | Calendar ID to check. OMIT to use the agent's configured default calendar (or primary). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| working_hours_only | No | If true, only show free slots during working hours (9 AM - 6 PM). OMIT to show all free time (the default). | |
| min_duration_minutes | No | Minimum duration in minutes for free slots. Filters out short gaps. Default: 30 minutes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds what the result contains ('busy periods and free slots'), but says nothing about auth requirements, rate limits, or pagination. With annotation coverage, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action. Slight redundancy between 'check when you have free time' and 'shows busy periods and free slots' keeps it from a 5, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only query tool with seven fully documented optional parameters and no output schema, the description covers purpose, output nature, and use case adequately. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with defaults, timezone format, and omission semantics all documented in the schema itself. The description adds no parameter-level meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource ('Check ... free time in Google Calendar') and clarifies the return shape ('busy periods and free slots'), which distinguishes it from calendar_list_events. It stops short of naming a sibling tool explicitly, so it loses the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Useful for finding meeting times or checking schedule conflicts' gives implicit usage context, but there is no explicit when-to-use vs. when-not and no routing to alternatives such as calendar_list_events or calendar_create_event.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_create_eventAInspect
Create a new event in Google Calendar. Specify the title, start time, end time, and optionally invite attendees. Use ISO 8601 format for dates (e.g., 2024-12-15T14:00:00).
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | Event end time in ISO 8601 format. If not provided, defaults to 1 hour after start. Also accepts 'end_time' as alias. | |
| start | No | Event start time in ISO 8601 format (e.g., 2024-12-15T14:00:00). Also accepts 'start_time' as alias. | |
| title | No | Alias for summary - event title. | |
| summary | No | Event title/summary. Required. Also accepts 'title' as alias. | |
| end_time | No | Alias for end - event end time. | |
| location | No | Event location (physical address or virtual meeting link). | |
| timezone | No | Timezone for the event (e.g., 'America/New_York', 'UTC'). | |
| attendees | No | List of attendee email addresses to invite. | |
| start_time | No | Alias for start - event start time in ISO 8601 format. | |
| calendar_id | No | Calendar ID to create the event in. OMIT to use the agent's configured default calendar (or primary). | |
| description | No | Event description/notes. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| add_google_meet | No | If true, automatically creates a Google Meet link for the event. OMIT to skip Meet link. | |
| conference_data | No | Conference data for Google Meet. Alternative to add_google_meet flag. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-read-only, non-destructive, non-idempotent, open-world behavior, so the safety profile is covered. The description adds only the ISO 8601 input convention; it never says whether invitees actually receive notifications, whether the event is immediately live on shared calendars, or what side effects creation has.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, purpose front-loaded, followed by the required inputs and one concrete format example. No filler; it could be stronger only by ordering inputs to match required vs optional usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 14 mostly aliased parameters, nested conference_data, no output schema, and full annotation coverage, the description covers the core invocation path adequately. It omits what the call returns (created event identifier/link) and does not reconcile the stated 'required' title with an empty required list, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so every parameter including the alias fields (title/summary, start/start_time, end/end_time, calendar_id, add_google_meet) is already documented. The description restates title/start/end and an ISO 8601 example but adds no meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Explicit verb+resource: 'Create a new event in Google Calendar.' The create verb cleanly separates it from calendar_update_event, calendar_delete_event, and calendar_list_events. It does not, however, name any sibling or scope constraints explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied through the field checklist ('Specify the title, start time, end time, and optionally invite attendees') and format guidance. There is no statement of when to use this instead of calendar_update_event or calendar_check_availability, and no prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_delete_eventBDestructiveIdempotentInspect
Delete an event from Google Calendar. This action cannot be undone. Use with caution.
| Name | Required | Description | Default |
|---|---|---|---|
| event_id | Yes | ID of the event to delete. Required. | |
| calendar_id | No | Calendar ID containing the event. OMIT to use the agent's configured default calendar (or primary). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| send_notifications | No | Whether to send cancellation notifications to attendees. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered structurally. 'This action cannot be undone' reinforces irreversibility in plain language, but the description says nothing about notification side effects, permissions, or recurring-event scope, so it adds only modest value beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the operation, and no wasted preamble. 'Use with caution' is largely redundant with the preceding irreversibility warning and the destructiveHint annotation, costing a point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple destructive call whose annotations carry the safety profile and whose schema is fully documented. However, it omits meaningful gaps for a calendar deletion tool — behavior on recurring events, whether attendee notifications are suppressed, and confirmation of success (no output schema exists).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (event_id, calendar_id, in_workspace, send_notifications) are already fully documented in the schema. The description adds no parameter-level meaning, which is acceptable at this coverage level but earns only the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete an event from Google Calendar'), which cleanly separates it from the sibling calendar_create_event, calendar_update_event, calendar_list_events and calendar_check_availability. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use with caution' is a warning, not usage guidance — it does not say when to delete vs. update, what prerequisites exist, or what to do for recurring events. No alternatives or exclusions are named, so the agent gets no routing help beyond the obvious verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_list_eventsARead-onlyIdempotentInspect
List events from Google Calendar. Shows upcoming events by default. Can filter by date range and search query.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Free text search query to filter events. | |
| date_to | No | End date/time to query (YYYY-MM-DD or ISO 8601 format). Defaults to 7 days from now. Alias: time_max. | |
| date_from | No | Start date/time to query (YYYY-MM-DD or ISO 8601 format). Defaults to now. Alias: time_min. | |
| calendar_id | No | Calendar ID to list events from. OMIT to use the agent's configured default calendar (or primary). | |
| max_results | No | Maximum number of events to return. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, non-openWorld behavior, so the safety profile is covered. The description adds genuinely useful context beyond that, namely the default scope (upcoming events) and default date window, which is exactly the kind of default-behavior disclosure that helps an agent invoke it correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with zero filler; the primary purpose leads and the defaults follow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations carrying the safety profile and a fully-described schema, the description supplies the default-scope behavior an agent needs. No output schema exists, but for a simple list tool the coverage is essentially complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (query, date_from, date_to, calendar_id, max_results, in_workspace) is already documented in the schema. The description only repeats the date-range and query filtering at a high level, adding little beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List events from Google Calendar') that clearly separates it from the create/update/delete siblings. It doesn't explicitly name or differentiate from calendar_check_availability, which is the closest ambiguous sibling, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The note that it 'shows upcoming events by default' and can be filtered gives implied usage context, but there is no explicit when-to-use versus calendar_check_availability or any exclusions stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_update_eventAInspect
Update an existing event in Google Calendar. Can modify title, time, location, description, and attendees. Only specified fields will be updated.
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | New end time in ISO 8601 format. Optional. | |
| start | No | New start time in ISO 8601 format. Optional. | |
| summary | No | New event title/summary. Optional. | |
| event_id | Yes | ID of the event to update. Required. | |
| location | No | New event location. Optional. | |
| attendees | No | New list of attendee emails. Replaces existing attendees. | |
| calendar_id | No | Calendar ID containing the event. OMIT to use the agent's configured default calendar (or primary). | |
| description | No | New event description. Optional. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, covering the safety profile. The description adds one genuinely useful behavioral fact beyond the annotations: 'Only specified fields will be updated' (partial-update semantics). It does not mention attendee notification, recurring-event behavior, or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and resource, then scope, then the partial-update guarantee. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter mutation tool with full schema coverage and no output schema, the description covers purpose, scope, and partial-update behavior adequately. Minor gaps remain around side effects such as attendee notifications and recurring-event handling, but nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter including calendar_id fallback and in_workspace semantics. The description echoes the field list (title, time, location, description, attendees) without adding format or default details, and it omits calendar_id and in_workspace entirely. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Update an existing event in Google Calendar') and enumerates the mutable fields. It is clearly distinguishable from calendar_create_event, calendar_delete_event, and calendar_list_events by the verb alone, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: use this to modify an existing event that was already created via calendar_create_event. There is no explicit when-not guidance (e.g. do not use for creating events, or for deleting), so routing relies on the reader inferring from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calls_agent_duelAInspect
Start an agent-vs-agent VOICE test call: two AI voice agents share one LiveKit room — a 'caller' persona agent pursues a task brief against the 'callee' business agent under test. Use to evaluate booking flows, latency, and conversation quality without a human caller.
The CALLER agent should have an EMPTY voice_greeting (it must stay silent until the callee greets) and voice_filler_enabled=false.
Afterwards inspect both call_ids with agents.traces_list / calls.get_transcript. A subscribe-only listen token is returned for listening in live.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | The caller's brief — objective, persona details (name, phone), and when to end the call. Woven into its prompt as call instructions. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| max_duration_s | No | Hard cap on the call in seconds (30-900, default 300). | |
| callee_agent_id | Yes | Agent under test (answers and greets first). From agents.list. | |
| caller_agent_id | Yes | Customer-persona agent that places the call. Must be a different agent, active, with empty voice_greeting. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly=false, destructive=false, idempotent=false). The description adds real operational context the annotations don't carry: the caller must be a different, active agent with empty voice_greeting and voice_filler_enabled=false (it stays silent until greeted), plus a subscribe-only listen token for live monitoring.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then a second paragraph of prerequisites and follow-up. Dense but every sentence earns its place; minor tightening possible in the caller-config sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the full workflow: pre-call caller configuration, live monitoring via the returned listen token, and post-call inspection of both call_ids. With no output schema, the mention of returned tokens and call_ids is valuable; a bit more on what a completed duel produces would complete it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces the caller-agent constraints (empty greeting, filler disabled) but adds little syntactic meaning beyond the already-documented parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Start an agent-vs-agent VOICE test call' with the two roles (caller persona vs callee business agent under test) and the LiveKit room mechanism spelled out. This clearly distinguishes it from single-agent siblings like calls_make and calls_dispatch_agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use to evaluate booking flows, latency, and conversation quality without a human caller' gives a clear when-to-use condition. It also names the follow-up tools (agents.traces_list / calls.get_transcript) but stops short of explicitly contrasting with the nearest siblings (calls_make, agents_simulate_inbound).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calls_dispatch_agentARead-onlyIdempotentInspect
Send a workspace AI agent into a live call — ONE tool for every surface; target decides where: a Google Meet or Microsoft Teams link → meeting bot; a Telegram @group / t.me link / chat_id / t.me/call slug → Telegram group voice chat or conference; 'new' → creates a NEW Telegram conference call and returns its shareable link; a WhatsApp group JID (digits@g.us) → WhatsApp group call; a live session UUID (from calls.list_active), or thread_id of a running call → WAKES the agent on the bot already in that call instead of spawning a second one. For a meeting link the router also checks for a bot already in that meeting and wakes it rather than double-joining. Always pass agent_id (from agents.list). For a live TRANSLATOR use calls.dispatch_translator instead.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | target='new' only: label for the new conference. | |
| target | No | Where to send the agent: a Google Meet or Microsoft Teams link, a Telegram @group / t.me link / chat_id / t.me/call slug, 'new' for a fresh Telegram conference, a WhatsApp group JID (digits@g.us), or a live session UUID to wake the agent on a running call. Omit only when passing thread_id. | |
| agent_id | Yes | ID of an active agent in this workspace (from agents.list). Any active agent can be dispatched — a voice trigger is NOT required. | |
| greeting | No | First line the agent speaks on a 1:1 call. Omit to use the agent's configured default greeting. | |
| thread_id | No | Inbox thread id of a RUNNING call — wakes the agent on the bot already in it. Alternative to a session UUID target. | |
| vision_mode | No | Screen-share capture mode (Meet + Telegram spawns): 'off', 'on_demand' (agent can call vision_query), 'continuous_0_3fps' (ambient scene captures each turn). OMIT to use 'off' (the default). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| instructions | No | Task brief for the agent, e.g. 'take notes and answer questions about the roadmap'. On a spawn it is woven into the voice system prompt; on a wake it is the kickoff turn the agent responds to. OMIT for a generic listening/greeting agent. | |
| start_immediately | No | Spawn routes only: if true the agent starts talking as soon as it joins instead of waiting to be addressed. OMIT (default false) to stay silent until addressed. Wake routes always speak immediately. | |
| channel_account_id | No | Telegram routes: workspace Telegram account that joins/founds the call. Optional with exactly one Telegram account; required with several. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description states the tool 'creates a NEW Telegram conference call' and sends agents into live calls, which contradicts the annotation readOnlyHint=true. Three other hints (openWorld, idempotent, non-destructive) are consistent with the described wake-instead-of-double-join behavior, but the readOnly conflict is a genuine safety-signal mismatch that could lead an agent to dispatch without confirming side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and the routing decision, then qualified with wake-vs-spawn, the required agent_id, and the translator alternative. It is dense rather than padded, though it re-lists the `target` values almost verbatim from the schema, which is redundant for an agent that can read both.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter, 6-route dispatch tool with no output schema, the description covers every routing branch, the wake-vs-spawn distinction, and the one return value it exposes ('returns its shareable link'). It omits failure/prerequisite behavior (e.g., whether the call must be live, permission requirements, what happens when the target is unreachable), which keeps it short of a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 10 parameters, including the full `target` enumeration, and the description largely restates that same routing list. It adds a little genuine guidance (agent_id provenance, thread_id as an alternative to a session UUID, translator routing), but not enough to exceed the baseline when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('Send a workspace AI agent into a live call') and immediately disambiguates the router with 'ONE tool for every surface; `target` decides where'. The enumerated surfaces (Meet/Teams link, Telegram group/chat_id/t.me/call, 'new', WhatsApp JID, session UUID, thread_id) let an agent distinguish this tool from siblings like calls_dispatch_translator without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use and when-not-to-use: 'For a live TRANSLATOR use calls.dispatch_translator instead', 'Always pass agent_id (from agents.list)', and it points to calls.list_active as the source of a session UUID. It also states the wake-vs-spawn rule and its fallback ('the router also checks for a bot already in that meeting and wakes it rather than double-joining'), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calls_dispatch_translatorARead-onlyIdempotentInspect
Send a live speech translator to a call. target decides where: a Google Meet or Microsoft Teams link → meeting bot; a Telegram @group / t.me link / chat_id → Telegram group voice chat; 'new' (or omitted) → creates a native DialogBrain meeting with translation on and returns the join + guest links; a native meeting call_id or /meeting/ URL → enables translation on that running meeting. Always pass target_language (ISO code); optional app_languages (extra subtitle-only languages), sentence_length (short|medium|long, native only), silent (subtitles without voice), source_language (the meeting's spoken language, ISO code — improves recognition; omit for autodetect), tts_provider + tts_voice (the translator's voice; omit for the workspace default).
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | native_new only: meeting title. | |
| silent | No | Meeting routes only (Google Meet / Teams): true = subtitles without speaking into the call. OMIT for a normal speaking translator. | |
| target | No | Where to send the translator: a Google Meet or Microsoft Teams link, a Telegram @group / t.me link / chat_id, 'new' for a fresh native meeting, or an existing native meeting call_id or /meeting/ URL. Omit for a new native meeting. | |
| agent_id | No | Meeting routes ONLY (Google Meet / Teams), and required there: the active agent the bot session is recorded against (calls.send_to_meet needs one). Get it from agents.list. Telegram and native meetings ignore it — they arm translation on a call that already exists. | |
| tts_voice | No | Specific voice id for tts_provider (e.g. 'alena', 'nova'). Omit for the provider default. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| tts_provider | No | Translator VOICE provider (cartesia, openai, yandex, deepgram, ...). Omit for the workspace default translation voice. Meeting routes (Google Meet / Teams) + native new meetings. | |
| app_languages | No | Extra subtitle-only languages (max 4). | |
| comeback_phrase | No | Attention-recall phrase. | |
| sentence_length | No | Native meetings only: short|medium|long buffering (default medium). | |
| source_language | No | The call's spoken language (ISO code, e.g. 'en', 'ru'). When set, speech recognition runs in that language's dedicated mode for better accuracy. REQUIRED in practice for languages autodetect does not cover (e.g. 'vi', 'th', 'id', 'tl'). Omit when participants may speak multiple languages (autodetect). All routes. | |
| target_language | Yes | Primary spoken translation target (ISO code, e.g. 'th'). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, yet the description says the 'new'/omitted route 'creates a native DialogBrain meeting with translation on' and that a native call_id 'enables translation on that running meeting' — both are environment mutations. The conflicting signal is severe for an agent deciding whether this call has side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the routing rule and structured with backticked parameter names and arrows, so the branching is scannable despite being one long paragraph. Some sentences duplicate schema text, which is waste given 100% coverage, but nothing else is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter tool with no output schema, the description covers routes, defaults, conditional applicability ('native only', 'Meeting routes only') and notes the 'new' route returns join + guest links. It omits failure behavior and how to subsequently stop/mute translation, but the core calling contract is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema descriptions are themselves rich (route restrictions, 'OMIT' semantics, examples). The prose restates rather than extends the schema — it adds no format, unit, or validation detail beyond what the properties already document, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Send a live speech translator to a call') and immediately differentiates the four routing surfaces for `target`. An agent can tell exactly what capability this exposes without opening the schema, and the routing table is far more specific than any sibling (e.g. calls_set_translation_language).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description is essentially a decision tree: Meet/Teams link → meeting bot; Telegram handle/link/chat_id → Telegram voice chat; 'new'/omitted → fresh native meeting; call_id or /meeting/ URL → arm translation on a running meeting. It also gives when-not guidance implicitly ('OMIT for a normal speaking translator', 'Telegram and native meetings ignore it'), which removes almost all inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calls_get_transcriptARead-onlyIdempotentInspect
Get the structured transcript and final state of a voice call by call_id. Returns per-turn rows in chronological order, call status (active/completed/failed/abandoned), duration, and an outcome field telling whether the recipient picked up (answered/no_answer/busy/declined/failed/unknown). answered_at is non-null once the recipient picked up. Returns active turns if the call is still in progress.
| Name | Required | Description | Default |
|---|---|---|---|
| call_id | Yes | Call ID returned by calls.make in _meta.call_id. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful behavioral detail beyond that: it enumerates the status values, explains the `outcome` pickup semantics, and defines the `answered_at` non-null condition. Auth and rate-limit behavior remain undisclosed, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the core purpose, then progressively detailing return contents. Every sentence carries information, though the enumeration of status/outcome values is dense. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately carries the return-value burden and does so thoroughly (per-turn rows, status, duration, outcome, answered_at). For a read-only two-parameter getter this is nearly complete; only cross-tool routing guidance is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already fully documented in the schema (including where call_id comes from). The description only restates 'by call_id' and says nothing about in_workspace, adding no meaning beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the structured transcript and final state of a voice call') plus the access key ('by call_id'). This is clearly distinguishable from the list-oriented siblings calls_list_history and calls_list_active, which the agent can infer without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the note that it 'Returns active turns if the call is still in progress' tells the agent it can be used on live calls, but there is no explicit when-to-use guidance and no routing to alternatives (e.g. when to prefer calls_list_history). Adequate but with a clear gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calls_hangupARead-onlyIdempotentInspect
Hang up an active voice call by call_id. Use after calls.make when the agent decides to terminate before the callee does, or to abort a stuck call. Idempotent: returns success if the call is already terminal.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | Short internal reason for ending the call (e.g. 'campaign timeout'). Stored on voice_sessions.metadata. | |
| call_id | Yes | Call ID returned by calls.make in _meta.call_id. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations declare readOnlyHint=true, but the description says it 'hangs up an active voice call' and 'terminates' a call, which is an action that changes the environment/state of the call. This conflicts with a read-only declaration. (idempotentHint/idempotency and destructiveHint=false are consistent and add no contradiction.)
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and identifier, followed by usage and the idempotency guarantee. No filler; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple action tool with full schema coverage and no output schema, the description covers purpose, usage, and idempotency. However, side effects (e.g., whether the callee is notified, permission/auth needs) are absent, and the description is silent on behavioral impact that is muddled by the readOnly annotation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so call_id, reason, and in_workspace are already fully documented with examples. The description only restates call_id as the key and implies the post-calls.make source, adding little beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb + resource + identifier: 'Hang up an active voice call by call_id.' This clearly distinguishes it from sibling call-control tools like calls_transfer, calls_mute_translation_tts, and calls_make.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear triggering context: use after calls.make to terminate before the callee does, or to abort a stuck call. It references a sibling (calls.make) as the prerequisite but does not name explicit alternatives/exclusions (e.g., transfer vs. hangup) or state when NOT to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calls_list_activeARead-onlyIdempotentInspect
List active voice calls in this workspace. Use before calls.make on a Telegram account (only one MTProto call per account at a time) to check whether the line is free.
| Name | Required | Description | Default |
|---|---|---|---|
| channel | No | Filter by voice channel. OMIT to include all channels. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| channel_account_id | No | Filter by channel_account.id (the calling Telegram account or Twilio number). Combine with channel for a per-line busy check. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly, idempotent, non-destructive behavior, so the bar is lower; the description adds genuinely useful domain behavior — the one-MTProto-call-per-account limit and the pre-flight role of the call — which no annotation conveys. Return payload shape is not described, but that is a minor omission for a filtered list.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the listing scope front-loaded and the usage rationale immediately after. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-parameter read-only list tool, the description covers purpose and the exact decision it supports. With no output schema, the returned call fields are left unstated, a small gap but not one that blocks correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the channel, in_workspace, and channel_account_id descriptions already explain filtering and the per-line busy check. The description adds no parameter syntax or semantics beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'List active voice calls in this workspace.' The qualifier 'active' distinguishes it from history-style siblings, and the description ties it to a concrete decision (is the line free) rather than restating the name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use: run it before calls.make on a Telegram account, and explains the governing constraint (only one MTProto call per account at a time). It does not name when-not to use it or point to a listing alternative such as calls_list_history, which keeps it short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calls_list_historyARead-onlyIdempotentInspect
Search historical voice calls in this workspace by participant name, contact_id, thread, channel, source, and/or date range. Returns one row per call (NOT per turn) with call_id, duration_seconds, outcome, direction, started_at, source, channel_label, and parent_thread_id (the originating chat thread for Telegram-group / Twilio-outbound / Meet calls). Pair with calls.get_transcript(call_id) for the full per-turn transcript. Use this instead of messages.read_history for cross-thread call queries — group calls and Meet sessions live on per-call sub-threads, not on the parent chat thread.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum calls to return (default 20, max 100). | |
| since | No | ISO date or datetime lower bound (inclusive). Default: 90 days ago. Naive timestamps are interpreted as UTC. | |
| until | No | ISO date or datetime upper bound (inclusive). Default: now. | |
| source | No | Filter by voice_sessions.source: 'telegram' (1:1 + group), 'whatsapp' (native WhatsApp voice), 'twilio' (PSTN), 'meet' (Google Meet bot), 'livechat' (in-app voice), 'android' (Android device). OMIT to include all sources. | |
| channel | No | Filter by message-level channel of the call thread: 'telegram' (1:1 voice or group call sub-thread), 'twilio_voice', 'meet_voice', 'livechat_voice', 'whatsapp' (native WhatsApp calls — these coalesce into the contact's messaging thread). OMIT to include all voice channels. | |
| thread_id | No | Restrict to calls on this thread OR with this thread as their originating parent (Telegram group → call sub-thread back-link, Twilio outbound source_thread_id back-link). | |
| contact_id | No | Filter by exact entity_id (from contacts.find). Mutually exclusive with participant_name when both target the same person. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| participant_name | No | Filter to calls whose parent thread has a participant matching this name (substring match against entity.title). Resolves group calls via the parent group's roster, not the per-call thread's speaker list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower; the description adds real value by disclosing the return granularity ('one row per call, NOT per turn') and the field list, which matters because no output schema exists. It does not cover pagination behavior beyond the limit param or auth requirements, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the core action front-loaded and zero filler. The middle sentence enumerates return fields and is dense but each item earns its place given the absent output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully carries the return-value burden by naming the row granularity and fields, and it closes the routing question against messages.read_history. For a 9-param read tool this is complete enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter already carries its own semantics and enum meanings. The description restates the filterable dimensions (participant name, contact_id, thread, channel, source, date range) but adds no syntax or format detail beyond the schema; baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (historical voice calls) plus the filter dimensions, and explicitly distinguishes itself from siblings calls.get_transcript and messages.read_history. An agent can tell it apart from list_active and transcript tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to prefer this over messages.read_history ('cross-thread call queries — group calls and Meet sessions live on per-call sub-threads') and names the complementary tool to pair with (calls.get_transcript). Both the alternative and the selecting condition are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calls_makeDestructiveInspect
Place an outbound AUDIO/VOICE phone call via Twilio (PSTN) or Telegram (MTProto 1:1 call). Use this any time the user asks to 'call', 'ring', 'phone', 'dial', or have a spoken conversation. Do NOT use messages.send when the user asks to call someone — a call is real-time voice, not a text message. You conduct the conversation as the voice agent using the provided greeting and instructions.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | Android package that places the call (generic — any app with a dispatch recipe, or manual UI dispatch). Example: 'com.whatsapp'. Used only when channel='android'. Ignored for other channels. | |
| channel | No | Voice transport: 'twilio' or 'telnyx' (phone via PSTN — both require phone_number in E.164; pick the carrier the workspace has connected), 'telegram' (MTProto 1:1 call — requires telegram_user_id, NOT a phone number or thread_id), 'maxru' (Max.ru voice call — requires maxru_user_id), 'android' (Android device voice call — requires phone_number in E.164), 'whatsapp' (WhatsApp voice call via the workspace's connected WhatsApp account — requires phone_number in E.164), 'whatsapp_business' (a Meta Cloud API number; requires phone_number in E.164; the contact must have allowed calls, otherwise a permission request is sent instead). OMIT to auto-select based on the current thread (e.g. inside a Telegram DM → uses 'telegram'). | |
| greeting | Yes | The first sentence the agent speaks immediately when the call connects. ALWAYS provide a greeting — without it the caller hears silence. Keep it short and natural. Example: 'Hi, this is Diana calling from DialogBrain. Do you have a moment to chat?' | |
| contact_id | No | LINE only: the contact entity ID for storing the confirmed display name after placement. Ignored for other channels/apps. | |
| report_back | No | When to re-invoke you after the call ends. 'on_answer' (default) = only if the call was answered, 'always' = even on missed/failed calls, 'never' = fire and forget. Transcript is always stored regardless of this setting. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| instructions | No | What to do during the call — objective, questions, tone. The AI generates a natural opening and guides the conversation. Example: 'Call about invoice #1234. Ask if they received it and when payment is expected. Be friendly and professional.' | |
| phone_number | No | Destination phone number in E.164 format (e.g., '+15551234567', '+66812345678'). Required when channel='twilio'. | |
| viber_number | No | Destination Viber number in E.164 (e.g., '+15551234567'). Required when channel='viber' — the call is placed from the workspace's connected Viber account. | |
| maxru_user_id | No | Destination Max.ru user/chat ID. Required when channel='maxru'. | |
| skip_dispatch | No | Set true after you placed the call manually with the android_* UI tools — skips the automatic recipe and attaches straight to the already-live call. OMIT for the normal automatic dispatch. Used only when channel='android'. Ignored for other channels. | |
| voice_agent_id | No | ID of the agent that conducts the call (an `id` from agents.list). If omitted, uses the workspace's default voice-capable agent when one exists. Pass this when the call fails with 'No voice agent configured'. | |
| telegram_user_id | No | Destination Telegram user ID (decimal int64 as string, e.g. '123456789'). Required when channel='telegram'. The caller account must have had prior interaction with this user — a cold contact cannot be reached via voice. | |
| target_display_name | No | LINE only: the contact's LINE display name to place the voice call to (resolved by the planner from the contact). Ignored for other channels/apps. | |
| permission_request_text | No | Only for channel='whatsapp_business': the text of the call-permission request WhatsApp sends when the contact has not yet allowed calls. Omit for a generic default. |
calls_meet_browserARead-onlyIdempotentInspect
Attach to a Google Meet bot's live browser to diagnose and recover a bot that isn't visibly joining. Pass the meet session's call_id; returns a page_id. Then drive the bot's Meet page with the generic browser tools (browser.snapshot / browser.click / browser.take_screenshot / browser.evaluate / browser.console_messages / browser.network_requests) using that page_id — read the snapshot to see whether the bot is in the lobby, blocked, or admitted, and click guest-side controls to recover a stalled join. Note: host admission ('Admit') happens in the host's own browser and is not present on the bot's page.
| Name | Required | Description | Default |
|---|---|---|---|
| call_id | Yes | The meet session's call_id (UUID), e.g. from calls.send_to_meet's session_id or calls.list_active. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/openWorld/non-destructive, so the bar is lower. The description still adds real context: it returns a page_id, interaction happens through generic browser tools rather than this call, and recovery means clicking guest-side controls. The caveat that host-side 'Admit' is not on the bot's page is a genuinely useful behavioral boundary.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and every sentence carries information — the return value, the driving tools, the recovery intent, and the host-admission caveat. It is dense and slightly long, but nothing is wasted or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must cover returns and it does (page_id). It also explains the downstream workflow and the critical caveat about host admission, giving an agent everything needed to call and use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (call_id, in_workspace) are fully documented in the schema. The description reinforces the required call_id ('Pass the meet session's call_id') but adds no syntax or format detail beyond the schema, matching the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Attach) and resource (a Google Meet bot's live browser) with an explicit goal (diagnose and recover a bot that isn't visibly joining). It is clearly distinguishable from generic browser tools because it targets a bot's Meet page and returns a page_id for follow-up. An agent can identify the tool's role without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger ('a bot that isn't visibly joining') plus the follow-on workflow, naming the exact browser.* tools to drive the page. It lacks an explicit 'do not use when X' or a direct routing statement against the close sibling browser_attach_meet, so it stops just short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calls_mute_translation_ttsARead-onlyIdempotentInspect
Mute or unmute an active call's live-translation TTS. scope='call' (default, Google Meet sessions ONLY): the translator goes silent INSIDE the call — participants hear no bot voice, in-flight audio is cut immediately, but translation audio + subtitles remain available in the app (listen overlay). scope='everywhere' (any channel): stop synthesizing entirely from the next utterance — subtitles only, no audio anywhere. muted=false reverses either. Translation itself keeps running throughout.
| Name | Required | Description | Default |
|---|---|---|---|
| muted | Yes | true to mute, false to unmute. | |
| scope | No | 'call' (default; Google Meet only): silent in the call only — app audio + subtitles keep flowing. 'everywhere' (any channel): no TTS synthesis at all (subtitles only). OMIT to use the 'call' default. | |
| call_id | Yes | The active call / Meet session call_id (UUID) — e.g. from calls.send_to_meet's session_id or calls.list_active. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover idempotency and non-destructiveness, yet the description adds real behavioral detail: in-flight audio is cut immediately, subtitles and app-side listen overlay survive the 'call' mute, and translation itself keeps running. The only friction is the mild tension with readOnlyHint=true for an operation that toggles live audio behavior, though no persistent data is mutated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose in the first clause, then elaborates scope behavior and reversibility. Dense but each sentence carries an actionable distinction between the two scopes, with little filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description covers both modes, the default, reversibility, and the ongoing-translation behavior, which is what an agent needs to call it correctly. Minor omissions (prerequisites like an active translation session, invalid call_id behavior) keep it just under complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents call_id, muted, scope, and in_workspace. The description largely restates the scope semantics rather than adding syntax or edge-case meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (mute/unmute) and a precise resource (an active call's live-translation TTS), immediately distinguishing it from sibling tools like calls_set_translation_language or calls_hangup. The scope semantics make the boundary of the operation explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use conditions for each mode: scope='call' is Google Meet sessions ONLY, scope='everywhere' works on any channel. It also states muted=false reverses either, so the agent knows both directions. It does not name a sibling alternative or an explicit when-not-to-use case, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calls_set_translation_languageBRead-onlyIdempotentInspect
Change the target language of an active voice/Meet call's live translator on the fly — no hangup or re-dispatch (also arms translator mode if it isn't already on). Pass the call_id and an ISO language code, e.g. 'th' (Thai), 'ru' (Russian), 'es' (Spanish), 'en' (English). Takes effect within ~10ms — speak and the translation switches to the new language.
| Name | Required | Description | Default |
|---|---|---|---|
| call_id | Yes | The active call / Meet session call_id (UUID) — e.g. from calls.send_to_meet's session_id or calls.list_active. | |
| language | Yes | Target language ISO code: 'th' (Thai), 'ru' (Russian), 'es' (Spanish), 'en' (English), etc. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, yet the description describes an operation that mutates live call state: it changes the target language and 'arms translator mode if it isn't already on'. This directly contradicts the read-only claim, which would lead an agent to treat the call as untouched when the tool actually alters what participants hear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and immediately followed by the key constraint (no hangup/re-dispatch) and the latency expectation. Slightly padded by duplicating the language examples already present in the schema, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with full schema coverage and no output schema, the description covers the important runtime behavior: it works on an in-flight call, arms translator mode, and takes effect in ~10ms. It omits error behavior for an invalid language code and the exact response, but those are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so call_id and language are already documented, and the description merely repeats the same ISO-code examples ('th', 'ru', 'es'). It adds no syntax, validation, or failure semantics beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource+scope: changing the target language of an active voice/Meet call's live translator. It implicitly separates itself from the dispatch path ('no hangup or re-dispatch'), but never addresses the near-identically named sibling calls_set_translation_languages (plural), which is the most likely confusion point for an agent choosing a tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use (an active call whose translator is running) and states that it also arms translator mode if it is off, so the agent knows it works pre-emptively. It implicitly rules out the alternatives ('no hangup or re-dispatch') but never names a sibling tool or an explicit when-not condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calls_set_translation_languagesARead-onlyIdempotentInspect
Set the EXTRA target languages an active voice/Meet call's translator produces for the APP ONLY (subtitles + listenable audio in the inbox — NOT spoken into the call). Operators switch between them in the UI. Pass call_id and app_languages (ISO codes, e.g. ['de','fr']); pass [] to drop all extras. Max 4. The primary spoken language is managed by calls.set_translation_language. Takes effect within ~10ms; new languages translate NEW speech only.
| Name | Required | Description | Default |
|---|---|---|---|
| call_id | Yes | The active call / Meet session call_id (UUID) — e.g. from calls.send_to_meet's session_id or calls.list_active. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| app_languages | Yes | Extra app-only target languages (ISO codes), e.g. ['de','fr','es']. Pass [] to remove all extras. Max 4. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is a mutating setter ('Set the EXTRA target languages', 'Takes effect within ~10ms'), yet the annotations declare readOnlyHint=true. A write operation cannot be read-only, so the description directly contradicts the structured safety signal and could mislead an agent about side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the key distinction (EXTRA app-only vs primary spoken) and the mutating semantics, then constraints and timing. Dense but every clause carries information; only the trailing timing sentence is somewhat tangential.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a call-scoped config setter with no output schema, it covers effect timing, propagation scope, the 4-language cap, clearing semantics, and the sibling responsible for the primary language. The main gap is behavioral (permissions/state impact), which is muddied by the contradictory readOnly annotation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents call_id, in_workspace, and app_languages (ISO codes, [] to clear, max 4). The description largely restates that (ISO codes, [] to drop, max 4) and adds only the behavioral note that new languages translate new speech, so it does not meaningfully exceed the schema on parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Set) plus resource (extra app-only translation languages) with the scope nailed down: subtitles + listenable audio in the inbox, NOT spoken into the call. It explicitly separates this from calls.set_translation_language, so an agent can distinguish it from its closest sibling without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly states the context of use (extra app-only target languages for an active voice/Meet call) and names the alternative for the primary spoken language via calls.set_translation_language. It also tells the agent how to clear extras ([]). It stops short of an explicit 'use X when / don't use Y when' framing, but routing is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calls_transferARead-onlyIdempotentInspect
Transfer the current phone call to a person on another number (e.g. 'let me put you through to a manager'). Only works on a live PHONE call, and only to a number the operator has pre-approved for this agent. Say one short sentence to the caller first ('connecting you now'), THEN call this. It rings the destination and returns when they pick up — which can take most of a minute — and only then do you leave the call. If nobody answers you are told so and are still on the line with the caller.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | The phone number to transfer the caller to, E.164 (e.g. '+15551234567'). Must be pre-approved for this agent. | |
| call_id | No | Which call to transfer. Omit inside a live voice turn — it is taken from the call. Required over MCP. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already covering readOnly/idempotent/open-world, the description adds substantial call-flow detail they cannot express: it rings the destination, blocks until pickup (up to 'most of a minute'), only then does the agent leave the call, and if nobody answers the agent is informed and remains on the line. The pre-approval prerequisite is also disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the action and example, then layers constraints, ordering, latency, and failure behavior in short clauses. Every sentence carries actionable information (timing, no-answer handling, speak-before-calling), with no redundant restatement of the name or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description explains the return semantics anyway (returns on pickup, reports no-answer while staying on the line) plus the latency an agent must plan around. For a 3-parameter live-call tool this leaves nothing material unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so `to`, `call_id`, and `in_workspace` are already documented with their formats and constraints. The description reinforces the pre-approval rule for the destination number but adds no syntax, range, or default information beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Transfer the current phone call to a person on another number') and scopes it tightly ('Only works on a live PHONE call'). The phrase 'to a person on another number' implicitly separates it from agent_handoff (handing to an AI) and calls_make (originating a call), so an agent can route correctly without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong operative conditions: only on a live PHONE call, only to a pre-approved number, and the sequencing rule ('Say one short sentence to the caller first ... THEN call this'). It does not name a sibling alternative or state a when-not-to-use beyond the live-call/pre-approval limits, so it stops short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calls_waitARead-onlyIdempotentInspect
Block until a voice call ends (status changes from 'active') or timeout elapses. Returns ended=true with final state when the call has ended; ended=false on timeout (re-issue to keep waiting). The returned state includes outcome so callers can branch on pickup vs. no-answer (answered/no_answer/busy/declined/failed/unknown). Default timeout 90s; cap 110s — bounded by nginx proxy_read_timeout 120s on /mcp.
| Name | Required | Description | Default |
|---|---|---|---|
| call_id | Yes | Call ID returned by calls.make in _meta.call_id. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| timeout_seconds | No | Max seconds to wait. Default 90, cap 110 (bounded below nginx 120s proxy_read_timeout). On expiry returns ended=False with status='active' so the caller can re-issue to keep waiting. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnly, idempotent, non-destructive). The description adds substantial behavioral context beyond that: the blocking/hold semantics, the ended=true/false return contract, the branchable outcome enum values, and the timeout default/cap tied to the nginx proxy_read_timeout limit. For a long-polling tool with no output schema, this is exactly the disclosure needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core blocking behavior, then return semantics, then outcome values, then timeout specifics. Efficient and well-ordered, though the timeout default/cap is restated in both the description and the schema, a minor redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully carries the return contract (ended flag, final state, outcome enum) and the timeout/expiry behavior. An agent has everything needed to call and branch on results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents call_id, in_workspace, and timeout_seconds. The description reinforces the timeout default/cap and the re-issue pattern, but adds little else. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource: 'Block until a voice call ends (status changes from active) or timeout elapses.' This clearly distinguishes it from siblings like calls_list_active, calls_hangup, and calls_make, which an agent can rule out without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains when the tool applies (wait for call completion) and what to do on timeout ('re-issue to keep waiting'). It gives clear operational context but does not explicitly contrast with an alternative like polling calls_list_active or using calls_get, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
campaigns_cancelCRead-onlyIdempotentInspect
Cancel a campaign — no further numbers are dialed.
| Name | Required | Description | Default |
|---|---|---|---|
| campaign_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, but canceling a campaign changes its state by stopping dialing, which is a mutation. The description ('Cancel a campaign') therefore contradicts the readOnlyHint annotation, and no additional behavioral context (reversibility, effect on history) is provided to resolve the conflict.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler. The core action and its consequence are stated immediately, and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-changing tool with no output schema and misleading annotations, the description omits whether cancellation is reversible, whether it deletes campaign data or call history, and how it relates to pause. It also provides nothing to correct the readOnlyHint contradiction an agent would otherwise trust.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: in_workspace is documented in the schema, while campaign_id is self-evident as the required identifier. The description adds no parameter-level detail (e.g., format, scope, or side effects of in_workspace), so it does nothing beyond the partially-covering schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Cancel) and resource (campaign) and adds the operational consequence that no further numbers are dialed. It does not, however, distinguish itself from the sibling campaigns_pause, which is the distinction an agent most needs when deciding between stopping and pausing a campaign.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance or mention of alternatives such as campaigns_pause or campaigns_resume. The title of the sibling set implies a pause/cancel contrast exists, but the description never tells the agent which condition selects cancel over pause.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
campaigns_createBRead-onlyIdempotentInspect
Start a bulk outbound call campaign: dial a list of phone numbers with a voice agent, at a bounded concurrency, retrying no-answers. Phone network only (twilio/telnyx). It begins dialing on its own — no separate start.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Campaign name | |
| channel | No | Carrier (default telnyx). | |
| agent_id | No | Voice agent that makes the calls. | |
| greeting | No | First line the agent says on pickup. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| instructions | No | What the agent should do on the call. | |
| max_attempts | No | Retries per number on no-answer (default 3). | |
| phone_numbers | Yes | Numbers to dial, E.164 (e.g. '+15551234567'). | |
| max_concurrent | No | Simultaneous calls (default 2). | |
| retry_delay_minutes | No | Wait before a retry (default 30). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description clearly describes a write/mutation action that creates a campaign and starts dialing places calls, while annotations declare readOnlyHint=true and idempotentHint=true. This is a direct contradiction: a tool that 'begins dialing on its own' is not a read-only, idempotent operation. Flagged as an annotation contradiction, so the score is capped at 1.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action. The trailing clause 'It begins dialing on its own — no separate start' earns its place by preempting a likely follow-up question. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation tool with full schema coverage and no output schema, the description covers purpose and key runtime behavior (auto-dial, retries, concurrency). It omits permission/auth requirements and any note about phone-number validation or failure handling, leaving modest gaps given the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all ten parameters (name, phone_numbers, agent_id, max_concurrent, max_attempts, etc.). The description restates a few of them (list of numbers, voice agent, bounded concurrency, retrying no-answers) but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Start a bulk outbound call campaign') and specifies scope precisely: dialing a list of numbers with a voice agent at bounded concurrency, phone-network only. An agent can immediately tell this apart from single-call siblings like calls_make and from campaigns_pause/resume/status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage context ('Phone network only (twilio/telnyx)', 'begins dialing on its own') but never states when to use this versus alternatives such as calls_make for a single call, or the sibling campaign lifecycle tools. Usage is implied rather than explicit, and no when-not guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
campaigns_listBRead-onlyIdempotentInspect
List this workspace's campaigns, newest first.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered. The description adds the ordering trait ('newest first'), which is genuine behavioral context, but says nothing about pagination behavior or the workspace-override semantics of in_workspace.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the resource and the ordering constraint front-loaded and zero wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description should describe what a campaign entry looks like and how limit/offset pagination behaves. It supplies neither, leaving an agent to guess at return shape and paging for a list tool with 3 parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only in_workspace is documented), and the description mentions no parameters at all. limit and offset remain undocumented in both places, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb (List) and resource (this workspace's campaigns) plus an ordering guarantee ('newest first'). It is distinguishable from the write-oriented siblings (cancel/create/pause/resume/status), though it does not explicitly contrast itself with campaigns_status, which also reads campaign state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no when-to-use guidance, no exclusions, and no mention of alternatives. An agent must infer that this is the enumeration tool from the name alone; nothing routes it away from campaigns_status for a single campaign's state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
campaigns_pauseCRead-onlyIdempotentInspect
Pause a campaign — stop starting new calls; in-flight ones finish.
| Name | Required | Description | Default |
|---|---|---|---|
| campaign_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, but pausing a campaign changes its runtime state (no new calls are started), directly contradicting the read-only claim. The description's detail about in-flight calls is useful, but the conflict between the stated mutating behavior and the annotation is a serious inconsistency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the effect front-loaded and the in-flight caveat second. Every word earns its place and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, no-output-schema tool, the description covers the main behavioral question (what happens to running calls). It is incomplete on permissions, immediacy, and how pause relates to resume/cancel, and it leaves campaign_id unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: in_workspace is documented in the schema, but campaign_id has no description anywhere. The tool description adds no parameter guidance whatsoever, so it fails to compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (pause) and resource (campaign) plus the operational meaning: new calls stop, in-flight calls finish. This implicitly separates it from a hard 'cancel', though the description never names campaigns_cancel or campaigns_resume explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the name and by 'stop starting new calls; in-flight ones finish', which hints at when pause is preferable to cancel. However, there is no explicit when-to-use/when-not or named alternative among the campaigns_* siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
campaigns_resumeCRead-onlyIdempotentInspect
Resume a paused campaign.
| Name | Required | Description | Default |
|---|---|---|---|
| campaign_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description states a state-changing operation ('Resume a paused campaign') while the annotations declare readOnlyHint=true, which directly contradicts the described effect. Because the structured fields are actively misleading here, the description cannot be credited for behavioral disclosure and the inconsistency must be flagged.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short, front-loaded sentence with no filler or repetition. It is efficient, though the brevity edges toward under-specification for a state-mutating tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, a contradictory readOnlyHint, one fully undocumented required parameter, and no usage guidance, the definition is materially incomplete. An agent cannot confidently determine prerequisites, effects, or how to source campaign_id.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: in_workspace is documented in the schema, but campaign_id carries no description anywhere and the tool description adds nothing about where to obtain it or what form it takes. With half the parameters undocumented, the description fails to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Resume') and resource ('campaign') plus the precondition that the campaign must be paused, so an agent knows what the call does without opening the schema. It does not, however, distinguish itself from the nearby siblings campaigns_pause, campaigns_cancel, or campaigns_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage signal is the implied precondition 'paused campaign'. There is no explicit when-to-use guidance, no statement of what to do if the campaign is already active or cancelled, and no pointer to the sibling tools that manipulate the same resource.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
campaigns_statusBRead-onlyIdempotentInspect
Progress of one campaign: status and per-number counts.
| Name | Required | Description | Default |
|---|---|---|---|
| campaign_id | Yes | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds only that the result contains status and per-number counts; it says nothing about freshness, polling behavior, or what 'per-number' refers to.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though the fragmentary 'Progress of one campaign' phrasing is slightly terse even for a compact tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only status tool with no output schema, the description is roughly adequate, but 'per-number counts' is left ambiguous (counts of what, per which number type), and no return-value shape is hinted at. An agent knows it is a read but not much about the payload.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: in_workspace is documented in the schema, but campaign_id has no description in either place. The phrase 'one campaign' implies campaign_id selects a single campaign, so the description adds marginal meaning but does not compensate for the undocumented required parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource ('one campaign') and the data returned ('status and per-number counts'), which meaningfully distinguishes it from the sibling campaigns_list. It lacks an explicit verb and does not name any sibling, but the scope is clear enough to route correctly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no when-to-use guidance, no prerequisites, and never mentions alternatives like campaigns_list or campaigns_pause. An agent must infer usage entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
channels_connectAInspect
🔌 Connect a messaging channel to this workspace, or find out how.
Call with no channel for the catalogue: every channel, its connect_mode and its link.
Call with channel only to learn what to ask the person for (credential_fields).
Call with channel + credentials to connect a credentials channel: email over IMAP with an app password, telegram_bot with a bot token, line with the channel access token and channel secret, avito.
Channels that need a QR scan, a phone code or a browser sign-in (telegram, whatsapp, instagram, facebook, email (Google sign-in), youtube, tiktok, x, threads, linkedin, wechat, zalo, maxru, the voice carriers) cannot be connected from here: the tool answers success with needs_person=true and the link https://dialogbrain.com/connect/, which opens the connection dialog for that channel. Give the person that link; do not tell them to look for it in settings. A website widget is not connected here at all: use widgets.create.
| Name | Required | Description | Default |
|---|---|---|---|
| channel | No | Channel id, e.g. 'email', 'telegram_bot', 'whatsapp'. Omit for the catalogue. | |
| credentials | No | The fields credential_fields asked for. Never echoed back. | |
| display_name | No | Name for the account in the channel list. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare this is a non-read-only, open-world, non-idempotent write, but the description adds critical behavior the annotations cannot: that unsupported channels still return success with needs_person=true plus a link, that credential_fields drive what to ask the user, that credentials are never echoed back, and that in_workspace stores nothing. This is far beyond annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then progressively details the three modes and the exclusions. Dense and information-rich, with the only mild cost being the long enumeration of unsupported channels, which is nonetheless load-bearing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the burden of describing returns (needs_person, the connect link, credential_fields) and does so explicitly. Combined with the mode breakdown and exclusions, an agent has everything needed to call this correctly and advise the user afterward.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so baseline is 3, but the description adds real semantic value by explaining how the `channel` and `credentials` parameters behave across the three modes and what `credentials` must contain (the fields credential_fields asked for). It complements rather than merely restates the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (connect a messaging channel to this workspace) and immediately distinguishes itself from siblings like channels_list and channels_get_profile by framing itself as both a connector and a catalogue lookup. The catalogue/learn/connect tri-mode framing makes the tool's scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly enumerates the three calling modes (no channel, channel only, channel+credentials) and spells out which channels cannot be connected here, redirecting QR/phone/browser-signin channels to the hosted dialog and website widgets to widgets.create. This is textbook when/when-not/alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
channels_connect_telegram_botAInspect
🤖 Connect a Telegram Bot (Bot API) channel from its bot token.
When to use:
After a bot was created via @BotFather and you have its token.
The token alone is sufficient — no Telegram user account is needed.
Validates the token, creates the channel account, and registers the webhook so the bot starts receiving messages immediately.
| Name | Required | Description | Default |
|---|---|---|---|
| bot_token | Yes | Telegram bot token from @BotFather (e.g. '123456789:ABCdef...'). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false. The description goes beyond them by disclosing the concrete effects: token validation, channel account creation, and immediate webhook registration so messages start flowing. It omits failure handling for an invalid token and whether a retry produces a duplicate channel, which is the main remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action in the first line, followed by a tight two-bullet usage block and one sentence on effects. Efficient and scannable, though the 'When to use:' heading is slightly ceremonial for two bullets.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covered well enough for a two-parameter connection tool: it explains prerequisites, the credential needed, and the side effects (account creation, webhook registration). With no output schema, the description could still say what is returned (e.g. channel id) or what an invalid token yields, but all call-critical information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with bot_token already documented including an example format, so the baseline is 3. The description's note that 'the token alone is sufficient — no Telegram user account is needed' adds contextual assurance about the required parameter, but nothing about the optional in_workspace parameter or token format beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb, resource, and protocol variant: 'Connect a Telegram Bot (Bot API) channel from its bot token.' This distinguishes it from the generic sibling channels_connect and from other channel tools without needing to open either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'When to use' section gives explicit preconditions (bot created via @BotFather, token in hand, no Telegram user account required). It does not, however, name the alternative tool (e.g. channels_connect) or state when this tool should NOT be chosen, so it is clear context without explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
channels_contactsAInspect
👥 Load people into a custom channel account, so its conversations are titled with names and each person has a card in the workspace's contacts.
contacts is a list of {contact_id, display_name, conversation_id?, avatar_url?, details?}. contact_id is the platform's id for the person. details may carry email, phone, about, company, job_title, linkedin, website. Give conversation_id for people the account already has a conversation with.
import_history: 'answered' reads into the inbox only the conversations in which the other side wrote; 'all' reads every conversation; omit it to load names only. The import runs in the background; a second request while one runs is queued, not lost. Check progress with channels.list(account_id=...): history_import.skipped counts the conversations left out and not_imported names the first 100 with the reason. Up to 5000 contacts per call; loading again is safe.
| Name | Required | Description | Default |
|---|---|---|---|
| contacts | Yes | People to load: {contact_id, display_name, conversation_id?, avatar_url?, details?}. | |
| account_id | Yes | The custom channel account (from channels.list). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| import_history | No | Also read their existing conversations: 'answered' or 'all'. OMIT to load names only. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations: the import runs in the background, a second request is queued rather than lost, progress is checked via channels.list(account_id=...), history_import.skipped/not_imported report exclusions, and the call is safe to repeat. These behaviors are not derivable from the readOnlyHint=false/destructiveHint=false annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the outcome statement, then grouped into a contacts paragraph, an import_history paragraph, and operational limits. Every sentence carries actionable information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 4 parameters and no output schema, the description covers what the tool does, how to structure input, the async processing model, and how to verify results externally. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100% (baseline 3), the description expands meaning well beyond the schema: it enumerates the contact structure fields, explains what 'details' may carry, tells when to supply conversation_id, and unpacks the import_history enum values in operational terms. This is genuine added value, not repetition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Load people into a custom channel account') plus the observable outcome (conversations titled with names, contacts get cards in the workspace). This is clearly distinguishable from siblings like contacts_sync or channels_update_profile at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage rules for the import_history parameter ('answered' vs 'all' vs omit for names only) and practical conditions (up to 5000 per call, re-loading is safe). It stops short of routing the agent away from sibling tools, so it is strong but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
channels_get_profileARead-onlyIdempotentInspect
🪪 Read a connected channel's current bot profile: display name, long description, short description, the / command menu and the persistent menu button.
Pass language_code to read back a localized variant instead of the default one.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| language_code | No | IETF language code (e.g. 'en') to read a localized variant of the text fields instead of the default. OMIT for the default variant. | |
| channel_account_id | Yes | The channel account to read, from channels list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds the returned field set and the localization behavior of `language_code`, but says nothing about auth requirements, rate limits, or behavior when the profile is empty or the channel is unreachable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the resource and its returned fields. The second sentence is largely a restatement of the `language_code` schema description, which is mildly redundant, and the leading emoji adds no information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully compensates by naming the fields the call returns, and it explains the one non-obvious optional parameter. For a simple read-only tool this is nearly complete; only error/empty-profile behavior is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including the `in_workspace` session-scoping note and the `language_code` IETF-code guidance, so the baseline is 3. The description's callout that `language_code` reads a localized variant instead of the default echoes the schema rather than extending it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ("Read a connected channel's current bot profile") and enumerates exactly what is returned: display name, long description, short description, command menu and persistent menu button. It is clearly a read operation, which implicitly sets it apart from channels_update_profile, but it never names a sibling the way a top-scoring definition would.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the agent can infer this is the read-side counterpart of channels_update_profile, and the second sentence gives a real condition for `language_code` (use it to read a localized variant, omit for the default). There is no explicit when-to-use/when-not guidance or comparison against channels_list or channels_update_profile.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
channels_listARead-onlyIdempotentInspect
📋 The messaging channels connected to this workspace and whether they are working.
One row per channel account: channel, display_name, status, error, and reconnect_url when the account needs repair. A channel in status 'error' is NOT receiving messages — say so plainly and give the person reconnect_url; it opens the repair directly, so do not tell them to look for it in settings and do not build a link of your own.
Pass only_broken=true to check for outages without listing the healthy channels.
| Name | Required | Description | Default |
|---|---|---|---|
| account_id | No | One account in detail. For a custom channel this adds whether it is listening or polling, why inbound is down, how many contacts it has, the progress of a history import, a health check and the agents that would answer on it. OMIT to list the workspace's channels. | |
| only_broken | No | true = only the accounts that need attention (status 'error'). OMIT to list every channel of the workspace. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only/idempotent/non-destructive, yet the description adds substantial behavioral context beyond them: exactly what each row contains, that status 'error' means the channel is NOT receiving messages, and that reconnect_url opens repair directly so the agent must not fabricate its own link. This is precisely the kind of guidance annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short paragraphs, front-loaded with the resource and then the actionable error-case instructions. Every sentence carries weight; nothing is redundant or decorative beyond the leading emoji.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-param read tool with full schema coverage, the definition covers output columns, error semantics, remediation behavior, and the filtering option. No output schema is needed because the return shape is described inline.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description reframes only_broken as an outage-checking use case rather than a bare filter flag, which adds a little meaning beyond the mechanical schema text, though account_id and in_workspace are left entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('list the messaging channels connected to this workspace') plus the health status dimension. An agent can distinguish it from channels_connect, channels_get_profile, and channels_contacts without opening the schema, and the row layout is spelled out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear condition for the only_broken path ('to check for outages without listing the healthy channels') and clear operational guidance for the error case. It does not name an alternative sibling to use instead, so it stops just short of fully explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
channels_update_profileAInspect
🪪 Update a connected channel's bot profile: display name, long description (the "What can this bot do?" card), short description, the / command menu, and the persistent menu button next to the input field.
Every field is optional — pass only what you want to change; any field you omit is left exactly as it is.
menu_button with type="web_app" is what puts a persistent button next to the input field that opens a Mini App page bound to this bot (pass the page's https URL as menu_button.url).
Two things this tool CANNOT set, because the Bot API does not expose them — both stay BotFather-only: the bot's AVATAR, and registering a NAMED Mini App (t.me/bot/appname). menu_button.web_app and any web_app message button work off a direct URL and need no such registration.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Bot display name (<= 64 characters). OMIT to leave unchanged. | |
| commands | No | The `/` command menu: a list of {command, description} objects (command: lowercase letters/digits/underscores, <= 32 chars; description: <= 256 chars; <= 100 commands total). OMIT to leave unchanged. | |
| description | No | Long description shown on the "What can this bot do?" card before Start (<= 512 characters). OMIT to leave unchanged. | |
| menu_button | No | The persistent button next to the input field: {type: "commands"|"default"|"web_app", text?, url?}. "commands" shows the `/` menu, "default" removes the custom button, and "web_app" opens a Mini App at `url` (must be https://) with the button labelled `text`. OMIT to leave unchanged. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| language_code | No | IETF language code (e.g. 'en') to set a localized variant of the text fields instead of the default. OMIT for the default variant every user without a matching locale sees. | |
| short_description | No | Short description shown on the profile page (<= 120 characters). OMIT to leave unchanged. | |
| channel_account_id | Yes | The channel account to update, from channels list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true). The description goes well beyond them by disclosing partial-update semantics (omitted fields are untouched) and hard capability limits (avatar and named Mini App registration stay BotFather-only), which prevents the agent from attempting operations that will silently fail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by partial-update semantics, the menu_button detail, and then limitations. Every paragraph carries actionable information, though the parenthetical asides and quoted card names make it slightly longer than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter nested mutation tool with no output schema, the description covers fields, omit-to-leave-unchanged behavior, and capability limits well. It does not mention the return shape, permission/auth requirements, or the precondition that the channel must already be connected, which are the remaining gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema carries the parameter documentation and baseline is 3. The description still adds real meaning: it explains the interplay of menu_button.type="web_app" with menu_button.url and text, and clarifies that language_code sets a localized variant rather than the default text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Update a connected channel's bot profile") and enumerates exactly which fields are affected: display name, long/short description, command menu, and persistent menu button. This cleanly separates it from the sibling read tool channels_get_profile and the other channels_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong conditional guidance — "pass only what you want to change; any field you omit is left exactly as it is" — and spells out when to use menu_button type="web_app". It does not explicitly route to channels_get_profile for reading current values first, but the context is otherwise clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collections_add_fileAInspect
Add a file to a knowledge collection.
The file must be uploaded and indexed first (files_upload + files_ingest). If the file was previously removed, it is re-enabled.
| Name | Required | Description | Default |
|---|---|---|---|
| file_id | Yes | ID of the file to add (from files_upload) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| collection_id | Yes | ID of the collection |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the bar is lower. The description adds real value on top: it discloses the required upload/index state and the notable side effect that a previously removed file is re-enabled rather than duplicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero waste; the primary action is front-loaded and the supporting conditions follow. Nothing here could be trimmed without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter add operation with no output schema, the description covers the essential prerequisites and state behavior an agent needs before calling. It stops short of describing failure modes or what a successful add returns, but the core is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so collection_id, file_id, and the workspace override are all documented in the schema itself. The description adds nothing about parameter formats or constraints, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Add a file to a knowledge collection.' An agent can immediately tell it apart from collections_remove_file or files_ingest by the action and target. It stops short of naming sibling tools, so it doesn't reach the top of the scale.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete precondition: the file must be uploaded and indexed first via files_upload + files_ingest, which tells the agent when this call is valid. It doesn't name alternative tools for adjacent operations, but the sequencing order is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collections_add_websiteAInspect
Add a website to a knowledge collection. Its pages are indexed in the background (sitemap first, else links on the same host) and kept current.
Pass path_prefix (e.g. '/help') to index only one section. Check progress with collections.list_websites.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Website address, e.g. 'https://example.com' | |
| max_pages | No | Maximum pages to index (1-500, default 100) | |
| path_prefix | No | Only index pages under this path, e.g. '/help'. OMIT for the whole site. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| collection_id | Yes | ID of the knowledge collection | |
| refresh_interval | No | How often to re-index automatically (default weekly) | weekly |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantive behavior beyond the annotations: indexing happens in the background, sitemap is tried first with same-host links as fallback, and the index is kept current. Annotations already cover the safety profile (non-read-only, non-idempotent, non-destructive), and the description usefully supplements them rather than contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short paragraphs, front-loaded with the action and its effect, then the scoping hint and progress-check pointer. No filler sentences, though the second paragraph could be folded more tightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers what happens after the call (background indexing), how to vary scope, and where to monitor progress. It omits error/duplicate handling and expected timing, but is otherwise adequate for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both required and optional parameters are already documented, including `path_prefix` and `refresh_interval`. The description largely restates the path_prefix semantics already in the schema and adds no format or boundary detail beyond it, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Add a website to a knowledge collection') and explains the indexing consequence, which distinguishes it from collections_add_file and collections_remove_website nearby. It does not, however, explicitly contrast itself with sibling tools such as collections_refresh_website.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: use `path_prefix` to scope indexing to one section, and check progress with collections.list_websites. It lacks any explicit when-not-to-use guidance, such as what to do if the site is already added.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collections_assign_agentAInspect
Assign a knowledge collection to an AI agent.
Once assigned, the agent's knowledge.query will automatically scope RAG search to files in its assigned collections.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | ID of the AI agent | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| collection_id | Yes | ID of the collection to assign |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false, idempotentHint=false). The description adds genuinely useful behavioral context beyond that: assignment automatically changes the agent's knowledge.query scoping. It does not clarify whether repeated/duplicate assignments are additive or error, which the idempotentHint=false makes relevant.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler. The action is front-loaded, and the follow-on sentence explains the consequence rather than restating the action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with full schema coverage, no output schema, and annotations covering safety, the description is nearly complete. Only the semantics of re-assignment (additive vs. replacing, idempotency) are left unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so collection_id, agent_id, and in_workspace are all documented in the schema itself. The description adds no parameter-level detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (assign) and resource (a knowledge collection to an AI agent), which is unambiguous and inherently contrasted with the sibling collections_unassign_agent. An agent can tell exactly what this does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence explains the effect (the agent's knowledge.query RAG search becomes scoped to assigned collections), which implies the reason to use it, but there is no explicit when-to-use guidance, no mention of the collections_unassign_agent alternative, and no prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collections_createAInspect
Create a named knowledge collection.
Collections group files for RAG search. After creating, add files with collections.add_file and assign to agents with collections.assign_agent.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Collection name (must be unique per user) | |
| description | No | Optional description of the collection | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it is a non-read-only, non-destructive, non-idempotent, closed-world write, so safety is covered. The description adds the RAG grouping semantics, but with no output schema it never says what the call returns (presumably a collection id needed by the follow-up calls it recommends) or that duplicate names fail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler: the first defines the action, the second gives the follow-up tool names. Front-loaded and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter create tool with full schema coverage and no output schema, the description covers purpose, concept, and next steps well. The only real gap is the return value/handle that downstream calls require, which the description leaves implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and each of the three parameters carries its own description (name uniqueness, optional description, workspace scoping), so the schema already does the work. The description adds no parameter-level information beyond the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Create a named knowledge collection") and immediately explains the domain concept ("Collections group files for RAG search"), which separates it from sibling mutators like collections_add_file or collections_assign_agent that operate on an existing collection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit workflow sequencing: after creating, add files via add_file and attach agents via assign_agent. What's missing is any when-not guidance (e.g., that a collection must exist before the add/assign calls), but for a single-purpose create tool the routing is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collections_deleteADestructiveIdempotentInspect
Delete a knowledge collection.
If the collection is assigned to agents, prompts, or channels, pass force=true to delete anyway. CASCADE removes all assignments automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | Force delete even if collection is in use. OMIT for the safe default (refuse to delete in-use collections). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| collection_id | Yes | ID of the collection to delete |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so safety is covered. The description adds real behavioral context beyond that: an in-use collection is refused by default and force=true overrides it, and CASCADE silently removes all assignments. It does not describe auth requirements or a return payload, but the key destructive consequence is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the core action leads. The force condition and the CASCADE consequence are packed into the follow-up without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden, but for a delete action the outcome (removal plus cascade of assignments) is the essential information and it is present. Annotations cover the safety profile, making this largely complete for the agent's decision.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description goes further by explaining the 'why' behind force=true (in-use by agents/prompts/channels) rather than merely restating the schema's 'OMIT for the safe default', adding meaningful semantic weight.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Delete) and resource (knowledge collection) in the opening sentence. An agent can immediately distinguish it from sibling management tools like collections_remove_file, collections_remove_website, or collections_unassign_agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when to pass force=true (collection assigned to agents, prompts, or channels), which is the key conditional for this action. It stops short of naming alternatives, but no sibling offers an equivalent collection-deletion path, so the guidance is contextually clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collections_listARead-onlyIdempotentInspect
List all knowledge collections in the workspace.
Collections are named groups of files used for RAG search. Auto-created collections (per-agent, per-prompt) are hidden by default.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| include_inactive | No | Include inactive collections. OMIT to list only active collections (the default). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description still adds real behavioral context beyond that: auto-created per-agent/per-prompt collections are hidden by default, which is not stated anywhere in the schema. It omits pagination/ordering, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and scope, followed by the one piece of context an agent needs to interpret the result. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, no-required-parameter listing tool with full schema coverage and no output schema, this is nearly complete: it explains what is listed and what is filtered out. Only the return shape (fields per collection) is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented inline, making 3 the baseline. The description's 'hidden by default' note touches the include_inactive behavior but does not clarify the distinction between inactive and auto-created-hidden collections, so it adds little beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('knowledge collections in the workspace') and then defines what a collection actually is ('named groups of files used for RAG search'). This makes it unambiguous versus mutation siblings like collections_create, collections_delete, and collections_add_file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: it conveys the default scope (auto-created collections hidden) but never states when to reach for this tool versus collections_list_files or collections_list_websites, nor any prerequisite. Adequate but leaves selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collections_list_filesARead-onlyIdempotentInspect
List all files in a knowledge collection with their indexing status and chunk counts. Each returned file has a file_id (integer) that can be passed to messages.send as attachments=[file_id] to send the file to a contact, or to files.read to read its text content.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| collection_id | Yes | ID of the collection |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is low. The description adds real value beyond them: it discloses what is returned (indexing status, chunk counts, file_id) and how the file_id plugs into messages.send and files.read. No auth or pagination detail is mentioned, but the return/behavior disclosure is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, and the second sentence's clauses each earn their place by explaining how to use the returned file_id. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description correctly describes return contents (indexing status, chunk counts, file_id) and integration points. For a two-parameter read tool with full schema coverage and annotations, nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both collection_id and in_workspace are already documented in the schema. The description adds no parameter syntax or format detail, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (List) plus resource (files in a knowledge collection) plus scope (with indexing status and chunk counts). An agent can tell this apart from collections_list (which lists collections) or files.read (which reads content) without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear downstream routing: file_id can be passed to messages.send as attachments or to files.read. It does not, however, state exclusions or explicitly contrast this with sibling tools like collections_list_websites. Clear context, no explicit when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collections_list_websitesARead-onlyIdempotentInspect
List the websites of a knowledge collection with indexing status, pages indexed / found, last indexed time and last error.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| collection_id | Yes | ID of the knowledge collection |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so the safety profile is covered. The description adds genuinely useful behavioral context by disclosing what each row contains (indexing status, pages indexed/found, last indexed time, last error), which matters because no output schema exists. It omits pagination/result-size behavior, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that names the action first and appends the returned fields. Every clause carries information; there is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with full annotation coverage and a 100% documented schema, the description is nearly complete, and it compensates for the absent output schema by listing the returned fields. It is only slightly short on operational details such as ordering, pagination, or how often indexing status updates.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (collection_id and in_workspace) are fully documented in the schema. The description only echoes the collection scoping implicitly and adds nothing about parameter format or side effects, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List the websites of a knowledge collection') and enumerates the fields it returns, which implicitly separates it from siblings like collections_list_files and collections_add_website. It does not name any sibling explicitly, so it stops short of full differentiation, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the way to inspect a collection's websites, versus collections_add_website or collections_refresh_website. There is no statement of when to prefer this call, no prerequisite (e.g. the collection must exist or that indexing is asynchronous), and no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collections_refresh_websiteAInspect
Queue a website of a collection for re-indexing now. Unchanged pages are skipped.
| Name | Required | Description | Default |
|---|---|---|---|
| source_id | Yes | ID of the website source (from collections.list_websites) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| collection_id | Yes | ID of the knowledge collection |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false and destructiveHint=false. The description adds real behavioral context beyond that: 'Queue ... now' signals an asynchronous enqueue rather than a synchronous re-index, and 'unchanged pages are skipped' tells the agent work may be partially skipped. It still omits auth requirements and how completion is observed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and resource and followed by the one behavioral caveat. No filler; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation tool with no output schema and adequate annotations, the description covers purpose and the key skip behavior. It does not say that the call returns immediately after queueing or how the agent learns when re-indexing finishes, which is a meaningful gap for a queued operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with all three parameters documented in the schema itself (including the source_id pointer to collections.list_websites). The description adds no parameter detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: queue a collection's website for re-indexing. 'Re-indexing' implies the website source already exists, which implicitly separates it from collections_add_website, but no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative is given. The agent must infer from 're-indexing' that a website must already have been added via collections_add_website before this call makes sense; no prerequisites are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collections_remove_fileAInspect
Remove a file from a knowledge collection.
The file itself is not deleted — only the collection membership is removed.
| Name | Required | Description | Default |
|---|---|---|---|
| file_id | Yes | ID of the file to remove | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| collection_id | Yes | ID of the collection |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=false and readOnlyHint=false. The description adds genuinely useful context by explaining exactly what is and is not destroyed — the file survives, only membership is removed. It does not cover auth needs or idempotency (annotation says idempotentHint=false), so it stops short of full disclosure but clearly adds value beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and immediately followed by the most important clarifying constraint. No filler; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple membership-removal tool with full schema coverage and annotations covering the safety profile, the definition is nearly complete, and the key behavioral nuance (file not deleted) is present. It is slightly thin on edge cases (e.g., behavior if the file is not in the collection) but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so file_id, collection_id, and in_workspace are already documented in the schema. The description adds no parameter-level meaning beyond what the schema provides, making the baseline of 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (remove a file from a knowledge collection) and names the exact scope of the effect. An agent can distinguish it from siblings like collections_add_file, collections_remove_website, and collections_delete without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clarifies the operation's scope (membership only, not deletion), which implicitly helps route away from a true file-delete tool. However, it never explicitly states when to use this versus alternatives such as files_delete or collections_delete, nor any preconditions. Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collections_remove_websiteADestructiveIdempotentInspect
Remove a website from a knowledge collection and DELETE every page file indexed from it. Cannot be undone.
| Name | Required | Description | Default |
|---|---|---|---|
| source_id | Yes | ID of the website source (from collections.list_websites) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| collection_id | Yes | ID of the knowledge collection |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is partly covered. The description goes beyond them by specifying exactly what is destroyed (every page file indexed from the website) and reinforcing irreversibility, which is the highest-value context for this call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the action and its destructive consequence front-loaded before the irreversibility warning. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description plus annotations cover what the agent needs: the action, its side effects on indexed page files, and that it cannot be undone. Minor gaps remain around post-removal collection state and auth requirements, but nothing essential for invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with source_id even pointing at collections.list_websites, so the schema carries the parameter burden. The description adds no syntax or format detail beyond it, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Remove a website from a knowledge collection') and adds the crucial scope detail that every page file indexed from the website is also deleted. That detail cleanly separates it from the sibling collections_remove_file, which removes a single file rather than an entire website source.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The destructive scope and 'Cannot be undone' warning imply this is the terminal cleanup action, but there is no explicit when-to-use guidance, no mention of the sibling collections_remove_file, and no note on what happens to the collection if this was its last source. Usage must be inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collections_unassign_agentAInspect
Remove a knowledge collection from an AI agent.
The collection and its files are not deleted — only the agent assignment is removed.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | ID of the AI agent | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| collection_id | Yes | ID of the collection to unassign |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false. The description adds genuine value beyond them by clarifying scope: the collection and its files survive, only the assignment is removed. It does not cover auth requirements or whether re-running is safe, keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed by the key scope caveat. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With full schema coverage, annotations covering safety hints, and no output schema, the description supplies the one piece of context an agent most needs: that this is not a destructive delete. It is complete enough to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (agent_id, collection_id, in_workspace) are already documented in the schema. The description adds no additional meaning about parameter usage or format, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: 'Remove a knowledge collection from an AI agent.' This clearly identifies the assignment relationship being severed. The second sentence distinguishes it from a delete operation, though it does not name the inverse sibling (collections_assign_agent) explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use clause or named alternative. The clarification that deletion is not involved implicitly tells the agent this is the right tool when only the assignment should go, but usage must be inferred rather than read.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_add_channelAInspect
🔗 Link a channel identity (email, phone, Telegram/WhatsApp handle) to an existing contact.
When to use:
User learns a contact's email or phone and wants to save it
Adding a second channel for an existing person
Email/phone are stamped as the contact's identity keys, so channel sync attaches the matching conversation to the same contact automatically. If the identity already belongs to a DIFFERENT contact, the call fails unless merge_if_linked=true — which IRREVERSIBLY merges the two contacts into one. Requires contact_id (entity_id) from contacts.find.
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | Email address, phone number, or username for this channel | |
| channel | Yes | Channel type to add | |
| contact_id | Yes | entity_id from contacts.find/contacts.sync; when sync returned no entity_id, pass its person_id instead | |
| display_name | No | Optional display label for this identity | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| merge_if_linked | No | If the identity already belongs to a DIFFERENT contact, merge that contact into this one (IRREVERSIBLE — the other contact record is deleted). Default false: the call fails instead of merging. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already declaring readOnly=false/destructive=false, the description still adds substantial context: identity-key stamping, automatic channel sync, fail-by-default behavior, and the irreversible merge. Note the merge path deletes a contact record while destructiveHint=false — the description is the more accurate source, so the annotation is arguably under-declared rather than contradicted (the destructive action is opt-in via merge_if_linked).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, followed by tight when-to-use bullets, then a clearly delineated caveat paragraph. Every sentence carries weight and the most dangerous behavior (IRREVERSIBLE merge) is called out rather than buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with no output schema, the description covers purpose, usage, identity semantics, and the critical merge caveat thoroughly. The only gap is that it does not say what the call returns on success (e.g., updated contact/identity), which is a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all six parameters including merge_if_linked and contact_id sourcing. The description reinforces identity-key meaning and the contacts.find provenance, but adds little the schema lacks, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb (link/add) and resource (channel identity: email, phone, Telegram/WhatsApp handle) scoped to an existing contact. It clearly distinguishes this from generic contact mutation siblings like contacts_update or contacts_merge by focusing on attaching a single identity channel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The explicit 'When to use' bullets give real context (learning an email/phone, adding a second channel), and the description explains the merge-vs-fail fork. It stops short of naming alternative tools for adjacent cases (e.g., contacts_update for other fields), so it is clear rather than exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_capture_leadAInspect
📝 Save the current contact's details as a structured lead (contact + Contacts tab).
When to use:
In a website-chat OR voice-call conversation, as soon as the person has given at least ONE way to reach or identify them — a name, a phone number or an email (plus company / use case if mentioned) — e.g. when booking a demo or a class. A phone number with no name is still a lead: capture it, then ask for the name if you need it.
Call it once you have the details; then continue (e.g. share the booking link).
Creates/links a contact record and a lead entry in the workspace's lead inbox. Works in any conversation thread (website chat, voice call, DM).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | The visitor's full name, if they gave it. Omit when unknown — a phone or email is enough on its own. | |
| No | The visitor's email address. | ||
| phone | No | Phone number, if the visitor provided one. | |
| company | No | The visitor's company / organization, if mentioned. | |
| use_case | No | Their main use case / what they want to do with DialogBrain, if mentioned. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the safety profile is supplied. The description adds the substantive side effects beyond the annotations: it creates/links a contact record plus a lead entry in the lead inbox, and works across any thread type. It omits the duplicate-creation risk implied by idempotentHint=false, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is front-loaded with the core action before the 'When to use' list, and every sentence carries information (trigger, counter-example, follow-up). The bullets keep it scannable despite covering several scenarios, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering the mutation profile and no output schema required, the description is nearly complete for a capture tool: it explains the trigger, what is created, and cross-field requirements. The only meaningful gap is that it does not describe what the call returns or how duplicates are handled (relevant given idempotentHint=false).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each field is already documented, making 3 the baseline. The description still adds cross-parameter meaning the per-field schema does not: at least one of name/phone/email suffices, name can be omitted, and company/use_case are optional context. That is genuine semantic value above the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first line gives a specific verb and resource: 'Save the current contact's details as a structured lead (contact + Contacts tab),' and later clarifies it 'Creates/links a contact record and a lead entry in the workspace's lead inbox.' The purpose is unambiguous, but it never explicitly contrasts itself with the many contacts_* siblings (contacts_update, contacts_find, contacts_merge), so it earns a 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
A dedicated 'When to use' block states the triggering context (website-chat or voice-call conversation, as soon as one identifier is given) with concrete examples (booking a demo/class) and even handles the edge case 'a phone number with no name is still a lead.' It stops short of naming when NOT to use it or pointing to a relevant alternative like contacts_update, so it is strong but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_discoverARead-onlyIdempotentInspect
Search for a contact on a live channel (Telegram, WhatsApp, etc.) before adding them. Use this to look up a person by username or phone number before calling contacts.sync. This is the right tool when asked to add or find a specific person by @username or phone (use contacts.sync afterwards to actually add them) — not group_discovery.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Username, phone, or name to search for | |
| channel | Yes | Channel name: telegram, whatsapp, etc. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds that the search happens on a live channel and is a lookup that precedes an add, but says nothing about rate limits, behavior on no match, or result shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core instruction is front-loaded well, but the parenthetical instructs the same thing twice ('before calling contacts.sync' and 'use contacts.sync afterwards to actually add them') and repeats the 'before adding' idea, wasting the middle of the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a live-channel search with no output schema, the description should signal what a successful lookup yields or what happens when nothing is found; neither is covered. Annotations plus the sync follow-up make it usable, but an agent cannot predict the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents query, channel and in_workspace; baseline is 3. The description adds only that query is a username or phone number (schema also allows a name), which is marginal value over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
State a specific verb (search/look up) plus resource (a person on a live channel by username or phone) and explicitly routes away from group discovery. However it never distinguishes itself from the near-identical sibling contacts_find, which an agent could easily confuse with this lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use ('before calling contacts.sync', 'when asked to add or find a specific person by @username or phone') and names contacts.sync as the follow-up action plus an exclusion against group discovery. The absence of any guidance versus contacts_find is the one meaningful gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_findARead-onlyIdempotentInspect
👤 Search for contacts in your address book by name or username.
When to use:
User asks 'find contact X' or 'who is Y?'
User wants to know someone's username or ID
Before sending a message to verify contact exists
To get contact's channel reference for messaging
Examples: ❓ User: 'find contact named [name]' → contacts_search(query='[name]', limit=5)
❓ User: 'who is [full name]?' → contacts_search(query='[full name]', limit=1)
❓ User: 'search for @username' → contacts_search(query='username', limit=10)
Returns: name, username, channel, channel_ref, similarity_score, match_type. Plus:
entity_id: local DB key — pass to contacts.profile. Null for live-discovered contacts (skip contacts.profile for those).
telegram_user_id (when channel='telegram'): the Telegram user ID — pass to calls.make / messages.send. NOT entity_id.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return | |
| query | Yes | Name or username to search for (supports partial matches) | |
| channel | No | Filter by channel. OMIT to search across all channels. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds genuinely useful behavioral context the annotations cannot: the entity_id vs telegram_user_id distinction, which consumer each ID feeds (contacts.profile vs calls.make/messages.send), and that entity_id is null for live-discovered contacts. It does not cover edge cases like zero-result behavior or result ranking beyond the returned similarity_score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the one-line purpose, then cleanly sectioned into 'When to use', 'Examples', and 'Returns' — every block is useful. It is somewhat verbose and the examples invoke a stale tool name (contacts_search) that does not match this tool (contacts_find) or any sibling, which is a minor structural flaw rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the full burden of explaining returns, and it does so thoroughly: field list plus precise downstream routing semantics for entity_id and telegram_user_id. Combined with the usage triggers, an agent has everything needed to call and consume this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents query, limit, channel, and in_workspace, making 3 the baseline. The examples do implicitly demonstrate limit sizing (limit=5/1/10) and partial-match querying, but the description never explains the channel filter or in_workspace parameter, so it adds only marginal meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search for contacts in your address book by name or username'), with scope ('by name or username') that distinguishes it from contacts_discover (find new contacts), contacts_profile (fetch one contact), and contacts_research. An agent can select it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
An explicit 'When to use' block gives four concrete triggers (find contact X, who is Y, verify before messaging, get channel reference), and the worked examples map natural-language user intents to concrete calls with parameters. This is stronger routing guidance than most sibling tools provide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_mergeAInspect
🧬 Merge two contacts into one: all channel identities, scores and summaries of the second contact move onto the first, and the second contact record is deleted.
When to use:
contacts.find shows the same person twice (e.g. a messenger contact and an email contact)
The user explicitly asks to merge two specific contacts
Both ids are entity_id values from contacts.find. The merge is recorded in Review Duplicates and can be undone there (undo is heuristic, not guaranteed). Never merge on name similarity alone — confirm the pair is really one person first.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| contact_id_keep | Yes | entity_id (from contacts.find/sync) of the contact to KEEP; a person_id is accepted when sync returned no entity_id | |
| contact_id_merge | Yes | entity_id (from contacts.find/sync) of the contact to merge INTO the kept one; a person_id is accepted too |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is rich about behavior — what moves, that the second record is deleted, that the merge is logged in Review Duplicates and undo is 'heuristic, not guaranteed'. However, it directly conflicts with the structured annotation destructiveHint=false while explicitly describing a record deletion, so the agent gets contradictory signals about safety. Credit for the substantive disclosure, penalty for the conflict.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the operation and its effect, then a scannable 'When to use' block, then the id-source and reversibility caveats. Every sentence carries new information (moved fields, deletion, undo location, name-similarity warning); nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers what changes, the reversibility path, and how to source the ids. It stops short of stating return/confirmation shape and permission or workspace prerequisites (the cross-workspace in_workspace behavior lives only in the schema), which keeps it just under complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three parameters (in_workspace, contact_id_keep, contact_id_merge) are already documented. The description restates that both ids are entity_id values from contacts.find and notes person_id fallback, adding little beyond the schema; the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (merge two contacts) and spells out the exact semantics: channel identities, scores and summaries move to the kept contact, and the second record is deleted. This is clearly distinguishable from siblings like contacts_update, contacts_find, and contacts_add_channel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Has an explicit 'When to use' list (duplicate shown by contacts.find, or explicit user request), names contacts.find as the id source, and adds a concrete exclusion: 'Never merge on name similarity alone — confirm the pair is really one person first.' Both the trigger and the anti-trigger are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_profileARead-onlyIdempotentInspect
👤 Get full profile for a contact: all channel identities, notes, role, capabilities, birthday.
When to use:
After contacts.find to get complete info about a specific person
To see all channels a contact is reachable on
To read notes, role, or capabilities for a contact
Requires contact_id (entity_id) from contacts.find.
| Name | Required | Description | Default |
|---|---|---|---|
| contact_id | Yes | entity_id from contacts.find | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds real value by disclosing the returned fields, which matters because there is no output schema; it does not contradict any annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line summary followed by a scannable bulleted 'When to use' list and a single prerequisite line. Every sentence carries information; nothing is redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a simple two-parameter read tool whose safety profile is fully covered by annotations, the description supplies the return-field summary that the missing output schema would otherwise provide, plus the id's source. Nothing needed to call it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds provenance for the required id ('entity_id from contacts.find'), telling the agent where the value must come from rather than just its type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get full profile for a contact') and enumerates the payload (channel identities, notes, role, capabilities, birthday), so an agent can distinguish it from contacts_find without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
A dedicated 'When to use' block gives three concrete scenarios and an explicit prerequisite (run after contacts.find, using the id it returns). It stops short of naming a when-not or an alternative tool, but the routing context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_researchARead-onlyIdempotentInspect
Research a person before outreach: returns a synthesized profile (current role, company, location, career history, education, LinkedIn/social URLs) plus a candidates array for disambiguating namesakes. Use to personalize a first touch or brief before a meeting. Does NOT return contact channels (email/phone/telegram) — use contacts.discover to add a reachable channel, or the LinkedIn URL from the result for a connection request.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Who to research: full name plus any disambiguating context you have (company, role, location), e.g. '[full name], founder of [company]'. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safe-read profile (readOnly, idempotent, openWorld, non-destructive), so the bar is lower; the description adds real value by disclosing the return shape and the namesake-disambiguation behavior via the candidates array. It stops short of mentioning cost, latency, or caching behavior for an open-world lookup, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three clauses, zero waste: purpose and return shape front-loaded, then the negative boundary and its alternative. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return-value burden and does so by enumerating the profile fields and the candidates array. Combined with the explicit boundary about contact channels and the usage triggers, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both query and in_workspace with examples. The description adds no syntactic or formatting detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (research) and resource (a person) and enumerates the synthesized fields returned (role, company, location, career history, education, social URLs) plus the disambiguation candidates array. It also differentiates itself from siblings by explicitly naming contacts.discover and contrasting the return shape.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use triggers (before outreach, to personalize a first touch or brief before a meeting) and an explicit when-not (does not return contact channels), routing the agent to contacts.discover for a reachable channel. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_syncAInspect
Add a discovered contact and open a conversation thread. Returns thread_id for the new conversation, plus entity_id — the canonical handle to pass as contacts.update's contact_id. Call contacts.discover first to verify the contact exists.
| Name | Required | Description | Default |
|---|---|---|---|
| channel | Yes | Channel name: telegram, whatsapp, etc. | |
| identifier | Yes | Username or phone number to add | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| from_account_id | No | Which of the workspace's accounts on this channel opens the conversation — i.e. WHO the contact will see as the sender. OMIT and the platform picks the account that already holds most of this workspace's threads. Pass it whenever the sender matters: a workspace with several numbers on one channel usually has one that must not start conversations (a personal line, or one under a spam restriction). An account that is not active in this workspace is refused, never silently swapped for another. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a mutating, non-idempotent, open-world operation, so the safety profile is covered. The description adds real value beyond that: it discloses the return values (thread_id, entity_id), the downstream linkage (entity_id is the handle for contacts_update's contact_id), and a precondition. It does not warn that repeat calls create duplicate threads, but that is a modest gap given the annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: action first, return values second, prerequisite third. Every sentence carries distinct information and nothing is redundant with the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining returns and does so adequately (thread_id, entity_id plus its downstream use). Prerequisites and the mutation profile are covered by annotations, but the description is silent on whether opening a conversation actually sends a visible message to the contact — relevant given openWorldHint=true.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so channel, identifier, in_workspace, and from_account_id are already fully documented in the schema, including the detailed from_account_id sender-selection guidance. The description adds no parameter-level detail beyond the return-value linkage, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Add a discovered contact and open a conversation thread.' The phrase 'discovered contact' plus the reference to contacts.discover implies this registers a pre-existing contact rather than creating one from scratch, partially distinguishing it from siblings like contacts_add_channel or contacts_capture_lead, though it never names those alternatives directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear ordered prerequisite: 'Call contacts.discover first to verify the contact exists.' That tells the agent when this tool is appropriate in the workflow. It stops short of stating when NOT to use it or naming which sibling to prefer for other contact operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_updateAInspect
✏️ Update a contact's profile: name, notes, role, capabilities, birthday, preferred channel.
When to use:
User wants to add notes about a contact
User wants to set/update role or capabilities for a contact
User wants to rename a contact or update birthday
Requires contact_id — the entity_id returned by contacts.find or contacts.sync. At least one optional field must be provided.
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | Contact role (e.g. developer, client, partner). Empty string clears role. | |
| notes | No | Free-text notes/context about this contact. Empty string clears notes. | |
| contact_id | Yes | entity_id from contacts.find or contacts.sync | |
| birthday_day | No | Birth day 1-31 (must be set together with birthday_month) | |
| capabilities | No | List of capabilities (e.g. ['backend', 'design']) | |
| display_name | No | New display name (max 255 chars) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| birthday_year | No | Birth year 1900-2100 (optional, standalone) | |
| birthday_month | No | Birth month 1-12 (must be set together with birthday_day) | |
| preferred_channel | No | Preferred channel for contacting this person. OMIT to leave the preferred channel unchanged. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-destructive, non-idempotent, closed-world mutation, so the safety profile is covered. The description adds one genuine behavioral rule not present in the annotations ("At least one optional field must be provided") and partial-update framing, but says nothing about permissions, whether the update is reversible, or how unspecified fields behave beyond what the schema already states.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the verb and affected fields, then a scannable "When to use" list, then the required-parameter constraint — a sensible ordering with no filler prose. The field enumeration in the opening sentence is partially duplicated by the bullets below, and the leading emoji is decorative, keeping it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter partial-update tool with annotations covering the safety profile and no output schema, the description supplies the essentials: what is mutable, how to obtain contact_id, and the at-least-one-field rule. It does not describe the response or failure modes (e.g., invalid or unknown contact_id), which is a minor gap given no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description's field list and the contact_id provenance note largely restate what the schema already documents, including the clearing semantics ("empty string clears") and the omit-to-leave-unchanged rule for preferred_channel. No additional syntax, coupling, or interaction detail is contributed beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ("Update a contact's profile") and enumerates the mutable fields (name, notes, role, capabilities, birthday, preferred channel), which separates it from sibling writers like contacts_add_channel or contacts_merge. It also pins down the entity being mutated via the contact_id provenance note. It stops short of an explicit "not this tool, use X instead" statement, so it lands just under top marks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
A dedicated "When to use" block lists three concrete user intents (add notes, set role/capabilities, rename or set birthday), and a precondition is stated: contact_id is required and at least one optional field must be supplied. What's missing is the negative side — no guidance on when to prefer contacts_merge, contacts_profile, or contacts_capture_lead over this update.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
db_createAInspect
Provision this workspace's own database (one per workspace). Returns the connection string and password ONCE — relay them to the user immediately; they cannot be retrieved later. Errors if the workspace already has a database. Use db.schema / db.query / db.execute to work with it afterwards.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only convey the safety profile (readOnly=false, idempotent=false, destructive=false). The description adds substantially more: the return value contains a connection string and password revealed exactly ONCE and unrecoverable, plus the failure mode when a database already exists. That one-time-secret disclosure is critical behavioral context an agent could not get from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each carrying distinct load: purpose, the irreversible credential handoff, and the error condition plus follow-up tools. The most urgent information (relay the password immediately) is surfaced rather than buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument mutation tool with no output schema, the description covers everything needed: what is created, what comes back, that it cannot be re-fetched, when it fails, and which tools to use next. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single optional parameter and schema description coverage is 100%, with the schema text itself explaining the workspace override and its non-persistent effect. The description adds no parameter-level meaning, so the baseline 3 for fully-documented schema is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Provision this workspace's own database') with an explicit cardinality constraint (one per workspace), which immediately separates it from the sibling read/write tools db_query, db_execute and db_schema. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear precondition ('Errors if the workspace already has a database') and routes the agent to the correct follow-up tools (db.schema / db.query / db.execute). What it lacks is an explicit when-to-use-vs-workspace_create or a 'use this only if you have not yet provisioned' framing, so the routing is strong but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
db_executeAInspect
Run SQL (CREATE TABLE / INSERT / UPDATE / DELETE / ALTER) against this workspace's own database. Use db.schema first to see what exists, and db.query to read.
| Name | Required | Description | Default |
|---|---|---|---|
| sql | Yes | SQL DDL/DML, e.g. CREATE TABLE records (id serial primary key, name text) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations declare destructiveHint=false, yet the description explicitly advertises DELETE and ALTER — irreversible data loss and schema mutation against a live database. An agent reading both gets directly conflicting safety signals about whether this call can destroy data, which is the highest-stakes fact for this tool. The dot-notation references ("db.schema", "db.query") also do not match the real sibling names db_schema/db_query.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the operation scope front-loaded ahead of the routing advice. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter write tool with no output schema and full schema coverage, the description covers purpose, scope and sequencing adequately; return behavior need not be explained. It stops short of the completeness bar because it leaves the destructive/irreversible consequences of DELETE/ALTER unstated and conflicts with the safety annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are documented in the input schema, including the in_workspace isolation semantics ("Nothing is stored; other sessions are not affected"). The description adds no parameter-level detail of its own, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Run SQL) plus the exact resource and scope: DDL/DML against "this workspace's own database," and enumerates the operations (CREATE TABLE / INSERT / UPDATE / DELETE / ALTER). It also positions itself against db_query (read) and db_schema (inspect), so an agent can distinguish it from its closest siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit sequencing and routing: "Use db.schema first to see what exists, and db.query to read." This is a when-to-use and when-to-use-something-else statement in one sentence, which is exactly what an agent needs to pick between db_execute, db_schema and db_query.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
db_queryARead-onlyIdempotentInspect
Read from this workspace's own database. Pass either a natural-language question (e.g. 'how many records were added this week') or a raw SQL SELECT. Read-only; returns rows plus a chart hint.
| Name | Required | Description | Default |
|---|---|---|---|
| sql | No | Raw SQL SELECT (alternative to question) | |
| question | No | Natural-language question about the data | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the bar is lower, yet the description adds real value: it confirms read-only behavior, restricts input to a SELECT, and discloses the return shape ('rows plus a chart hint'). It stops short of noting rate limits, permissions, or result limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with zero filler: destination, input modes, and constraints/returns in order of importance. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param read-only tool with no output schema, the description covers purpose, input modes, read-only constraint, and a rough return shape. It is nearly complete, lacking only explicit differentiation from sibling data tools and any limits on result size.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning by framing sql and question as mutually exclusive alternatives and giving a concrete example of the natural-language question. The workspace-scoping nuance of in_workspace is left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (read) and scoped resource (this workspace's own database), and clarifies the two accepted input modes. It implicitly contrasts with the write-oriented siblings db_execute/db_create via 'Read-only', but never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent how to call it ('Pass either a natural-language question ... or a raw SQL SELECT') but gives no when-to-use vs alternatives guidance, e.g. when to prefer db_query over analytics_query, db_schema, or db_execute. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
db_schemaARead-onlyIdempotentInspect
List the tables and columns of this workspace's own database. Call this before db.query or db.execute to learn what exists.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered externally. The description adds genuine context beyond that: the workspace-scoped read behavior ('this workspace's own database') and the recommended sequencing relative to db.query/db.execute. It does not describe the return shape or any size limits, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the purpose front-loaded ahead of the routing advice. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-param, read-only introspection tool with full annotation coverage and full schema coverage, the description is nearly sufficient. The only mild gap is that no output schema exists and the description does not sketch the return shape (e.g., nested tables/columns structure), though 'tables and columns' implies it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single optional param in_workspace is already well documented in the schema (including its non-persistent, session-isolated behavior). The description adds no parameter-level detail, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (tables and columns of this workspace's own database), and the phrase 'this workspace's own database' disambiguates it from the sibling management tools db_create/db_query/db_execute. An agent can tell immediately what this returns and how it differs from the DML/DQL siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternatives ('Call this before db.query or db.execute') and states the condition that selects it ('to learn what exists'). This is textbook when-to-use guidance pointing at the exact sibling tools an agent would otherwise reach for first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
documents_createAInspect
Render a document (PDF / HTML / PPTX / DOCX) and save it to the workspace.
This tool has two input pipelines — pass exactly one of content_html or content_markdown.
Pipeline A — content_html (canonical for decks, proposals, designed pages)
You author full HTML+CSS. A baked-in design-system preamble ships first
(<style> with Inter/Manrope as data-URI fonts, CSS-variable palette tokens,
8px spacing scale, and pre-styled layout helpers); your markup and any of
your own <style> blocks land after the preamble so you can override
anything. Chromium renders the assembled document into a static PDF —
JavaScript is disabled and DNS is blackholed, so external font / image /
script fetches will fail by configuration.
Required when this pipeline is used:
title— human-readable, used for PDF metadata and the saved filename.content_html— the<body>and any custom<style>blocks. The renderer wraps this in<html>…</html>and injects the preamble + a canonical<meta charset>+<title>. Do NOT emit<script>,<iframe>,<object>,<embed>,<meta>,<link>,<base>,<form>, or event handlers — the sanitizer strips them.output_type—"pdf"or"html". ("pptx"and"docx"requirecontent_markdownsince they need structured markdown intermediates.)
Optional:
page_preset—"slide_16_9"(default for any deck),"a4"(default for flowing documents — used if omitted),"letter", or"none"(you declare your own@pagerule). For a web-styled page (dark background, full-bleed sections) use"none"and declare@page { margin: 0 }, set the background onhtmlas well asbody, and addprint-color-adjust: exact— the a4/letter presets keep 24mm paper margins, which paint as a white frame around dark designs.design_tokens— flat dict overriding the preamble's CSS variables. Whitelisted keys:brand_primary,accent,surface_dark(hex color),font_display,font_body(font name from ['Inter', 'Manrope', 'monospace', 'sans-serif', 'serif', 'system-ui', 'ui-monospace', 'ui-sans-serif', 'ui-serif']).language— BCP-47 tag (default"en"). Drives<html lang>.
Slide structure (page_preset="slide_16_9")
Each slide is <section class="slide …">…</section>. The base .slide
class is what sizes it to the viewport and forces the page break — do
not drop it. Composable variants (apply alongside .slide):
.slide-cover— gradient hero, big display title..slide-split— two equal columns, image + narrative..slide-stats— three-up KPI cards (use<div class="stat">with.stat-value+.stat-labelinside)..slide-quote— centered pull quote +<cite>attribution.
Layout helpers (work in any preset): .grid-2, .grid-3, .split,
.stack, .cluster, .callout, .muted, .kbd.
Speaker notes
<aside class="notes">…text…</aside> inside a <section class="slide">.
The sanitizer strips them from the rendered PDF and returns them as
slide_notes[] (parallel to slide order). Orphan notes outside any slide
are dropped with a warning.
Images
Only these src schemes resolve:
file:NNN— workspacefile_id.data:image/...;base64,...— inline.https://<host>where<host>∈DOCUMENTS_MEDIA_URL_ALLOWLIST. Other URLs are dropped and replaced with an HTML comment placeholder.
Pipeline B — content_markdown (invoice / contract only)
Required:
title,content_markdown,output_type.
Optional:
theme—"invoice"or"contract". Triggers the corresponding exemplar styling and (for invoices) the arithmetic validator that fail-closes on missing or mismatched totals.language— BCP-47 (default"en").
Delivery contract (CRITICAL)
After this tool returns file_id, deliver the file with
messages.send(attachments=[file_id], text="<short caption>"). Embedding
the file_id in a markdown link, sandbox: URL, or /api/files/<id>/download
text will render as plain text on the recipient's channel — the
attachments parameter is the only way the file actually attaches.
Exemplars
INVOICE (English):
Invoice INV-{YYYYMMDD-HHMMSS}
From: {Issuer Legal Name}, {Address}, {Tax ID} To: {Customer Name}, {Customer Address}, {Customer Tax ID} Issue date: {YYYY-MM-DD} Due date: {YYYY-MM-DD}
Description | Qty | Unit price | Total |
{Service 1} | 1 | 1500.00 | 1500.00 |
{Service 2} | 2 | 500.00 | 1000.00 |
Subtotal: USD 2500.00 Tax (20%): USD 500.00 Total: USD 3000.00
Payment: {bank details OR crypto wallet — never both}
INVOICE (Russian):
Счёт-фактура № INV-{YYYYMMDD-HHMMSS}
От: {Юридическое название организации}, {Адрес}, ИНН {Tax ID} Кому: {Название клиента}, {Адрес клиента}, ИНН {Tax ID} Дата: {YYYY-MM-DD} Срок оплаты: {YYYY-MM-DD}
Описание | Кол-во | Цена | Сумма |
{Услуга 1} | 1 | 1500.00 | 1500.00 |
{Услуга 2} | 2 | 500.00 | 1000.00 |
Подытог: USD 2500.00 НДС (20%): USD 500.00 Итого: USD 3000.00
Реквизиты: {банковские реквизиты ИЛИ криптокошелёк — не оба сразу}
CONTRACT (English):
Service Agreement
Between: {Provider Legal Name}, {Address} ("Provider") And: {Client Legal Name}, {Address} ("Client") Effective date: {YYYY-MM-DD}
1. Scope of services
{Concise description of what Provider agrees to deliver.}
2. Term
This Agreement begins on the Effective date and continues until {termination condition or end date}.
3. Compensation
Client pays Provider {amount and currency} according to {payment schedule}.
4. Confidentiality
Both parties agree to keep proprietary information of the other party confidential during and after the term of this Agreement.
5. Termination
Either party may terminate with {N} days' written notice.
6. Governing law
{Jurisdiction}.
Provider: ____________________ Client: ____________________ {Provider signatory name} {Client signatory name}
CONTRACT (Russian):
Договор оказания услуг
Между: {Юридическое название Исполнителя}, {Адрес} ("Исполнитель") И: {Юридическое название Заказчика}, {Адрес} ("Заказчик") Дата вступления в силу: {YYYY-MM-DD}
1. Предмет договора
{Краткое описание услуг, которые Исполнитель обязуется оказать.}
2. Срок действия
Договор вступает в силу с указанной даты и действует до {условие прекращения или дата окончания}.
3. Стоимость и порядок оплаты
Заказчик оплачивает услуги Исполнителя в размере {сумма и валюта} в порядке {график платежей}.
4. Конфиденциальность
Стороны обязуются сохранять конфиденциальность сведений, полученных в ходе исполнения настоящего Договора, в течение срока его действия и после его прекращения.
5. Расторжение
Любая из сторон вправе расторгнуть Договор, направив письменное уведомление не менее чем за {N} дней.
6. Применимое право
{Юрисдикция}.
Исполнитель: ____________________ Заказчик: ____________________ {ФИО подписанта Исполнителя} {ФИО подписанта Заказчика}
| Name | Required | Description | Default |
|---|---|---|---|
| theme | No | Invoice or contract styling for content_markdown. Rejected with content_html (use design_tokens + your own CSS instead). OMIT for default (unthemed) styling. | |
| title | Yes | Short human-readable title for the document. | |
| language | No | BCP-47 language tag (e.g. 'en', 'ru', 'zh', 'ja'). Drives <html lang> and (markdown path) font fallback for non-Latin scripts. | en |
| output_type | Yes | Renderer target: 'pdf' | 'pptx' | 'docx' | 'html'. PPTX/DOCX require content_markdown. | |
| page_preset | No | Page geometry for content_html. 'slide_16_9' = 1280x720 deck, 'a4'/'letter' = flowing document, 'none' = LLM declares its own @page. Defaults to 'a4' inside the html branch when omitted. Rejected with content_markdown. | |
| content_html | No | Full HTML body (with optional <style> blocks) for the canonical Chromium pipeline. Mutually exclusive with content_markdown. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| design_tokens | No | Flat dict of CSS-variable overrides for content_html. Whitelisted keys: brand_primary, accent, surface_dark (hex color), font_display, font_body (Inter|Manrope|system-ui|ui-sans-serif|ui-serif|ui-monospace|sans-serif|serif|monospace). Unknown keys / invalid values are dropped with a warning. Rejected with content_markdown. | |
| content_markdown | No | Markdown body for the invoice/contract pipeline. Mutually exclusive with content_html. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false and openWorldHint=false, so the write/safety profile is covered. The description adds genuine behavioral context beyond annotations: JS disabled and DNS blackholed, sanitizer strips script/iframe/link tags, image src allowlist, and speaker notes stripped from the PDF and returned as slide_notes[]. It does not, however, discuss rate limits, quotas, or cost, so it stops short of a 4-5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and pipeline rules are front-loaded, headers segment the material logically, and the exemplars genuinely earn their space as few-shot guidance. The definition is very long and the four localized exemplars make it heavier than strictly necessary, but almost every block adds actionable instruction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so: it tells the agent the tool returns file_id (and slide_notes[]), and it closes the loop with the attachment delivery contract. Given the two pipelines, nested design_tokens object, and four output formats, nothing an agent needs to call this correctly appears missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds real meaning on top: it explains why a4/letter presets paint a white 24mm frame around dark designs (motivating page_preset='none'), clarifies theme is rejected with content_html, and details the slide/notes/image conventions the schema only names. Some content (design_tokens whitelist, page_preset defaults) duplicates the schema, keeping it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb (render) plus the exact output formats and the side effect (save to workspace), so an agent instantly knows what the tool does. It is clearly distinguishable from siblings like artifacts_export_pdf or images_generate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Routes the agent explicitly: 'pass exactly one of content_html or content_markdown', with named conditions ('canonical for decks, proposals, designed pages' vs 'invoice / contract only'). It also spells out the required delivery step (messages.send with attachments=[file_id]) and warns what will NOT attach.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
feedback_saveAInspect
Save a behavioral rule, preference, or correction that should guide future agent behavior. Use this when the user gives explicit guidance like 'always reply in Russian', 'don't suggest meetings before 11am', or 'invoice link goes via email, not chat'. Structure the rule as: the rule itself, why it matters (if stated), and how to apply it. Scope: 'workspace' for org-wide rules, 'agent' for per-agent overrides, 'person' for per-contact preferences. Prefer feedback.save over notes.save for anything that's instructive rather than informational.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Short identifier for this rule (e.g. 'reply_language', 'meeting_hours'). Must not start with '__' (reserved). | |
| why | No | Why this rule matters (optional but recommended for the distiller). | |
| rule | Yes | The rule itself, in imperative form. Required. | |
| scope | Yes | Scope of the rule. 'workspace' for org-wide rules; 'agent' for per-agent overrides; 'thread' for conversation-specific guidance; 'person' for per-contact preferences. 'global' accepted as deprecation alias for 'agent'. | |
| how_to_apply | No | When/how to apply the rule (optional). Helpful for conditional rules like 'apply when speaking to Russian-speaking customers'. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| scope_ref_id | No | Required for scope='thread' (thread_id) and scope='person' (person_id). | |
| target_agent_id | No | Target agent. In agent mode optional (defaults to self); required from MCP. Ignored when scope='workspace'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the write/safety profile (readOnlyHint=false, destructiveHint=false), so the bar is lower. The description adds that saved rules persist and "guide future agent behavior," which clarifies the lasting side effect beyond the annotations. It does not address the non-idempotent hint (what happens on repeated saves), so it is not fully complete but adds real value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, followed by usage examples, structuring advice, scope, and sibling preference in a sensible order. Every sentence is useful, though the scope sentence partially duplicates the schema enum descriptions, adding mild redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description covers purpose, trigger conditions, structuring, and scoping, which is most of what an agent needs. It omits any note on the non-idempotent behavior and the in_workspace/target_agent_id semantics, though those are documented in the schema itself.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds structural guidance ("the rule itself, why it matters, and how to apply it") that maps to the rule/why/how_to_apply parameters, going beyond the schema's field-level text. However, the scope explanation largely restates the enum's own descriptions, limiting the added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Save") and resource ("behavioral rule, preference, or correction") with concrete examples, and explicitly contrasts with the sibling notes.save. An agent can tell exactly what this does and how it differs from alternatives without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger conditions ("when the user gives explicit guidance like 'always reply in Russian'") and names the alternative with a selecting condition ("Prefer feedback.save over notes.save for anything that's instructive rather than informational"). Both when-to-use and when-to-use-something-else are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_complete_uploadAInspect
Finish an upload started with files_create_upload_url, after the bytes have been PUT to the upload_url.
Verifies the object actually landed and matches the size (and sha256, when one was declared), then makes the file usable: from here file_id works in messages_send, messages_send_email attachments, documents.create, and the rest.
Returns: file_id, status, byte_size, mime_type.
| Name | Required | Description | Default |
|---|---|---|---|
| file_id | Yes | file_id returned by files_create_upload_url | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly=false, destructive=false, idempotent=false), and the description adds real substance beyond them: it verifies the object landed and matches declared size/sha256, and it explains that the file only becomes usable afterwards. The remaining gap is failure behavior — what happens if verification fails (error? partial object?) and whether re-invoking after a failed verify is safe, which matters given idempotentHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and its precondition, then the verification behavior, then the downstream usability, then a compact Returns line. With no output schema, the Returns enumeration earns its place rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, no-output-schema tool, the description supplies everything an agent needs: the prerequisite call, the PUT step between, the verification semantics, the return fields, and where the file_id becomes valid. Nothing material is left to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so file_id and in_workspace are already documented in the schema, and the description does not elaborate on either. It does add context that the sha256/size are checked against what was declared at upload-URL creation time, which clarifies why a mismatch matters, but that is not parameter-level detail. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (finish the file upload) and explicitly ties itself to its predecessor, files_create_upload_url, so the agent can place it in the upload lifecycle without opening a schema. It is clearly separable from the sibling files_upload / files_ingest / files_create_upload_url tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The precondition is unusually explicit: run it only after the bytes have been PUT to the upload_url from files_create_upload_url. It also routes the agent forward by naming messages_send, messages_send_email attachments, and documents.create as consumers of the resulting file_id. It does not, however, state when to prefer files_upload or files_ingest over this two-step flow, so the alternative-selection guidance is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_create_upload_urlAInspect
Get a URL to upload a local file DIRECTLY to storage, without passing its bytes through this conversation.
Use this for any file you already have on disk — it is the only way to send one without spending tokens on its contents, and the only way to send anything above ~50 KB at all.
Three steps:
files_create_upload_url(filename, mime_type, size_bytes, sha256)
curl -X PUT --data-binary @/path/to/file (send every header from
headersverbatim)files_complete_upload(file_id) -> the file is ready
Pass sha256 when you can: if the workspace already holds that exact file you get its file_id back immediately with deduplicated=true and no upload at all.
Returns: file_id, upload_url, method, headers, expires_at.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | Optional display title | |
| sha256 | No | Hex sha256 of the file (`sha256sum <path>`). Optional but recommended: enables dedup and verifies the upload arrived intact. | |
| filename | Yes | Filename with extension (e.g. 'photo.jpg') | |
| mime_type | No | MIME type (e.g. 'image/jpeg'). Guessed from the filename when omitted. | |
| size_bytes | Yes | EXACT byte length of the file (`stat -c %s <path>`). Checked before anything is transferred, so an oversize file is refused here rather than after the upload. Limit: 100 MB. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only give a thin safety profile (readOnlyHint=false, destructiveHint=false). The description goes well beyond: it lays out the full three-step lifecycle, discloses the dedup short-circuit (file_id returned with deduplicated=true, no upload), the 100 MB pre-transfer rejection, URL expiry, and warns to send headers verbatim. This is exactly the behavioral context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core distinction before the mechanics, then uses a clean numbered three-step block, one sentence on dedup, and a one-line returns list. No filler or repetition; each section carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the return fields (file_id, upload_url, method, headers, expires_at). Combined with the workflow and dedup behavior, an agent has everything needed to invoke and follow through correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents sha256 (dedup + integrity), size_bytes (exact length, 100 MB limit, checked pre-transfer) and mime_type guessing. The description mostly reuses that framing and adds little parameter-level meaning beyond the schema; it mainly adds workflow ordering. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a URL to upload a local file DIRECTLY to storage') and immediately contrasts it with the alternative path of sending bytes through the conversation. An agent can distinguish this from files_upload/files_ingest without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scopes usage: 'any file you already have on disk', 'the only way to send one without spending tokens', 'the only way to send anything above ~50 KB at all'. It also gives the condition under which the optional sha256 is worth passing (immediate dedup). Nothing about when-to-use is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_deleteADestructiveIdempotentInspect
Permanently delete files from this workspace by their IDs (generated documents, uploads, screenshots, etc.).
IRREVERSIBLE: removes the DB record, detaches every reference (knowledge collections, thread pins, message attachments), and deletes the stored blob. There is no undo.
Use to clean up leftover / superseded generated files. Only files that belong to this workspace are touched; unknown or other-workspace ids are returned under not_found. Max 20 per call.
| Name | Required | Description | Default |
|---|---|---|---|
| file_ids | Yes | List of file IDs to permanently delete (max 20). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the destructiveHint/idempotentHint annotations by spelling out exactly what is destroyed (DB record, knowledge-collection references, thread pins, message attachments, stored blob), that there is no undo, that only same-workspace files are touched, and that foreign/unknown IDs surface under `not_found`. That is precisely the extra context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The irreversible warning is front-loaded and capitalized, followed by the intended use and the edge-case behavior in three tight sentences. No filler, every sentence carries load.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive 2-param tool with no output schema, the description covers safety, scope, batch limit, and the partial-failure return bucket (`not_found`), which is the one return-value detail an agent actually needs. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (file_ids, in_workspace) are already documented in the schema. The description mainly restates the max-20 limit and the workspace scoping, adding little semantic detail the schema lacks; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (permanently delete), resource (files), scope (this workspace, by IDs) and enumerates what kinds of files it covers. It is trivially distinguishable from siblings like files_upload, files_read, or search_files without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use to clean up leftover / superseded generated files" gives a clear positive trigger and implicitly excludes active/referenced files. It does not name alternatives (e.g., collections_remove_file for detaching rather than deleting), so it stops just short of full when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_get_base64ARead-onlyIdempotentInspect
Download one or more files server-side and return their content as base64-encoded strings. Use this to inspect images, PDFs, or any binary file attached to messages when you cannot access presigned S3 URLs directly. Supports up to 5 files per call, max 15 MB each. For large files batch in groups of 1-2 to avoid oversized responses. For images, pass variants to get processed copies instead of the original: rotate (clockwise) a photo taken sideways or upside down, enhance a faded print, or crop and enlarge a region with small text. Ask for both rotate 90 and 270 in one call when the text runs vertically and you cannot tell which way is up. An MCP client shows at most 5 images per response; the rest are listed, not shown.
| Name | Required | Description | Default |
|---|---|---|---|
| file_ids | Yes | List of file IDs to fetch as base64 (max 5). Get IDs from files.info or message attachment_file_ids. | |
| variants | No | Images only, optional. Each variant is applied to every requested image and returned as its own entry (max 4 variants). Keys: rotate (0, 90, 180, 270; clockwise), enhance (true = grayscale + auto-contrast), crop ([x0, y0, x1, y1], fractions 0..1 of the image after rotation; the region is enlarged). Example: [{"rotate": 90, "enhance": true}, {"rotate": 270, "enhance": true}]. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), and the description adds real behavior beyond them: hard limits (5 files, 15 MB each), batching advice to avoid oversized responses, the MCP client display cap of 5 images with extras listed-but-not-shown, and how variants expand into separate result entries.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action in the first sentence, then layers limits, batching advice, variant guidance, and the display cap in a logical order. It is dense and runs long, but nearly every sentence carries actionable information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates well by explaining that results arrive as base64 entries and that image rendering is capped by the client. Combined with file-size limits, batching guidance, and variant semantics, an agent has everything needed to call this correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so a baseline of 3 is warranted, but the description adds genuine meaning: it explains why variants exist (sideways photos, faded prints, small text) and supplies a practical trick ('ask for both rotate 90 and 270... when you cannot tell which way is up') that the schema does not encode.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb, resource and output form: 'Download one or more files server-side and return their content as base64-encoded strings.' It further scopes itself to binary/message attachments when presigned S3 URLs aren't reachable, which separates it from the read-oriented siblings (files_read, files_info) without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear triggering context ('inspect images, PDFs, or any binary file... when you cannot access presigned S3 URLs directly') plus concrete operational guidance (max 5 per call, batch 1-2 for large files) and a usage heuristic for variant requests. It stops short of naming a sibling alternative explicitly, so it is strong but not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_infoARead-onlyIdempotentInspect
Get metadata and download URLs for files by their IDs.
When to use:
After messages_read_history returns attachment_file_ids
To get a presigned download URL to read a received file
Returns: filename, mime_type, byte_size, download_url (1-hour presigned URL).
| Name | Required | Description | Default |
|---|---|---|---|
| file_ids | Yes | List of file IDs (max 20) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive behavior, so the bar is lower. The description adds genuinely useful context beyond annotations: the returned artifacts and, importantly, that download_url is a 1-hour presigned URL, which signals expiry behavior an agent must plan around.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tight and front-loaded: a purpose sentence, a labeled 'When to use' list, and a compact 'Returns' line. No filler, and every clause carries operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although there is no output schema, the description enumerates the exact return fields (filename, mime_type, byte_size, download_url), covering the gap. Combined with the workflow trigger and full schema coverage, an agent has everything needed to invoke and consume it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both file_ids (max 20) and the in_workspace override are fully documented in the schema, including the 'nothing is stored' semantics. The description adds no parameter-level guidance, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (metadata and download URLs for files) with the key scoping detail 'by their IDs'. This cleanly separates it from siblings like files_get_base64, files_read, and files_ingest, which handle content rather than metadata/URLs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use triggers, notably 'after messages_read_history returns attachment_file_ids' and 'to get a presigned download URL to read a received file'. It does not, however, state when NOT to use it or point to alternatives like files_get_base64 or files_read for actually reading content.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_ingestAInspect
Save and index a file into the knowledge base. Use this when the user asks to save, store, or remember a document. The file will be processed (OCR if needed) and indexed for future search.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Optional list of tags for categorization (e.g., ['presentation', 'dextrade']). | |
| title | No | Human-readable title for the file (e.g., 'Project Presentation', 'Q1 Report'). If not provided, uses original filename. | |
| file_id | Yes | ID of the file to ingest (from attachment_file_ids in context). | |
| thread_id | No | Optional thread ID to associate the file with. If not provided, uses context thread. | |
| description | No | Optional description of the file contents. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the mutation profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=false), so the safety picture is covered. The description adds useful post-condition context, noting the file is processed with OCR if needed and indexed for future search. However, it does not surface the non-idempotent consequence (re-ingesting the same file may create duplicates), which is the most behaviorally significant trait here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, then the trigger, then the behavioral consequence. No filler and no repetition of structured fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter mutation tool with full schema coverage and annotations covering the safety profile, the description covers purpose, trigger, and post-processing well. The only gap is the lack of any idempotency or duplication warning, which matters given idempotentHint=false.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for all six parameters, including defaults and the file_id source ('from attachment_file_ids in context'), so the schema carries the parameter burden. The description adds no parameter-level detail beyond that, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Save and index a file into the knowledge base', which tells the agent exactly what happens. It does not, however, distinguish itself from closely-named siblings such as files_upload, files_create_upload_url, files_complete_upload, collections_add_file, or agents_add_file, so the agent must infer which ingest path applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger condition: 'Use this when the user asks to save, store, or remember a document.' This is solid usage context. It stops short of the 5 level because it names no alternative or exclusion (e.g., when to use files_upload or agents_add_file instead).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_readARead-onlyIdempotentInspect
Read text content of an attached file. Works for: .txt, .md, .json, code files, and PDFs (after files.ingest extracts text). DO NOT call on binary files — for IMAGES use files.get_base64, for AUDIO/VIDEO it cannot be transcribed via this tool, and for non-PDF DOCUMENTS run files.ingest first, THEN files.read. Calling on a binary mime-type returns an error — saves you a turn to read the routing hint before deciding.
| Name | Required | Description | Default |
|---|---|---|---|
| file_id | Yes | ID of the file to read (from attachment_file_ids in context). | |
| encoding | No | Text encoding to use (default: utf-8). | utf-8 |
| max_chars | No | Maximum characters to return (default: 10000). Use smaller values for large files. | |
| summarize | No | If true, generate AI summary instead of returning raw content. Use for 'summary', 'summarize', 'краткое содержание' requests. OMIT to return raw content (the default). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, so the safety profile is covered. The description adds genuine value beyond that by disclosing failure behavior ('calling on a binary mime-type returns an error') and the ingest-then-read ordering for documents. It stops short of describing truncation/pagination behavior around max_chars.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose in the first sentence, then the routing rules. Mostly tight, though the closing clause ('saves you a turn to read the routing hint before deciding') is rationale-flavored padding that could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with no output schema, the description tells the agent what content is returned (text), which formats work, what fails, and how to route elsewhere. Complete enough to call correctly; only minor gaps around truncation limits remain, which the schema's max_chars partially covers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (including encoding, max_chars, summarize, in_workspace) are already documented in the schema. The description adds no additional parameter semantics, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read text content of an attached file') and enumerates the supported formats (.txt, .md, .json, code, PDFs after ingest). It clearly distinguishes itself from files.get_base64 and files.ingest by naming them as the correct routes for other content types.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('read text content') and when-not-to-use guidance with named alternatives: images -> files.get_base64, non-PDF documents -> files.ingest first, audio/video unsupported. This is the model routing guidance an agent needs to pick the right sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_uploadAInspect
Upload a file to DialogBrain and get a file_id for use in messages_send.
When to use:
User wants to send a file/image to a contact
Before calling messages_send with an attachment
For a file that already exists on disk, use files_create_upload_url instead: content here travels through the conversation twice (once read, once written back), so it is only sensible up to about 50 KB. source_url is fine at any size — the fetch happens server-side.
Returns: file_id (integer) to pass to messages_send attachments parameter.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | Optional display title | |
| content | No | Base64-encoded file bytes. Suitable up to ~50 KB — above that the encoded bytes cost more context than the task; use files_create_upload_url for a local file, or source_url for a remote one. Either content OR source_url is required. | |
| filename | No | Filename with extension (e.g. 'photo.png') | upload |
| mime_type | No | MIME type (e.g. 'image/png', 'application/pdf') | application/octet-stream |
| source_url | No | Public URL to fetch file from. Either content OR source_url is required. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the mutation profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false). The description adds genuine context beyond that: the ~50 KB practical ceiling on `content` and the reason (double context cost), plus that `source_url` avoids it because the fetch is server-side. It stops short of covering duplicate-upload behavior or retention, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and return value, then grouped when-to-use bullets, then the size caveat, then a one-line returns note. No sentence is redundant and the decision-relevant constraint (50 KB) is easy to find.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description states the return contract (file_id integer, to be passed as an attachment to messages_send). Combined with the annotations covering the safety profile and 100% schema coverage, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema already documents all six parameters. The description still adds value by explaining the content-vs-source_url size tradeoff and the rationale for the 50 KB guidance, which is meaning the schema only gestures at.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Upload a file to DialogBrain') and names the output artifact (file_id) it produces. It explicitly distinguishes itself from the sibling files_create_upload_url, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use bullets (send a file to a contact, precede messages_send with an attachment) and an explicit when-not: for an on-disk file, use files_create_upload_url because content round-trips through the conversation. The alternative and the condition selecting it are both stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
folders_createAInspect
📁 Create a new inbox folder to organize threads.
When to use:
User wants to create a folder to group related conversations
User wants to organize threads by topic, project, or contact type
After creating a folder, use threads.update with folder_id to move threads into it.
| Name | Required | Description | Default |
|---|---|---|---|
| icon | No | Emoji icon for the folder (max 10 chars, optional) | |
| name | Yes | Folder name (max 100 chars) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and idempotentHint=false, so safety and non-idempotency are covered structurally. The description adds useful post-create workflow guidance (folder_id → threads.update) but says nothing about permissions, duplicate-name behavior, or what the folder object contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line purpose followed by a short when-to-use list and a single next-step sentence; nothing is padded. Slight overhead from the emoji prefix, but the structure is clean and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool with no output schema, the description covers intent, trigger conditions, and the follow-up action, and annotations cover the safety profile. The only meaningful gap is what the call returns (presumably a folder id, only implied by the threads.update instruction).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with name, icon and in_workspace all documented inline including length limits and workspace semantics. The description adds no parameter-level meaning beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new inbox folder to organize threads'), which is unambiguous and distinct from the sibling folders_delete. It does not, however, explicitly name or contrast any sibling, so differentiation is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Includes an explicit 'When to use' block with two concrete conditions (grouping related conversations, organizing by topic/project/contact type) plus a follow-up instruction to move threads via threads.update with folder_id. It gives clear context but no exclusions or alternative-tool routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
folders_deleteADestructiveIdempotentInspect
🗑️ Delete an inbox folder. Threads inside become unfiled (not deleted).
When to use:
User wants to remove a folder they no longer need
User wants to clean up their inbox organization
Threads inside the folder are NOT deleted — they simply move back to the inbox.
| Name | Required | Description | Default |
|---|---|---|---|
| folder_id | Yes | ID of the folder to delete | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, but the description supplies the non-obvious consequence that threads are reverted to the inbox rather than destroyed. It omits permission requirements, failure behavior for a missing folder_id, and whether the operation is reversible, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and its key side effect, and the bullet list is scannable. However, the 'threads are NOT deleted' fact is stated twice (opening line and closing line), which is mild redundancy rather than pure waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with no output schema, the definition covers what matters most: the destructive action, its safety profile via annotations, and the side effect on contained threads. Missing only edge-case handling (nonexistent folder, permissions), which is minor at this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so folder_id and in_workspace are already documented in the schema. The description adds no syntax, format, or constraint detail about either parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (delete an inbox folder) and immediately disambiguates the outcome: threads become unfiled rather than deleted. This clearly separates it from folders_create and from any thread-deletion tool such as threads_delete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'When to use' block gives two concrete triggering intents (remove unwanted folder, clean up inbox organization). It does not name alternatives or explicit when-not conditions, but there is no competing sibling that performs the same deletion, so the context is sufficient to choose this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group_addAInspect
Add a specific group to your discovery list by @username or invite link (t.me/...).
Groups and channels only — this does NOT add an individual person/contact. To add a person by @username (e.g. a customer or lead), use contacts.discover then contacts.sync instead.
When to use:
You already know the group's @username or invite link
Adding a known group without searching
Returns: group metadata including id, title, member_count.
| Name | Required | Description | Default |
|---|---|---|---|
| link | Yes | The group's @username or invite link (e.g. '@phuket' or 't.me/...') | |
| channel | Yes | Channel the group is on (e.g. 'telegram') | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=false, idempotent=false, destructive=false, openWorld=false, so the safety/dedup profile is covered. The description still adds real behavioral context beyond them: the target is a 'discovery list' (not actual group membership), and it discloses the return shape. It stops short of permissions, rate limits, or side effects, keeping it at a 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and scoping clause before the usage bullets and return note; the disambiguation sentences and bulleted when-to-use each carry distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-idempotent mutation with no output schema, it covers purpose, exclusions, alternatives, usage triggers, and an inline description of returned fields (id, title, member_count). Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents 'link', 'channel', and 'in_workspace'. The description restates the link format (@username or t.me/...) but adds no format detail or constraint the schema lacks, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Add a specific group to your discovery list') and immediately scopes it ('Groups and channels only'). It distinguishes itself from the many sibling group_* tools and from contacts_* by name, so an agent can route without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit 'When to use' section (known @username/invite link, adding without searching) and explicit exclusions ('does NOT add an individual person/contact'), plus the concrete alternative chain contacts.discover then contacts.sync for the person case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group_add_memberAInspect
Add a member to an existing group on Telegram or WhatsApp.
What this does:
Adds the specified member to the group
Resolves the member by username, phone number, or JID
Reports if the member is already in the group
Returns: success, chat_id, member, already_member.
| Name | Required | Description | Default |
|---|---|---|---|
| member | Yes | The member to add (format depends on channel: @username on Telegram, phone on WhatsApp) | |
| channel | Yes | Channel where the group exists (e.g., 'telegram', 'whatsapp') | |
| chat_id | Yes | ID of the group/channel to add the member to. Group id exactly as `group.list` reports it: a signed numeric id on Telegram, or a JID ending in @g.us on WhatsApp. A WhatsApp group has no numeric id, so this is a string either way. Read it from `group.list`; never compose one. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish that this is a non-read-only, non-idempotent, non-destructive write. The description adds real context beyond that: member resolution by username/phone/JID and the fact that an already-present member is reported rather than erroring, which is behaviorally important for a non-idempotent tool. It stops short of noting required permissions or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the one-line purpose, then a short bulleted breakdown and a returns line. The 'What this does:' header is mild padding, but each bullet carries distinct information and nothing is repeated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the explicit 'Returns: success, chat_id, member, already_member' line is valuable and fills that gap. For a four-parameter mutation tool with full annotation coverage, the description covers purpose, resolution, and return shape; only permission requirements and error behavior are left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema itself carries the detailed chat_id format rules ('never compose one', signed numeric id vs @g.us JID) plus per-channel member formats. The description's 'resolves the member by username, phone number, or JID' only lightly restates that, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb+resource+scope: 'Add a member to an existing group on Telegram or WhatsApp.' The 'existing group' qualifier separates it from group_create, and 'add a member' separates it from group_join. It never names those siblings explicitly, so an agent must infer the boundary, but the purpose itself is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the tool applies when a group already exists and you know a username/phone/JID. There is no statement of when to prefer this over group_join, group_add, or group_promote_admin, and no mention of prerequisites such as needing admin rights. Adequate but with a clear routing gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group_adminsARead-onlyIdempotentInspect
List administrators of a group on Telegram (or other channels).
Returns admin metadata including:
Admin rank (creator or admin)
Username and real name
Can they post (for restricted groups)
Bot status
Use this to understand group leadership before outreach or evaluation. Returns: group_id, title, admins list, total admin count.
| Name | Required | Description | Default |
|---|---|---|---|
| group_id | Yes | Group to act on. Accepts either the discovered-group id from group.search / group.list, or the platform's own group id (e.g. a Telegram chat id like -1001234567890). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real behavioral value beyond that: it enumerates the returned admin metadata (rank, username/real name, can-post, bot status), which the agent needs since there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then a scan-friendly bullet list of return fields and a short usage line. Minor redundancy between the 'Returns admin metadata including' list and the trailing 'Returns: group_id, title, admins list, total admin count'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description correctly carries the return-value burden by naming the fields returned and the metadata per admin. Combined with the annotations covering the safety profile, an agent has enough to invoke it correctly, though sibling routing remains unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both group_id and in_workspace are already well documented in the schema, including the flexible id formats. The description adds no parameter syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List administrators of a group') and platform context ('on Telegram (or other channels)'). An agent can tell what it returns, but there is no differentiation from siblings like group_promote_admin, group_list, or group_scan, which also touch group membership.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a use case ('to understand group leadership before outreach or evaluation'), which implies when the tool is useful. However there are no exclusions or named alternatives, so an agent must still infer how it differs from group_moderate or group_promote_admin.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group_createBInspect
Create a new group on a channel (Telegram or WhatsApp). Returns the new group's chat ID and invite link.
What this does:
Creates a new group with the specified title
Returns chat_id, invite_link, and channel_ref for further operations
Optionally registers the group in your inbox for monitoring
Returns: success, chat_id, channel_ref, title, thread_id.
| Name | Required | Description | Default |
|---|---|---|---|
| about | No | Optional description or about text for the group | |
| title | Yes | Title/name of the group to create | |
| channel | Yes | Channel to create the group on (e.g., 'telegram', 'whatsapp') | |
| group_type | No | Type of group to create. Options: 'supergroup' (default), 'basic'. Telegram-only; ignored on WhatsApp. | supergroup |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| participants | No | People to put in the group at creation: phone numbers or JIDs on WhatsApp, @usernames or phone numbers on Telegram. Without this the group is created empty and the creator is its only member, which on WhatsApp means nobody can be shown it — there is no invite link to send. Reported back per person in `participants_added` and `participants_failed`. | |
| register_in_inbox | No | Auto-register the created group in your inbox for monitoring. Default: true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the write profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false), so the safety burden is partly lifted. The description adds that the group is registered in the inbox by default and lists the returned identifiers, but it says nothing about authorization requirements, failure modes, or what happens to a partially created group when some participants fail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded well, but the body wastes space: the 'What this does' bullet 'Creates a new group with the specified title' merely restates sentence one, and the closing 'Returns: success, chat_id, channel_ref, title, thread_id' duplicates the earlier chat ID / invite link sentence without reconciling the two lists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating, no-output-schema tool with seven parameters, the description supplies the essential return contract (chat_id, invite_link, channel_ref) that an agent needs to chain follow-up calls, and annotations cover the safety profile. It is only missing error/edge behavior and any usage routing, which keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema itself is unusually rich (participants, group_type, in_workspace all fully documented), so the baseline is 3. The description only restates the title field and hints at register_in_inbox; it adds no semantics the schema does not already carry.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb+resource+scope: create a new group on a channel, limited to Telegram or WhatsApp. That is enough to separate it from join/leave/moderate siblings, but the description never names an alternative tool, so it stops short of explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance. The description never mentions group_join, group_add, group_add_member, or any condition that would select this tool over them, and it states no prerequisites (e.g. that the channel must already be connected).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group_get_invite_linkAInspect
Get an invite link for an existing group (Telegram, WhatsApp).
What this does:
Telegram: mints a FRESH link (e.g. t.me/+...) on every call; previously issued links stay valid
WhatsApp: returns the group's EXISTING link, the same one each time. Do not describe it to anyone as newly issued
Requires the connected account to be the group's owner or an admin with invite rights
When to use:
A member can't be added directly because their privacy settings block invites (group.add_member fails with a privacy error) — send them this link instead
Onboarding people to a private group
Returns: success, chat_id, invite_link.
| Name | Required | Description | Default |
|---|---|---|---|
| channel | Yes | Channel where the group exists (e.g., 'telegram') | |
| chat_id | Yes | ID of the group/channel to export an invite link for. Group id exactly as `group.list` reports it: a signed numeric id on Telegram, or a JID ending in @g.us on WhatsApp. A WhatsApp group has no numeric id, so this is a string either way. Read it from `group.list`; never compose one. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations mark it non-readonly and non-idempotent, and the description explains exactly why: Telegram mints a NEW link each call while prior links remain valid, WhatsApp returns the existing link (with a caution not to describe it as new). It also states the owner/admin-with-invite-rights requirement and the return fields — behavior beyond what the annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the one-line purpose, then scannable 'What this does' / 'When to use' / 'Returns' sections. Every sentence carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers platform-specific behavior, authorization prerequisites, usage triggers, and return shape. With no output schema, the explicit 'Returns: success, chat_id, invite_link' line closes the remaining gap, leaving nothing an agent needs missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents chat_id, channel, and in_workspace thoroughly (including the 'never compose an ID' caveat). The description adds no parameter-level meaning beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Get an invite link for an existing group') with platform scope (Telegram, WhatsApp) stated up front. An agent can immediately tell this apart from group_add_member or group_join without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'When to use' block gives explicit triggering conditions: privacy settings blocking direct adds (add_member fails with a privacy error) and onboarding to a private group. It effectively routes the agent away from group_add_member toward this tool in a concrete failure scenario.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group_joinAInspect
Join a group and start syncing its messages to your inbox. The group must be in your discovery list (use group.search or group.add first).
What this does:
Joins the group on Telegram (or other channel)
Creates a thread in your inbox for syncing messages
Optionally enables AI auto-reply drafts
Returns: success, thread_id, auto_reply_enabled.
| Name | Required | Description | Default |
|---|---|---|---|
| group_id | Yes | Group to act on. Accepts either the discovered-group id from group.search / group.list, or the platform's own group id (e.g. a Telegram chat id like -1001234567890). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| enable_auto_reply | No | Enable AI auto-reply drafts for messages in this group. Drafts can be reviewed and sent manually. Default: false (large public groups would otherwise generate drafts for every incoming message). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-readonly, non-destructive, non-idempotent, non-open-world operation. The description adds real value beyond that by enumerating the side effects (joins on the platform, creates an inbox thread, may enable auto-reply drafts) and the returned fields. It does not say whether a repeat call creates a duplicate thread or whether joining is reversible, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line summary followed by a short bulleted side-effect list and a returns line — every element earns its place and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by listing the return values (success, thread_id, auto_reply_enabled), and it covers prerequisites and side effects for a mutation tool. The only shortfall is not clarifying repeat-call/duplicate-thread behavior for a tool flagged non-idempotent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and each parameter (group_id, in_workspace, enable_auto_reply) is already well documented in the schema, including the auto-reply default and rationale. The description's 'Optionally enables AI auto-reply drafts' merely restates enable_auto_reply, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb ('Join') plus resource ('a group') and the immediate consequence ('start syncing its messages to your inbox'). The 'What this does' list further disambiguates it from siblings like group_add (which only adds to discovery), group_create, and group_leave.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a clear prerequisite ('The group must be in your discovery list') and routes the agent to the right alternatives by name ('use group.search or group.add first'). What is missing is the inverse case — what happens if the group is already joined, or when NOT to call this — so it stops short of a full when/when-not rule set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group_leaveAInspect
Leave a group and stop syncing its messages to your inbox. The group will be moved back to discovered state (can rejoin later).
What this does:
Leaves the group on Telegram (or other channel)
Stops syncing new messages
Disables any AI auto-reply drafts for this group
Keeps the group in discovery list (not deleted)
Returns: success, thread_id.
| Name | Required | Description | Default |
|---|---|---|---|
| group_id | Yes | Group to leave. Accepts either the discovered-group id from group.list (lifecycle='monitored'), or the platform's own group id (e.g. a Telegram chat id like -1001234567890). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the generic flags (readOnlyHint=false, destructiveHint=false, idempotentHint=false). The description goes well beyond them by enumerating concrete side effects: leaves the group on the platform, stops syncing, disables AI auto-reply drafts, and keeps the group in discovery (not deleted). This is exactly the side-effect disclosure a mutation tool needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then a scannable 'What this does' list and a 'Returns' line. Every bullet is informative, though the first bullet mildly restates the opening sentence ('stop syncing its messages' / 'Stops syncing new messages').
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-idempotent mutation tool with no output schema, the description covers the full picture: reversibility, side effects, and the return payload (success, thread_id). Nothing needed to invoke or reason about the call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema itself explains both group_id (discovered id or platform chat id) and in_workspace. The description adds no parameter-level detail, so the schema does the heavy lifting and baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Leave a group') and immediately scopes the effect ('stop syncing its messages to your inbox'). It is clearly distinguishable from siblings like group_join or group_add because it names the resulting state ('moved back to discovered state').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it explains the state transition from monitored to discovered and that the user 'can rejoin later', which tells an agent when this is the right action. It does not, however, explicitly name alternative siblings (e.g. group_join vs group_add) or state when NOT to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group_listARead-onlyIdempotentInspect
List groups in this workspace's DISCOVERY/SCORING list (the discovered_group table).
⚠️ This is NOT your messageable inbox. It is a subset used for discovery + quality scoring. Absence here does NOT mean a chat isn't joined or can't be messaged. Each row includes thread_channel_ref — the ref a synced thread carries — so you can match rows against search_threads results. To get a thread_id you can send to, use search_threads.
Lifecycle values:
discovered: found but not yet evaluated
bookmarked: saved for later
monitored: joined and actively syncing messages
dismissed: hidden
By default, dismissed groups are excluded. Returns: id, title, member_count, lifecycle, scan_status, overall_score, is_member, can_post, my_rank.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results (1-100, default 20) | |
| offset | No | Pagination offset. OMIT to start at row 0 (default). | |
| channel | No | Filter by channel (e.g. 'telegram'). Optional. | |
| lifecycle | No | Filter by state: discovered, bookmarked, monitored (=joined/syncing), dismissed. OMIT to include all states (dismissed excluded by default elsewhere). | |
| min_score | No | Minimum overall score (0.0-1.0). Optional. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only/idempotent/non-destructive profile, so the description's job is to add context beyond that. It does: the default exclusion of dismissed rows, the semantic meaning of each lifecycle value, and a return-field list (valuable since no output schema exists). It stops short of describing pagination behavior or cost, keeping it at a strong 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the scope warning and the not-your-inbox caveat before the lifecycle list and returns, which is the right ordering for a commonly misused tool. The lifecycle bullets are somewhat verbose but each carries distinct semantics; overall well-organized and appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating return fields (id, title, member_count, lifecycle, scan_status, overall_score, is_member, can_post, my_rank) and explaining ref-matching against search_threads. Combined with usage routing and lifecycle semantics, an agent has everything needed to call and interpret it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by explaining what each lifecycle value actually denotes ('discovered: found but not yet evaluated', 'monitored: joined and actively syncing') and reiterating the dismissed-by-default rule. It adds genuine interpretation on top of the enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List groups') and immediately scopes it to the workspace's DISCOVERY/SCORING list backed by the discovered_group table. It explicitly distinguishes itself from the messageable inbox and other thread-listing paths, so an agent can tell it apart from siblings like search_threads and group_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides both when-not ('This is NOT your messageable inbox... Absence here does NOT mean a chat isn't joined') and an explicit alternative with condition ('To get a thread_id you can send to, use search_threads'). It also documents the default scope behavior (dismissed excluded), leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group_moderateAInspect
Run ONE moderation action on a group the connected account administers (Telegram bot groups today).
Actions:
invite_link: mint a fresh invite link; older links stay valid. Optional params: name, member_limit, expire_date, creates_join_request
revoke_invite_link: params: invite_link
admins: list administrators (user_id, username, name, status)
member_count: how many members the group has
member_status: one member's status. params: user_id
promote / demote: params: user_id, optional rights {can_*: bool}
ban / unban: params: user_id, optional revoke_messages
mute / unmute: params: user_id, optional until_date (unix ts)
pin / unpin: params: message_id (omit on unpin for the latest), optional silent
set_title: params: title
set_description: params: description
⚠️ A Telegram BOT cannot add a member or create a group — the Bot API has no method for either. To bring someone in, mint an invite_link and send it to them.
Telegram rules that surprise callers (all verified against a live group):
A bot may only promote/demote admins IT promoted. Check
can_be_editedin member_status first: false means demote will answer 'not enough rights', and that is Telegram, not a failure on our side.promote additionally needs the bot to hold can_promote_members in the group; plain admin rights are not enough.
After unmute the status stays 'restricted' while every can_send_* is true. That is Telegram labelling anyone who carries an explicit permission set — it does NOT mean still muted. Read the permissions, not the status.
unban lifts the ban but does not bring the person back: the status becomes 'left' and they must rejoin.
Returns: success plus the action's own keys (invite_link, admins, member_count, status) and result, Telegram's own payload.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Which action to run (see the tool description for the list) | |
| params | No | Action-specific arguments, e.g. {'user_id': 123} for promote/ban, {'message_id': '456'} for pin, {'title': '...'} for set_title | |
| channel | Yes | Channel the group lives on, e.g. 'telegram_bot' | |
| chat_id | Yes | Group to act on: the platform chat id (e.g. -1004487018858), a channel_ref, or our numeric DB thread id | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Far beyond the annotations: it discloses that old invite links stay valid, that revoke/pin/unpin mutate state, that unban leaves the user 'left' and they must rejoin, and that post-unmute status reads 'restricted' even though all can_send_* are true. It also flags the hard limitation that a bot cannot add members or create groups at all, which is exactly the kind of surprise a caller needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, then a scannable action list, then caveats, then the return shape. Every bullet carries concrete information, though the density of caveats makes it on the long side for a tool an agent must parse on every call.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description closes the loop by naming the returned keys (invite_link, admins, member_count, status, result). It is complete for correct invocation, but the missing steer away from the narrower sibling moderation tools leaves a tool-selection gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the five top-level params, so the baseline would be 3; but the description carries meaning the schema cannot, since `action` has no enum and the schema only says 'see the tool description'. The bullet list supplies every valid action value and its action-specific `params` keys (user_id, message_id, title, rights, revoke_messages, until_date, member_limit, expire_date).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource ('Run ONE moderation action on a group the connected account administers') and enumerates every supported action with its parameters, so the scope is unambiguous. It does not, however, differentiate itself from overlapping siblings such as group_promote_admin, group_admins, and group_get_invite_link, which cover a subset of the same actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives real preconditions per action, e.g. 'Check `can_be_edited` in member_status first' and 'promote additionally needs the bot to hold can_promote_members', which tells the agent when an action will fail. It never states when to choose this dispatcher over the narrower sibling tools, so the alternative-selection guidance is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group_preview_messagesARead-onlyIdempotentInspect
Read recent public messages from a group without joining it. Only works for groups where can_preview_history=true.
Use this to manually evaluate message quality before deciding to join. For an automated quality score, use group.scan instead.
Returns: list of recent messages with sender identity (username, name, admin flag), text (150-char previews; pass full_text=true for untruncated), date, is_reply.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Number of recent messages to fetch (1-100, default 20) | |
| group_id | Yes | Group to act on. Accepts either the discovered-group id from group.search / group.list, or the platform's own group id (e.g. a Telegram chat id like -1001234567890). | |
| full_text | No | Return untruncated message text. OMIT for 150-char previews (keeps large batches consumable). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the read-only, idempotent, non-destructive safety profile, so the bar is lower. The description adds real behavioral context beyond that: the can_preview_history=true gating precondition and the 150-char preview default with full_text override.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose, then usage, then return shape, each in a distinct sentence or section. Every sentence adds information (scope, alternative, return format) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only preview tool with no output schema, the description compensates by enumerating the return fields (sender identity, text, date, is_reply). Annotations cover safety, and the precondition and alternative are noted, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, giving a baseline of 3 (the schema documents all four params). The description goes further by explaining the truncation behavior of full_text ('150-char previews; pass full_text=true for untruncated'), adding semantic clarity on why the flag exists.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read recent public messages from a group') plus a key scope constraint ('without joining it'). This clearly distinguishes it from siblings like group_join and group_scan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('to manually evaluate message quality before deciding to join') and names the alternative for a different need ('For an automated quality score, use group.scan instead'). The precondition (can_preview_history=true) is also stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group_promote_adminAInspect
Promote a member to admin in an existing group on Telegram or WhatsApp.
What this does:
Gives the specified member admin status in the group
On Telegram, this grants visibility of all group messages (even if not a bot)
Defaults to minimal/empty rights; specify custom rights if needed
Returns: success, chat_id, member.
| Name | Required | Description | Default |
|---|---|---|---|
| member | Yes | The member to promote (format depends on channel: @username on Telegram, phone on WhatsApp) | |
| rights | No | Optional admin rights dict (Telegram-specific). If not provided, defaults to minimal/admin status only. Example: {"post_messages": true, "edit_messages": true} | |
| channel | Yes | Channel where the group exists (e.g., 'telegram', 'whatsapp') | |
| chat_id | Yes | ID of the group/channel where the member will be promoted. Group id exactly as `group.list` reports it: a signed numeric id on Telegram, or a JID ending in @g.us on WhatsApp. A WhatsApp group has no numeric id, so this is a string either way. Read it from `group.list`; never compose one. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it is a non-readonly, non-destructive, non-idempotent mutation. Beyond that the description discloses a genuinely non-obvious side effect: on Telegram promotion grants visibility of all group messages even for non-bots, plus the minimal-rights default. It omits permission requirements and whether the change is reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the one-line purpose, then a tight bullet list and a returns line. Slight redundancy between the opening line and the first bullet, but nothing wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-idempotent mutation with no output schema, the description covers the effect, the platform-specific side effect, the rights default, and the return fields. Missing only permission/auth prerequisites needed to call it successfully.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five params including channel-dependent member formats and the rights example. The description adds only the rights-default nuance, which is already partly in the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first line gives a specific verb (promote) and resource (member to admin in an existing group) and names the two supported channels. It does not distinguish itself from near-neighbors like group_admins or group_add_member, so an agent must infer which sibling applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Defaults to minimal/empty rights; specify custom rights if needed' is real usage guidance for the optional rights param. However, there is no statement of when to prefer this over group_admins/group_moderate, no prerequisites (e.g., caller must already be admin), and no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group_scanAInspect
Scan a group to evaluate its quality before joining. Fetches recent messages, analyzes activity, spam, and engagement, then returns a quality score and plain-English verdict.
When to use:
After finding groups with group.search
Before deciding which groups to join
Returns: overall_score (0-1), is_disqualified, disqualify_reasons, individual scores, and a verdict string.
| Name | Required | Description | Default |
|---|---|---|---|
| group_id | Yes | Group to act on. Accepts either the discovered-group id from group.search / group.list, or the platform's own group id (e.g. a Telegram chat id like -1001234567890). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, so the agent knows this is a non-idempotent, non-destructive operation — yet the description describes only fetching and analysis, never explaining what side effect makes it non-read-only. The description does disclose the analysis pipeline (activity, spam, engagement) and return shape, which goes beyond annotations, but the mismatch between 'Fetches' and readOnlyHint=false is left unaddressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence of purpose, a two-item when-to-use list, and a compact returns line — front-loaded and entirely free of filler. Every element earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the returned fields (overall_score, is_disqualified, disqualify_reasons, verdict). The only missing piece is why the tool is flagged non-idempotent/non-read-only, which matters for a scan that evidently touches state.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (group_id and in_workspace) are fully documented in the schema, including the accepted id formats and the no-storage guarantee. The description adds no parameter-level meaning beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (scan), resource (group) and outcome (quality evaluation), and it distinguishes itself from group_search/group_list/group_join by framing the scan as a pre-join decision step. An agent can tell it apart from the discovery and join siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'When to use' section names the sequencer (after group_search) and the decision point (before deciding which groups to join), which is precisely the routing information an agent needs. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group_searchAInspect
Search for public groups or channels by topic on Telegram (or other channels). Returns matching groups with title, member count, and whether messages can be previewed.
Finds public groups/channels by topic — NOT individual people. To find or add a specific person by @username, use contacts.discover / contacts.find instead.
When to use:
Finding groups related to a topic or niche
Building a list of groups for outreach or monitoring
After searching, use group.scan to evaluate quality before joining.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return (1-50, default 20) | |
| channel | Yes | Channel to search on (e.g. 'telegram') | |
| keywords | Yes | Search keywords or phrase (e.g. 'crypto trading signals'). A list of keywords is also accepted and joined with spaces. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds real context beyond the annotation block: it describes the return payload ('title, member count, and whether messages can be previewed') and the follow-up workflow with group.scan. It does not cover auth/rate-limit or pagination behavior. Note the annotation block looks like uncustomized defaults (openWorldHint=false and idempotentHint=false are both questionable for a public Telegram search), so the read-only nature implied by 'Search' and the schema's 'Nothing is stored' sits awkwardly against readOnlyHint=false — but since the annotations are evidently unreliable defaults rather than an intentional counter-claim, this is not treated as a hard contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose and return shape, then a scannable 'When to use' block, and closes with the follow-up step. Slightly redundant between 'Search for public groups or channels by topic' and 'Finds public groups/channels by topic', but nothing is wasted overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with no output schema, it describes the returned fields and the recommended next action (group.scan), and clarifies the people-vs-group boundary. It omits any note on result ordering or pagination beyond the schema's limit, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters (limit, channel, keywords, in_workspace) including defaults and ranges. The description adds no parameter-level detail beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope ('Search for public groups or channels by topic on Telegram'), and explicitly distinguishes itself from sibling tools by declaring it does NOT find individual people and pointing to contacts.discover/contacts.find for that. An agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit 'When to use' list (finding groups by topic/niche, building outreach/monitoring lists), an explicit exclusion (not people → use contacts.discover/find), and a downstream step (use group.scan to evaluate before joining). Both the when and the alternative are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
images_generateAInspect
Generates a PNG image from a text prompt using Gemini 2.5 Flash Image. Returns a file_id consumable by messages.send(attachments=[...]) and other file-aware tools. Supports up to 12 reference image file_ids for subject-consistent edits and composition (use file IDs from the [ATTACHMENTS] block, files.search, or search.files). Latency: ~8-10s per image. Output: 1024×1024 PNG.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text description of the image to generate (3-4000 chars). | |
| aspect_ratio | No | Output aspect ratio. | 1:1 |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| reference_file_ids | No | Optional list of up to 3 file_ids whose images should be used as visual references (for edits, subject consistency, or composition). Files must be image MIME types (image/png, image/jpeg, image/webp, image/gif). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (readOnly=false, idempotent=false, destructive=false), and the description adds real operational context beyond them: latency (~8-10s per image), output dimensions (1024x1024 PNG), and the return value's downstream use. The only weakness is the misleading 'up to 12 reference file_ids' claim, but that is a schema conflict rather than an annotation conflict.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and return contract, and each sentence carries information. Minor redundancy in stating PNG twice (opening line and 'Output: 1024x1024 PNG') keeps it just short of exemplary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by describing the returned file_id and how to consume it, plus latency and dimensions. The only completeness gap is the inconsistent reference-count claim, which could lead an agent to pass an invalid number of references.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 and the description adds only marginal value (purpose of reference IDs, e.g. subject-consistent edits/composition). It also introduces a factual error: the description says 'up to 12 reference image file_ids' while the schema caps the array at 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a specific verb (Generates), resource (PNG image), input (text prompt) and underlying model, so an agent can distinguish it from videos_generate and images_search without opening the schema. Nothing tautological, no ambiguity about output type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains how the output is consumed (file_id usable in messages.send(attachments=[...])) and where to source reference IDs ([ATTACHMENTS] block, files.search, search.files), which is genuinely actionable. However it never states when to pick this over images_search or videos_generate, so sibling differentiation for usage is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
images_searchARead-onlyIdempotentInspect
Searches images in this workspace by visual content using vector embeddings (Voyage multimodal-3). Pass a text description; returns ranked file_ids with cosine scores and presigned download URLs. Up to 50 results.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max number of results. | |
| query | Yes | Text description of what you're looking for (3-4000 chars). | |
| mime_type | No | Optional — restrict to a specific image MIME (e.g. "image/png"). Filter is applied after RAG (same caveat as collection_id). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| collection_id | No | Optional — restrict to images attached to this collection. Filter is applied after RAG, so you may get fewer than `limit` results; pass a larger limit to broaden if needed. | |
| score_threshold | No | Minimum cosine similarity (0.0 returns all, higher = stricter). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, idempotent, and non-destructive behavior. The description adds valuable context beyond annotations: the embedding model (Voyage multimodal-3), the return structure (ranked file_ids, cosine scores, presigned download URLs), and a result cap of 50. It does not mention auth requirements or pagination, but the additions are useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences cover purpose, invocation, and return format. The information is front-loaded, and every sentence contributes meaning without filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately discloses the return format and result limit. It omits some caveats that live in the schema (post-RAG filtering, workspace override) and never routes to alternatives, but it covers what an agent minimally needs to call the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema documents all six parameters in detail. The description only restates 'Pass a text description' and 'Up to 50 results', adding little meaning beyond what the schema already provides, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Searches images') along with scope ('in this workspace by visual content using vector embeddings'). It does not explicitly distinguish itself from sibling tools such as vision_query or search_files, so no sibling differentiation is present.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implied usage: pass a text description to search images. No when-to-use or when-not-to-use guidance is provided, and no alternatives such as vision_query or search_files are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
instagram_list_mediaBRead-onlyIdempotentInspect
List photos and Reels on the connected Instagram Business/Creator account. Returns id, caption, media_type, permalink, thumbnail_url, timestamp.
| Name | Required | Description | Default |
|---|---|---|---|
| after | No | Pagination cursor from a previous call's next_cursor. | |
| limit | No | Page size, 1-50. Default 25. | |
| account_id | No | channel_account id of the Instagram account to use when the workspace has several (see channels.list). OMIT to use the workspace's connected Instagram account. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds one non-obvious behavioral fact — the account must be a Business/Creator type — plus the returned field set. It does not mention pagination flow or rate limits, so it adds only moderate value beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core purpose and followed by the return shape. Every clause carries information; nothing is padded or repeated from the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, listing the returned fields (id, caption, media_type, permalink, thumbnail_url, timestamp) is genuinely useful and covers the return-value gap. Pagination is handled in the schema via 'after'/'next_cursor', so the description is nearly complete; only a note that results are paged is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (after, limit, account_id, in_workspace) are already self-documented in the schema. The description adds no parameter syntax, defaults, or edge cases beyond that, which is the expected baseline 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List photos and Reels') plus scope ('connected Instagram Business/Creator account'), so an agent immediately knows this is a media read on the Instagram account. Siblings instagram_publish_media and instagram_update_media are different enough by verb that no explicit differentiation is needed, but the description never names them, keeping it just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance. It does not say how this differs from instagram_publish_media or instagram_update_media, when to supply account_id versus omitting it (that context lives only in the schema), or that results are paginated across calls. The agent is left to infer all routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
instagram_publish_mediaADestructiveInspect
Publish a photo (IMAGE), a video (REELS) or a carousel of 2-10 images from workspace files to a connected Instagram Business/Creator account. Pass file_id for one, file_ids for a carousel. Returns media_id + permalink. Instagram allows ~25 publishes per day.
| Name | Required | Description | Default |
|---|---|---|---|
| caption | No | Post caption (max 2200 chars). OMIT to publish without caption. | |
| file_id | No | Workspace files.id of the photo or video to publish. Use file_ids instead for a carousel. | |
| alt_text | No | Alternative text for a photo (accessibility; read by screen readers), up to 1000 characters. Photos only: Instagram ignores it for Reels. | |
| file_ids | No | Workspace files.id values for a CAROUSEL, 2 to 10 images, in the order a reader swipes them. Images only: a mixed carousel is refused rather than half-published. | |
| account_id | No | channel_account id of the Instagram account to use when the workspace has several (see channels.list). OMIT to use the workspace's connected Instagram account. | |
| media_type | No | 'auto' (default, detects from mime), 'image', or 'reels'. | auto |
| location_id | No | Facebook Place ID for location tag. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| share_to_feed | No | For Reels: also show on profile grid (default true). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real behavioral context beyond the annotations: the ~25 publishes/day rate ceiling, the guarantee that a mixed carousel is refused rather than half-published, and the returned media_id + permalink. Annotations already flag destructive/non-idempotent/openWorld, so this layer is a useful complement rather than a replacement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two front-loaded sentences: media shapes and required file_id/file_ids routing come first, then the return value and rate limit. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter publish tool with no output schema, the definition covers the call shape, return identifiers, and the rate ceiling. It leaves out failure/retry semantics and what happens to the workspace file after publishing, which keeps it short of fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents caption limits, alt_text photo-only behavior, file_ids ordering, media_type enum, and share_to_feed defaults. The description restates the file_id/file_ids rule but adds no format or constraint details the schema lacks, so the baseline 3 holds.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Publish) plus the resource and enumerates the supported shapes (photo/IMAGE, video/REELS, carousel of 2-10 images) to a connected Instagram Business/Creator account. This clearly separates it from the sibling instagram_list_media and instagram_update_media tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing rules: file_id for a single item, file_ids for a carousel, and account_id omitted to use the default connected account. It does not, however, state when to prefer instagram_update_media or what to do if the account is not yet connected, so the alternatives side is thinner than the selection side.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
instagram_update_mediaADestructiveInspect
Update the caption of a published Instagram photo or Reel. Only caption is editable after publish (Instagram limitation).
| Name | Required | Description | Default |
|---|---|---|---|
| caption | Yes | New caption (max 2200 chars). | |
| media_id | Yes | Instagram media ID (from list_media or thread metadata). | |
| account_id | No | channel_account id of the Instagram account to use when the workspace has several (see channels.list). OMIT to use the workspace's connected Instagram account. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true, idempotent=false, openWorld=true, and readOnly=false, so the safety profile is covered. The description adds a genuinely useful Instagram platform limitation (only caption is editable post-publish), but does not disclose that the existing caption is overwritten or what permissions/account binding is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences: the operation is front-loaded and the platform constraint follows immediately. Zero wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-field mutation tool with full schema coverage, full annotations, and no output schema, the description covers purpose and the key platform constraint. It would be fully complete only if it noted the overwrite behavior or account-context defaults already spelled out in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are documented in the schema itself (including caption length cap and account_id/in_workspace semantics). The description adds no parameter syntax or format detail beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update), a specific resource (Instagram media caption), and scope (published photo or Reel). An agent can distinguish it from siblings instagram_publish_media and instagram_list_media without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by noting that only the caption is editable after publish, which is useful context. However, it does not explicitly say when to use this tool versus instagram_publish_media (e.g., 'use publish_media to create new posts') or state prerequisites, leaving routing mostly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
integrations_add_endpointsAInspect
Add one or more API endpoints to an HTTP-API integration as callable tools, merged additively into the integration for base_url (created if none exists). Each endpoint becomes a tool with params + request/response schemas inferred from the samples you pass. When CREATING a new integration, provide auth: either identity (saved Browser Identity name/id) for cookie-session APIs, OR an auth block for token/header APIs, e.g. {type:'bearer', token:'...'} or {type:'api_key', token:'...', header_name:'X-API-Key'}. Updates keep the existing auth unless a new auth is passed. Returns the new tool count and names. Refresh the tools list afterwards to use them.
| Name | Required | Description | Default |
|---|---|---|---|
| auth | No | Auth block for a NEW token/header integration (or to change auth on update): {type: none|bearer|api_key|basic|custom_headers, ...}. bearer/basic → {token}; api_key → {token, header_name?}; custom_headers → {headers:{name:value,...}}. Token values are encrypted. Omit for browser_identity (use `identity`) or to keep existing auth. | |
| name | No | Display name. OMIT to keep an existing integration's name, or to name a NEW one after its host. | |
| extract | No | Explicit browser_identity auth extract for a NEW integration when auto-detect can't find one. One block, e.g. {kind:'cookie', name:'_session', value_format:'raw', header_name:'Cookie', header_template:'{{name}}={{value}}'} — or a LIST of blocks, merged, which is what a session-cookie API needs: the session in Cookie AND a CSRF header echoing one cookie back: {kind:'composite', blocks:[{kind:'cookies', origin:'https://app.example.com'}, {kind:'cookie', name:'xsrf_token', header_name:'x-xsrf-token'}]}. kind 'cookies' sends every live cookie for that origin (the only way to send an HttpOnly session cookie); its origin must cover the integration's base_url host. Updates reuse existing auth, so omit then. | |
| base_url | Yes | API base URL of the integration, e.g. https://api.boomnow.com | |
| identity | No | Saved Browser Identity name or numeric id — one way to auth a NEW cookie-session integration (updates reuse existing auth). For token/header APIs use `auth` instead. | |
| endpoints | Yes | Endpoints to add. Each: {method, path, query?(object), request_body?(object sample), response_body?(object/array sample)}. path is relative to base_url, e.g. /api/conversations/all. | |
| const_fields | No | Request-body fields that NEVER vary, as dotted paths into the sample body — e.g. ['query', 'operationName'] for a GraphQL endpoint, or ['envelope.version']. Their value is taken from the sample, pinned server-side, and HIDDEN from the model: it can neither see nor mistype them, and only the parts that actually change (variables, ids, dates) stay in the tool's arguments. Use for GraphQL documents, SOAP envelopes, API versions and tenant ids. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| keepalive_operation | No | operationId the session keep-alive should ping to hold a cookie session open, e.g. 'getHomeWidgetAlerts'. Must be a param-less GET on this integration. Set it when the API has no obvious session probe in its names (the sweeper looks for validate/session/me/whoami/health) AND its session idles out in hours: without it such an integration only gets the hourly origin warm and dies overnight. Carried over on later updates; OMIT to leave it unchanged. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the safety profile (readOnly=false, destructive=false, idempotent=false); the description adds substantial behavior beyond that. It discloses additive merging, that updates preserve existing auth unless overridden, that const_fields are pinned server-side and hidden from the model, and that it returns the new tool count and names.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the create/update flow are front-loaded, and every sentence carries information about auth, return values, or follow-up. It is dense and slightly long, but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, nested-object tool with no output schema, the description covers auth requirements, update semantics, side effects, and even the return value (new tool count and names). Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds relational meaning the schema lacks: the identity-vs-auth choice for new cookie-session vs token APIs, and the const_fields use cases (GraphQL documents, SOAP envelopes, tenant ids). It enriches parameter intent rather than merely restating it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (add API endpoints to an HTTP-API integration as callable tools), plus the merge semantics and create-if-absent behavior. It is clearly distinguishable from siblings like integrations_get_endpoints, integrations_remove_endpoints, and integrations_set_auth.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains create-vs-update behavior (provide auth/identity when creating; updates keep existing auth) and tells the agent to refresh the tools list afterward. It does not explicitly name sibling alternatives or give when-not-to-use exclusions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
integrations_capture_sessionAInspect
Restore an expired integration session by capturing the LIVE auth of an open browser page — works for COOKIE sessions (browser_identity) AND TOKEN/HEADER sessions (bearer / api_key / custom_headers, e.g. Devise access-token/client/uid). First call browser_open(url, identity_name) for the integration's site (the page must be logged in AND have called its API), then pass that page_id here with integration_id or base_url. For cookie sessions it captures the full storage_state (incl httpOnly cookies page JS can't read) into the bound identity; for token sessions it captures the auth headers the page sends to the API host and binds them. Re-connects the integration. The captured secret never leaves the server (returns counts only). Generic — any site, any auth scheme, no local Playwright.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | page_id from a prior browser_open of the integration's site (logged in). | |
| base_url | No | Or the integration's base URL, e.g. https://api.boomnow.com. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| integration_id | No | Integration id (from integrations.list). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare it is a non-readOnly, non-destructive, non-idempotent mutation. The description adds valuable context beyond that: it re-connects the integration, captures httpOnly cookies page JS cannot read, and that the secret never leaves the server (returns counts only). It omits required permissions, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first clause, then setup steps and session-type detail follow. It is a dense wall of text rather than bulleted, but nearly every sentence carries actionable content with little filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a multi-step capture tool with no output schema, the description covers the prerequisite sequence, both session types, the returned value (counts only), and the secrecy guarantee. Nothing an agent needs in order to invoke it correctly appears missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning: page_id must come from a prior logged-in browser_open, integration_id and base_url are alternatives that select the target, tying parameters to the workflow. This goes beyond raw schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource — restore an expired integration session by capturing LIVE auth from an open browser page — and distinguishes the two auth schemes (cookie storage_state vs token/header). This clearly separates it from siblings like integrations_set_auth or browser_attach_identity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit precondition flow: call browser_open(url, identity_name) first with a logged-in page that has called its API, then pass page_id with integration_id or base_url. It doesn't explicitly name a sibling as the alternative for the non-capture case, but the ordering and conditions are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
integrations_execute_toolAInspect
Run a connected integration tool by its ext_ slug (from integrations_search_tools) with the given arguments. Discovers + executes WITHOUT changing the tool list — so the catalog can grow without an MCP reload.
| Name | Required | Description | Default |
|---|---|---|---|
| arguments | Yes | Arguments for the tool (per its input schema) | |
| tool_name | Yes | ext<id>_<name> slug to execute | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, openWorldHint=true, idempotentHint=false. The description adds one useful non-annotation nuance: discovery happens without mutating the MCP tool list ('so the catalog can grow without an MCP reload'). It does not disclose that side effects depend entirely on the target slug, nor auth/error behavior, so it remains a modest addition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and the identifier source. The trailing clause about the catalog growing is slightly tangential but still informative, so it earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that can execute an arbitrary connected integration with unpredictable side effects and no output schema, the description covers the mechanics but omits return shape, failure modes, and the fact that behavior is fully determined by the target slug. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters including the nested 'arguments' object and the in_workspace override. The description adds only the slug-format/source hint, which the schema already restates; baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Run') plus the resource ('a connected integration tool') and pins down the exact identifier format (ext<id>_<name> slug). It also names the sibling tool that produces that slug (integrations_search_tools), so an agent can distinguish it from search/get-schema siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly tells the agent where the required tool_name comes from ('from integrations_search_tools') and that arguments are the target tool's own arguments. It gives clear usage context but states no explicit when-not or alternative (e.g. integrations_get_tool_schema to inspect first).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
integrations_get_endpointsARead-onlyIdempotentInspect
List every endpoint (operationId, method, path) registered on an HTTP-API integration. Identify it by integration_id or base_url. Use this to review what tools exist before integrations.remove_endpoints or add_endpoints.
| Name | Required | Description | Default |
|---|---|---|---|
| base_url | No | Or the integration's base URL, e.g. https://api.boomnow.com | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| integration_id | No | Integration id (from integrations.list). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety and caching profile is fully covered. Beyond that the description only adds the identity semantics (integration_id or base_url), which is parameter-level rather than behavioral; there is no mention of pagination, result size, or behavior when the integration is unknown.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and return fields, followed by the routing guidance. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, fully-annotated list tool with complete parameter documentation, the description covers what it returns, how to identify the target, and when to call it. Minor gaps remain: it does not mention pagination or behavior for an unknown/invalid integration, and it doesn't route against the other integrations read tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents base_url, in_workspace, and integration_id, including the 'Or the integration's base URL' either/or phrasing. The description's 'Identify it by integration_id or base_url' restates that without adding format or precedence details, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List every endpoint ... registered on an HTTP-API integration') and even enumerates the returned fields (operationId, method, path). It routes the agent away from the write siblings by naming integrations.remove_endpoints and integrations.add_endpoints, so it is distinguishable without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete usage scenario ('Use this to review what tools exist before integrations.remove_endpoints or add_endpoints'), which implies the read-first workflow. It does not, however, distinguish itself from adjacent read siblings such as integrations_list, integrations_search_tools, or integrations_get_tool_schema, so an agent could still pick the wrong read tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
integrations_get_tool_schemaARead-onlyInspect
Get the full input schema(s) for external tool slugs returned by integrations_search_tools, before executing them.
| Name | Required | Description | Default |
|---|---|---|---|
| tool_names | Yes | ext<id>_<name> slugs from integrations_search_tools | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds the sequencing constraint (call before execution) which is useful, but it never explains the surprising idempotentHint=false on what looks like a pure lookup, nor what the returned schema payload looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero filler, front-loading the action and resource before the workflow hint. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only lookup with full schema coverage and no output schema, the description covers the essential call context. It could go slightly further on the plural 'schema(s)' return shape and the non-idempotency, but nothing blocks correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (tool_names and in_workspace) are already documented in the schema. The description echoes the slug concept but adds no format or syntax detail beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get), resource (full input schema(s)), and scope (external tool slugs returned by integrations_search_tools). An agent can immediately distinguish this from integrations_search_tools (which finds slugs) and integrations_execute_tool (which runs them).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before executing them' establishes the workflow position relative to integrations_execute_tool, and it names integrations_search_tools as the source of valid slugs. It stops short of explicitly naming integrations_execute_tool as the alternative or stating when-not-to-use, so it is clear but not fully exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
integrations_listARead-onlyIdempotentInspect
List the workspace's external integrations (HTTP-API + MCP), each with its id, provider name, base_url, transport, status, and registered tool count. Use before integrations.add_endpoints to find the integration id / base_url.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered structurally. The description adds the shape of what is returned (id, provider, base_url, transport, status, tool count) and the workflow that consumes it, but says nothing about pagination, result limits, or what an empty list means. Given the annotation coverage, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both front-loaded: the first defines scope and output, the second defines the workflow use. Every clause carries information and there is no filler or restated title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with one optional parameter and no output schema, the description is nearly sufficient: it enumerates the returned fields, which substitutes for a missing output schema, and gives the sequencing rationale. It omits ordering/pagination guarantees and does not resolve ambiguity with the similarly scoped agents_list_integrations sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter (in_workspace) with 100% schema description coverage, so the baseline of 3 applies. The description adds no parameter-level detail beyond the schema, but none is needed given how fully the schema documents the workspace-override behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('List the workspace's external integrations') and even enumerates the returned fields (id, provider name, base_url, transport, status, tool count), which is unusually concrete. It does not, however, distinguish this tool from the similarly named sibling agents_list_integrations, so an agent could still hesitate between the two.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a clear use context and names the downstream tool: 'Use before integrations.add_endpoints to find the integration id / base_url.' That is a real routing hint rather than implied usage. It stops short of a full when/when-not statement (e.g. it never clarifies the relationship to agents_list_integrations or other integrations_* list-ish tools), so it does not reach 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
integrations_remove_endpointsAInspect
Remove endpoints (tools) from an HTTP-API integration — e.g. junk paths like static assets, /socket.io, or SPA routes that aren't real API calls. Identify the integration by integration_id or base_url, and the endpoints to drop by operation_ids (e.g. getSocketIo) and/or paths (e.g. /socket.io/). Re-registers the catalog so the removed tools disappear. Returns removed + remaining counts.
| Name | Required | Description | Default |
|---|---|---|---|
| paths | No | Exact spec paths to remove (all methods), e.g. ["/socket.io/"]. | |
| base_url | No | Or the integration's base URL. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| operation_ids | No | operationIds to remove, e.g. ["getSocketIo","getPieScreensMenu"]. | |
| integration_id | No | Integration id (from integrations.list). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations carry the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false), so the description only needs to add context, and it does disclose the catalog re-registration side effect and the returned removed/remaining counts. It does not state whether removal is reversible or what auth/permissions are needed, and 'the removed tools disappear' sits in mild tension with destructiveHint=false, leaving some ambiguity about permanence.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose and the junk-path rationale before the parameter mechanics and the return summary. Dense but every sentence carries information; slightly more repetition of schema examples than strictly needed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with no output schema, the description covers identification, selectors, side effect, and return shape (removed + remaining counts), which is close to complete. The remaining gap is permission/authorization requirements and reversibility.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds the alternative-selection semantics the schema leaves implicit: identify the integration by integration_id OR base_url, and the endpoints by operation_ids and/or paths. That either/or combination guidance is genuinely additive beyond the per-field schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Remove endpoints (tools) from an HTTP-API integration') with concrete examples of what qualifies as junk (static assets, /socket.io, SPA routes). It is easily distinguishable from integrations_add_endpoints and integrations_get_endpoints by name and scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use scenario — cleaning up non-API paths before/after catalog registration — and states the effect ('Re-registers the catalog so the removed tools disappear'). It does not explicitly name the complementary sibling (integrations_add_endpoints) or note when not to use it, so it falls short of a full routing statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
integrations_search_toolsARead-onlyInspect
Search your connected integrations for tools by intent. Returns matching tool slugs (ext_) + summaries + required params. Then call integrations_get_tool_schema for full input schemas, and integrations_execute_tool to run one. Use this instead of guessing tool names for external actions.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results (default 10) | |
| query | Yes | What you want to do, e.g. 'create a GitHub issue' | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is covered. The description adds return-shape context (tool slugs of the form ext<id>_<name>, summaries, required params) that the annotations do not convey. It stops short of mentioning result caps or pagination behavior, so it is useful but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose and return shape first, then the follow-up tool chain, then the anti-pattern. Front-loaded and waste-free.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by spelling out what comes back (slugs, summaries, required params) and how results feed the next two calls in the chain. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so query, limit and in_workspace are already documented. The description only restates the intent of the query parameter and adds nothing about limit or the workspace-scoped in_workspace override. Baseline 3 applies when the schema carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search your connected integrations for tools by intent') and immediately distinguishes itself from the sibling tools it hands off to (integrations_get_tool_schema, integrations_execute_tool). An agent knows this is discovery-only, not schema retrieval or execution.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the workflow alternatives and when to use each: search here, then call integrations_get_tool_schema for full input schemas, then integrations_execute_tool to run. It also states the negative case ('Use this instead of guessing tool names for external actions'), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
integrations_set_authAInspect
Set or update the authentication on an HTTP-API integration, generically. Identify it by integration_id or base_url. Pass auth = {type: bearer|api_key|basic|query|custom_headers|browser_identity|none, ...}: bearer/basic → {token}; api_key → {token, header_name?} (key rides in a HEADER); query → {token, param_name} (key rides in the QUERY STRING, e.g. param_name='api_key' → ?api_key=…); custom_headers → {headers:{name:value,...}}; browser_identity → {identity_id|identity_name, extract}. Token values are encrypted; nothing is stored in cleartext. To only refresh a browser_identity session, OMIT auth and pass cookies (and/or local_storage) — they are merged into the bound identity without touching its config. (oauth2 / service_account use the REST settings flow.)
| Name | Required | Description | Default |
|---|---|---|---|
| auth | No | Auth to set: {type, token?, header_name?, headers?(object), identity_id?, identity_name?, extract?(object)}. Omit to only refresh cookies on an existing browser_identity integration. | |
| cookies | No | Cookies to merge into the bound browser identity. Each {name, value, domain?, path?, expires?}. Existing names are overwritten; new ones added. | |
| base_url | No | Or the integration's base URL, e.g. https://api.boomnow.com. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| local_storage | No | localStorage entries to merge. Each {origin, name, value}. | |
| integration_id | No | Integration id (from integrations.list). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover safety (readOnly=false, destructive=false, idempotent=false), and the description adds real behavior: token values are encrypted and never stored in cleartext, and cookies/local_storage are merged into the bound identity without touching its config. It does not address whether overwriting an existing auth is reversible, but this is a meaningful addition beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded with the verb and resource, then the auth-type breakdown, then the refresh-only special case. Every clause carries information; slightly overpacked for a single paragraph but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param, all-optional mutation tool with no output schema, the description covers identification, payload variants, encryption, and merge semantics. A brief note on whether overwriting existing auth is reversible/idempotent would complete it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds genuine semantics: it explains where each key rides (`api_key` → HEADER, `query` → QUERY STRING with a concrete `?api_key=…` example) and the per-type payload shape, which the schema enumerates but does not contextualize.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (set/update) and resource (auth on an HTTP-API integration), and names the identification keys (`integration_id` / `base_url`). An agent can distinguish it from siblings like integrations_add_endpoints or browser_attach_identity without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: identify by id or base_url, and if only refreshing a browser_identity session, OMIT `auth` and pass `cookies`/`local_storage`. It also redirects oauth2/service_account to the REST settings flow, covering the when-not and alternative case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
integrations_set_realtimeAInspect
Add, replace or remove the realtime (websocket) connector of an HTTP-API integration, leaving the rest of its configuration and its credentials untouched. realtime is the whole block: {protocol: 'actioncable' | 'raw_json', ws_url, auth: {type, placement: 'header' | 'query', query_param, token_mint: {method, path, token_path}}, channel_template, actions: [{name, description, is_write, input_schema, frame}]}. Its actions become tools of the integration. To remove the connector pass remove=true and no block. Needs the admin or owner role.
| Name | Required | Description | Default |
|---|---|---|---|
| remove | No | true = delete the connector and its tools. OMIT to set a block. | |
| realtime | No | The whole realtime block. OMIT only together with remove=true. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| integration_id | Yes | The integration to change (from integrations.list). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations supply the safety profile (readOnly=false, openWorld=false, idempotent=false), and the description adds real value on top: the admin/owner role requirement, the guarantee that other configuration and credentials are untouched, and that removal deletes the connector together with its tools. Note the mild tension with destructiveHint=false, since remove=true performs a deletion; the description itself discloses this rather than hiding it, so this is disclosure, not contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the effect and the preservation guarantee, followed by the block spec and the removal rule. The inline block enumeration is dense but earns its place since the schema types the property only as a generic object; a sentence could be trimmed but nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating, nested-config tool with no output schema, the description covers the mutation semantics, permission requirement, and preservation guarantee. It does not describe the response payload, but with no output schema that is a minor gap and the key behavioral facts are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes beyond the schema by spelling out the internal shape of the `realtime` object (protocol values, ws_url, auth placement/query_param/token_mint, channel_template, actions with is_write and input_schema). That is useful because the schema itself only says 'the whole realtime block'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('set realtime connector of an HTTP-API integration') and immediately bounds the scope: it replaces only the realtime block and leaves configuration and credentials untouched. An agent can distinguish this from integrations_set_auth and integrations_add_endpoints without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger for the removal mode ('pass remove=true and no block') and states the non-obvious side effect that realtime actions become tools of the integration. It does not name sibling tools to prefer for adjacent tasks (e.g., auth or endpoint changes), so routing guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
integrations_sync_knowledgeAInspect
Refresh a knowledge collection from an integration's endpoints, for content too large to return in a tool response (a 300KB+ catalogue an agent could never fit in one call). Pass sources + collection to configure and run in one step — they are stored on the integration, so later runs need only integration_id. Every run is a FULL refresh; a source whose rendered text is unchanged is skipped (reported unchanged) rather than re-uploaded. Returns a per-source report: created / updated / unchanged / no_records / error.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | Re-upload and re-index every source even when its rendered text is unchanged. Use after a fix BELOW the text — chunking, extraction, embeddings — which the content digest cannot see (without this the only repair was deleting the synced files by hand). | |
| sources | No | Sources to store on the integration and then sync. OMIT to re-run whatever is already stored. REPLACES the stored list wholesale — pass every source you want, not just a new one, or the omitted ones stop syncing. Each item: {operation_id (from integrations.get_endpoints), arguments? (e.g. {"locale": "ru"}), title?, spec, endpoint?}. `endpoint` {method, path} lets the sync call an operation NOT registered on the integration (kept off the agent tool surface — e.g. a 300KB catalogue). `spec` is the projection: {records (REQUIRED dotted path to the row list), envelopes?, recurse?, title?, body?, meta?, drop?, max_chars?}. Same operation with different `arguments` is a separate document, which is how one endpoint serves four locales. | |
| collection | No | Knowledge collection to publish into, created if absent. Required when `sources` is given; ignored otherwise. | |
| description | No | Description for a newly created collection. Optional. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| integration_id | Yes | Integration id (from integrations.list) to sync. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the safety profile (readOnly=false, destructive=false, idempotent=false, openWorld=false); the description adds real behavior the agent cannot infer: every run is a FULL refresh, unchanged sources are skipped and reported as `unchanged`, `sources` REPLACES the stored list wholesale, and `force` exists because the content digest cannot see chunking/extraction/embedding changes. It also enumerates the per-source return categories.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and the reason it exists, then layered with the one-step-vs-later-run distinction and the return report. Dense but not padded; some sentences carry stacked parentheticals that slightly slow parsing, which keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description supplies the return shape (created/updated/unchanged/no_records/error), covers the full-refresh and skip semantics, the wholesale-replacement hazard, and the `force` escape hatch. For a 6-parameter mutation tool this is complete enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema descriptions for `force`, `sources`, `collection`, and `integration_id` are already highly detailed, so baseline is 3. The description adds the persistence model (sources/collection are stored on the integration, so later runs need only `integration_id`) and the fact that the same operation with different `arguments` is a separate document, but most parameter-level detail lives in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (refresh a knowledge collection from an integration's endpoints) and immediately gives the distinguishing use case: content too large to return in a tool response. An agent can tell this apart from integrations_execute_tool or collections_refresh_website without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditional guidance: pass `sources`+`collection` to configure and run in one step, later runs need only `integration_id`, and `force` is for repairs below the text layer. It stops short of naming sibling alternatives explicitly (e.g. when to use integrations_execute_tool or collections_add_website instead), so it is clear context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
job_completeAInspect
Mark the job as completed. This sanitizes PII from the context and records a completion summary. Use when all tasks in the job are done.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | Omit it: the run's own job is used. | |
| summary | No | Brief summary of what was accomplished | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-destructive, non-idempotent, non-open-world call. The description adds meaningful side-effect detail beyond that: it sanitizes PII from the context and records a completion summary. It does not note that calling it on an already-completed job (non-idempotent) may fail, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, then the side effects, then the usage trigger. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description covers the essential state transition and its effects, and all parameters are documented in the schema. It is nearly complete; only edge cases (re-completion, failure modes, what the completion summary is used for) are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so job_id, summary, and in_workspace are already documented. The description only loosely echoes the summary parameter and says nothing about in_workspace or the default job_id behavior, so it adds little beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Mark the job as completed') plus the two side effects (PII sanitization, completion summary). It is clearly distinct from job_read_context and job_update_context by name/action, though it does not explicitly contrast itself with the closest alternative, job_escalate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger: 'Use when all tasks in the job are done.' That is a clear usage condition, but there are no exclusions or alternatives (e.g., when to use job_escalate instead, or what to do if the job is abandoned rather than finished).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
job_escalateAInspect
Escalate the job to a human. Use when you cannot resolve an issue, someone is not responding, or a situation requires human judgment.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | Omit it: the run's own job is used. Never pass a thread or message id here. | |
| reason | Yes | Why escalation is needed | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (not read-only, not idempotent, not destructive, closed-world), lowering the bar. The description says when to escalate but nothing about what escalation actually does mechanically (notify whom, whether the job pauses, reversibility, latency), which is meaningful missing context for an action tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, purpose front-loaded, no filler. Every clause either states the action or the condition for using it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity action tool with one required param, full schema coverage, and no output schema, the description adequately conveys purpose and trigger conditions. It could be stronger by clarifying the effect of escalation or contrasting with agent_handoff, but nothing essential to calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including useful constraints like omitting job_id to default to the run's own job and not passing thread/message ids. The description adds no parameter-level meaning beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Escalate the job to a human'), so the agent knows exactly what action is performed. However, it does not differentiate itself from plausible siblings such as agent_handoff or agents_ask, which also route work elsewhere, so the boundary is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear triggering conditions ('when you cannot resolve an issue, someone is not responding, or a situation requires human judgment'). It does not name an alternative tool or state when NOT to escalate, so it stops short of explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
job_read_contextAInspect
Read the current job context. Returns the full state of your active job including assignments, escalations, and any data you previously stored.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | Omit it: the run's own job is used. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations exist (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=false), so the bar is lower. The description usefully discloses the returned payload (assignments, escalations, previously stored data), which is valuable since there is no output schema, but it says nothing about side effects or why a nominally read tool is marked non-read-only.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with purpose front-loaded before the return-value detail. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description needs to convey return value, and it does so adequately (full active-job state with assignments, escalations, stored data). For a two-optional-parameter read tool whose annotations cover the safety profile, this is close to complete; only side-effect/edge behavior is unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both optional parameters (job_id, in_workspace) are documented in the schema, so baseline 3 applies. The description adds no parameter-level meaning beyond the schema's own notes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (job context) plus what is returned: 'the full state of your active job including assignments, escalations, and any data you previously stored.' Clear and self-contained, but it does not name or distinguish itself from the context-mutating siblings (job_update_context, job_complete, job_escalate).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage ('Read the current job context') and enumerates the returned contents, so an agent can infer when it would want this state. However, it gives no explicit when-to-use framing, no exclusions, and never points to a sibling alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
job_update_contextAInspect
Update the job context by merging new data. Existing keys are preserved unless explicitly overwritten. Use this to record progress, update assignment statuses, or store intermediate results.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | Omit it: the run's own job is used. | |
| updates | Yes | Key-value pairs to merge into job context | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations establish this is a non-read-only, non-idempotent, non-destructive write. The description adds a genuinely useful trait beyond them: merge semantics where 'existing keys are preserved unless explicitly overwritten.' It omits concurrency/last-write-wins behavior, which for a non-idempotent writer would be worth stating.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both load-bearing: the first defines the operation and its merge behavior, the second enumerates use cases. Front-loaded with the action, no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with full schema coverage and no output schema, the description covers what the operation does, how merging works, and why to call it. Missing only edge-case behavior like concurrent update handling, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description's 'merging new data' echoes the schema's own 'to merge into job context' for `updates` and adds nothing about `job_id` or `in_workspace`, which the schema already documents well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Update the job context by merging new data') and clarifies the merge semantics versus replacement. It implicitly distinguishes itself from the sibling job_read_context, though it never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete when-to-use guidance: 'record progress, update assignment statuses, or store intermediate results.' That is clear context for invocation. It stops short of naming alternatives or exclusions (e.g., when to prefer job_read_context or job_escalate).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kg_find_entityARead-onlyIdempotentInspect
Find an entity by name in the Knowledge Graph.
USE WHEN user mentions a person, project, company by name and you need:
To resolve a name to entity_id for subsequent queries
'Кто работает над X?' → find X first
'Расскажи про Y' → find Y first
RETURNS entity_id for use in kg.get_relationships or kg.explore. ALWAYS use this as the FIRST step in KG query chains.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Entity name to search for. Can be in any language (Russian, English, etc.) - transliteration is automatic. | |
| limit | No | Maximum results to return (1-10). Default: 5 | |
| entity_type | No | Filter by entity type: - 'person': People, contacts - 'project': Projects, tasks - 'organization': Companies, teams - 'event': Meetings, deadlines - 'topic': Discussion topics - 'location': Places, addresses - 'document': Files, documents OMIT to include all entity types. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and closed-world behavior, so the safety profile is covered. The description adds meaningful behavioral context: it returns entity_id specifically for chaining into kg.get_relationships/kg.explore. It stops short of describing match behavior (ambiguity handling, ranking) or what happens when no entity is found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then uses labeled USE WHEN / RETURNS / ALWAYS sections. Every line earns its place and the downstream-chaining rule is prominent, aiding fast scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description explicitly states the return value (entity_id) and its intended downstream use. With all four parameters fully documented in the schema and the chaining role explained, an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with full inline descriptions for name, limit, entity_type, and in_workspace, so the schema carries the parameter burden. The description adds no parameter syntax or semantics beyond what the schema already provides, making the baseline of 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Find an entity by name in the Knowledge Graph') and clarifies its role as the entry point to KG query chains. It distinguishes itself from kg.get_relationships and kg.explore by positioning itself as the name-resolution step that precedes them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'USE WHEN' block names the trigger (a named person/project/company) and gives concrete examples including the 'find X first' pattern. It also names the downstream tools and states 'ALWAYS use this as the FIRST step in KG query chains,' leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kg_get_relationshipsARead-onlyIdempotentInspect
Get relationships for a specific entity from Knowledge Graph.
USE WHEN:
'Кто работает над X?' - filter by works_on
'С кем общался Y?' - filter by discussed_with
'Кто из компании Z?' - filter by member_of
'Что связано с W?' - no filter, get all
REQUIRES: entity_id from previous kg.find_entity step. Use: {{step_N.entity_id}} where N is the find_entity step number.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum relationships to return (1-50). Default: 20 | |
| direction | No | Relationship direction: - 'outgoing': Entity → Others - 'incoming': Others → Entity - 'both': All relationships (default) | both |
| entity_id | Yes | Entity ID from kg.find_entity step. Use {{step_N.entity_id}} reference. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| relation_types | No | Filter by relationship types (optional): People: works_on, works_for, member_of, manages, knows, client_of, provides_service Communication: discussed_with, participated_in, mentioned_in Org/Project: developed_by, funded_by, partnered_with, integrates_with, depends_on, part_of Document: issued_by, issued_to, signed_by, authored_by Other: uses, located_in, about, follows, owns, related_to |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds the dependency chain from find_entity, but nothing about return shape, pagination against the limit, or error behavior. With annotations carrying the safety burden, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line purpose followed by labelled USE WHEN and REQUIRES blocks that are skimmable. Every example earns its place by tying a question to a filter value, though the mixed-language phrasing is slightly quirky for a machine reader.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Full schema coverage plus annotations cover parameters and safety for this read-only tool. The one gap is the absence of any statement about the response shape (entities vs. edges) with no output schema to compensate, but the usage context is otherwise complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description rises above it by translating natural-language intents into concrete relation_types values, which the schema only lists abstractly. It does not, however, add anything beyond the schema for direction, limit, or in_workspace.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get relationships for a specific entity from Knowledge Graph') with clear scope. It distinguishes itself from the sibling kg_find_entity by naming it as the required upstream step. An agent can tell exactly what this returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
A 'USE WHEN' block maps concrete user intents to relation_types filters (works_on, discussed_with, member_of, or no filter for all). It also names the alternative/predecessor explicitly via 'REQUIRES: entity_id from previous kg.find_entity step', leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_gaps_answerAInspect
Answer a question from knowledge.gaps_list. Your title and answer are published as a knowledge article the asking agents will find next time; nothing is sent to the customer. Answering again replaces the previous answer.
| Name | Required | Description | Default |
|---|---|---|---|
| answer | Yes | The answer text (max 4000 characters), published with the title | |
| gap_id | Yes | Gap id | |
| question | Yes | The article title, in your own words (max 200 characters). It is published to the knowledge base the agents read, so do not copy the customer's wording from knowledge.gaps_list. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations only declaring it is a non-read-only, non-destructive, non-idempotent write, the description adds real behavioral context: the title and answer become a published knowledge article visible to other asking agents, nothing is sent to the customer, and a repeat call replaces the previous answer. That replace semantics directly informs how the non-idempotent hint should be interpreted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action, then side effects, then the re-answer rule. No filler and nothing repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple four-parameter write with no output schema, the description supplies the workflow source (gaps list), the side effect, the audience, and the repeat-call behavior. Nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (including the title-as-question mapping and the in_workspace override) is already documented in the schema. The description adds no syntax or format detail beyond what the schema provides, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Answer a question from knowledge.gaps_list') and names the sibling tool that produces the input, so the agent can distinguish it from knowledge_gaps_list and knowledge_query without opening a schema. The downstream effect (publishes a knowledge article) is also stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly establishes the precondition: you answer a question that originated in knowledge.gaps_list. It also notes that answering again replaces the prior answer, which is a usage-relevant condition. It stops short of naming explicit alternatives (e.g., knowledge_query or notes_save) or stating when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_gaps_listARead-onlyIdempotentInspect
List the questions agents could not answer from the knowledge base, merged by meaning and ordered by how often they were asked. Answer one with knowledge.gaps_answer.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | 1-200, default 20 | |
| status | No | open (default), answered or dismissed | open |
| agent_id | No | Only gaps this agent ran into | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered. The description adds genuinely non-obvious behavior: results are semantic merges of repeated questions and sorted by ask frequency, which an agent could not infer from the schema. It does not mention pagination behavior or what happens to a gap after it is answered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler. The core behavior (merged by meaning, ordered by frequency) is front-loaded and the follow-up action is placed second.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter, no-output-schema list tool this is nearly complete: it says what a gap is, how results are merged and ordered, and what to do next. The one real gap is that with no output schema, the description still does not hint at the shape of a returned item (id, question, ask count, status).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with limit, status, agent_id and in_workspace all documented inline, including the default status of 'open' and the workspace scoping caveat. The description adds no parameter-level meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (list knowledge gaps) plus two distinguishing behaviors: gaps are merged by meaning and ordered by frequency. An agent can tell this apart from knowledge_query or kg_find_entity without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to the follow-up action ('Answer one with knowledge.gaps_answer'), which is the natural next step after listing. It stops short of stating when not to use this tool (e.g. use knowledge_query when you have a real question rather than auditing unanswered ones), and the sibling is referenced as 'knowledge.gaps_answer' rather than the actual tool id knowledge_gaps_answer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_queryARead-onlyIdempotentInspect
Answer questions using knowledge base (uploaded documents, handbooks, files).
Use for QUESTIONS that need an answer synthesized from documents or messages. Returns an evidence pack with source citations, KG entities, and extracted numbers.
Modes:
'auto' (default): Smart routing — works for most questions
'rag': Semantic search across documents & messages
'entity': Entity-centric queries (e.g., 'Tell me about [entity]')
'relationship': Two-entity queries (e.g., 'How is [entity A] related to [entity B]?')
Examples:
'What did we discuss about the budget?' → knowledge.query
'Tell me about [entity]' → knowledge.query mode=entity
'How is [A] related to [B]?' → knowledge.query mode=relationship
NOT for finding/listing files, threads, or links — use search.files / search.threads / search.links for that.
| Name | Required | Description | Default |
|---|---|---|---|
| date_to | No | Filter messages until this date (ISO format: YYYY-MM-DD). | |
| file_ids | No | Specific file IDs to search within (for pinned files). Only ids returned by files.search or listed in the [ATTACHMENTS] block of THIS run; an id you have not seen in this workspace matches nothing and the search comes back empty. Leave it out to search the agent's collections. Numeric ids; numeric strings are accepted. | |
| question | Yes | The question to answer from user's knowledge base. Required even for entity queries. | |
| date_from | No | Filter messages from this date (ISO format: YYYY-MM-DD). Use for time-based queries like 'this week', 'last month'. | |
| thread_id | No | Limit search to a specific thread/chat | |
| max_sources | No | Maximum number of sources to consider (1-10) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| needs_aggregation | No | True if query asks for totals/sums/counts. | |
| include_relationships | No | Include KG relationships in answer (default: true for entity mode) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds behavioral value by disclosing the return shape ('evidence pack with source citations, KG entities, and extracted numbers') and the four routing modes, though it omits limits such as result caps or latency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then modes, then examples, then the exclusion — a logical order. The three example lines are somewhat redundant with the mode list, but each earns modest value and nothing is buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so ('evidence pack with source citations, KG entities, and extracted numbers'). Combined with the mode explanations, usage rules, and explicit sibling routing, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all nine parameters and the baseline is 3. The description adds meaning beyond the schema by defining the routing modes (auto/rag/entity/relationship) that aren't exposed as schema properties, which is useful context for invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Answer questions using knowledge base') and immediately scopes it (uploaded documents, handbooks, files). It explicitly distinguishes itself from siblings by naming search.files/search.threads/search.links as the tools for finding/listing, so an agent can select this without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('QUESTIONS that need an answer synthesized from documents or messages') plus an explicit when-not ('NOT for finding/listing files, threads, or links') with the exact alternatives. It even maps example queries to the right mode.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
linkedin_raw_requestARead-onlyIdempotentInspect
Send an arbitrary LinkedIn API request via Unipile's magic route. Only GET and POST methods are allowed. WARNING: This bypasses structured rate limiting and can perform destructive actions. Use this only when no other LinkedIn tool covers the needed functionality.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | Request body (for POST requests) | |
| method | No | HTTP method (only GET and POST allowed) | GET |
| request_url | Yes | Target LinkedIn API endpoint URL | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| query_params | No | URL query parameters |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description warns that this 'bypasses structured rate limiting and can perform destructive actions', which directly conflicts with annotations declaring readOnlyHint=true, destructiveHint=false, and idempotentHint=true. An agent trusting the annotations would assume a safe idempotent read while the description describes a potentially destructive, unthrottled write path. This is a serious annotation contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with zero waste: purpose first, then method constraint, then the destructive/rate-limit warning, then the fallback rule. Front-loaded and appropriately sized for a high-risk escape hatch.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-level open-world request tool with no output schema, the description covers the essential constraints and the risk profile. What it omits is what a raw call actually returns or how errors surface, and the behavioral picture it paints is at odds with the annotations rather than reconciled with them.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (request_url, method, body, query_params, in_workspace) are already documented in the schema. The description only restates the GET/POST restriction already carried by the method enum, adding no extra parameter meaning. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Send an arbitrary LinkedIn API request') plus the mechanism ('via Unipile's magic route'), and explicitly of the fallback nature relative to the other LinkedIn tools. An agent can identify this as the raw escape hatch without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule ('Use this only when no other LinkedIn tool covers the needed functionality') and a hard constraint (only GET and POST). This is a clear last-resort routing instruction against the many structured linkedin_* siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
linkedin_update_profileAInspect
Update the authenticated user's own LinkedIn profile. Supports adding/editing experience entries (role, company, skills, dates). Also supports updating location. Headline, summary, education are NOT supported by the API.
| Name | Required | Description | Default |
|---|---|---|---|
| location | No | Location to set on profile (requires LinkedIn location ID) | |
| experience | No | Add or edit a professional experience entry | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false and openWorldHint=false, so the safety profile is covered. The description adds real context beyond that: the operation is scoped to the authenticated user's own profile (an auth/ownership constraint) and it discloses hard API-side field limitations. It does not discuss reversibility, required scopes, or partial-update behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the operation and scope, then supported fields, then the exclusion list. Every sentence carries distinct information and the negative constraint in the final sentence earns its place by preventing wasted calls.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema but full annotation coverage, the description covers ownership scope, supported and unsupported fields, and the add-vs-edit intent. Minor gaps remain: no mention of what a successful response contains or whether a call with no arguments is a no-op, but these are secondary given the schema and annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the nested location/experience objects, the location ID requirement and the add-vs-edit id semantics are already fully documented in the schema. The description restates the same field set (role, company, skills, dates; location) without adding format, ordering, or interaction detail. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Update the authenticated user's own LinkedIn profile') and enumerates exactly which sub-resources are supported (experience entries, location) versus not (headline, summary, education). The scope qualifier 'authenticated user's own' pins it down well, though no sibling tool (e.g. linkedin_raw_request) is named as the fallback for the unsupported fields.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use conditions (adding/editing experience, updating location) and an explicit when-not ('Headline, summary, education are NOT supported by the API'). It stops short of naming the alternative tool an agent should reach for those unsupported fields, so routing is implied rather than complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
messages_deleteADestructiveIdempotentInspect
Delete a message from a thread. Supports Telegram, WhatsApp, and other connected channels. Note: Some channels have time limits on message deletion.
| Name | Required | Description | Default |
|---|---|---|---|
| thread_id | Yes | Thread/channel ID containing the message | |
| message_id | Yes | ID of the message to delete | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds genuinely new context: cross-channel scope (Telegram, WhatsApp, other channels) and channel-specific time limits, which an agent cannot get from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: action, channel scope, then caveat, with the core purpose front-loaded. No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, no-output-schema tool whose annotations already cover reversibility and safety, the description supplies channel scope and the time-limit constraint. Minor gaps remain (whether deletion is permanent per channel, permission needs), but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so thread_id, message_id, and in_workspace are already documented. The description adds no parameter-level detail (no format hints for IDs, no explanation of the workspace override), so the schema carries the full load and baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('Delete a message from a thread'), which clearly separates it from messages_edit, messages_forward, and threads_delete. It stops short of explicitly naming any sibling or scoping constraint, so it is clear but not maximally differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It notes a real usage caveat ('Some channels have time limits on message deletion'), which implies when deletion will fail, but it never states when to use this tool versus alternatives (e.g., threads_delete, messages_edit) or any permission prerequisites. Usage is implied rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
messages_editAInspect
Edit the text of a message you already sent, in place. Supported on Telegram, WhatsApp, LiveChat, Max.ru, LINE, WeChat and similar channels; NOT supported on Gmail, Instagram, or LinkedIn (their platforms forbid editing) — those return an explicit error. Note: channels impose their own limits (own messages only, edit time windows such as ~48h on Telegram / ~15min on WhatsApp).
Get the message_id from messages.read_history (each row's id).
| Name | Required | Description | Default |
|---|---|---|---|
| new_text | Yes | Replacement text for the message. | |
| thread_id | Yes | Thread ID containing the message (numeric DB id or channel_ref like 'telegram:-100123'). | |
| message_id | Yes | ID of the message to edit (from messages.read_history). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so safety is partly covered. The description goes beyond this by disclosing own-messages-only restriction, channel-specific edit windows (~48h Telegram / ~15min WhatsApp), and explicit errors on unsupported platforms. It stops short of describing what happens to the prior text/formatting on success, so a 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Core action is front-loaded in the first clause, with support matrix and constraints following. The channel enumeration is somewhat lengthy but each list element carries routing value; there is no filler sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description covers the important non-obvious risks: unsupported-platform errors, ownership restriction, and time windows. It does not describe the success return payload, but for an in-place edit that is a minor omission given how thoroughly the failure modes are covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented, and the description's pointer to messages.read_history for message_id duplicates the schema note ('from messages.read_history'). Baseline 3 applies since the schema does the heavy lifting and the prose adds no new syntax or format detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Edit the text of a message you already sent, in place'), which cleanly distinguishes it from messages_send, messages_forward, and messages_delete. An agent can tell exactly what operation this performs without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names supported channels (Telegram, WhatsApp, LiveChat, Max.ru, LINE, WeChat) and excluded ones (Gmail, Instagram, LinkedIn), with the reason ('their platforms forbid editing') and the behavior on failure ('return an explicit error'). It also gives the prerequisite path for obtaining message_id via messages.read_history.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
messages_forwardADestructiveInspect
Forward a message from one thread to another. Supports native Telegram forwarding (preserves original sender attribution) and text-based forwarding for cross-channel scenarios.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| dest_thread_id | No | Destination thread to forward into. Provide at least one of dest_thread_id or recipient_name. To forward into the active conversation, pass the current thread_id. (If both are provided, dest_thread_id wins and recipient_name is ignored.) | |
| recipient_name | No | Name of person to forward to (channel auto-resolved). Provide at least one of dest_thread_id or recipient_name. Use only when forwarding to a different contact than the current conversation. | |
| source_thread_id | Yes | Thread containing the message to forward (e.g., 'telegram:123456' or numeric DB ID) | |
| source_message_id | Yes | ID of the message to forward |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the mutation (readOnlyHint=false, destructiveHint=true, idempotentHint=false), so the safety profile is covered. The description adds real behavioral nuance beyond that: native forwarding 'preserves original sender attribution' while text-based forwarding does not, which matters to an agent picking a mode. It stops short of disclosing permissions or effects on the destination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the core action front-loaded and the mode distinction immediately after. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a forward tool with no output schema and full schema coverage of parameters, the description supplies the key behavioral distinction an agent needs. It omits prerequisites (e.g., that the destination channel/contact must be resolvable) but is otherwise sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters, including the dest_thread_id/recipient_name precedence rules. The description adds no parameter-level detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Forward a message from one thread to another.' This is clear and distinguishable from siblings like messages_send or messages_edit. However, it never names or contrasts those siblings, which is what a 5 requires.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description hints at usage context by contrasting native Telegram forwarding against 'text-based forwarding for cross-channel scenarios,' but it never says when to choose this tool over messages_send or how the mode is selected. Usage is implied rather than instructed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
messages_press_buttonADestructiveInspect
Press a button under a message in a Telegram chat, as the connected Telegram account: the inline menus bots such as @BotFather reply with. Name the button by its exact label.
Returns the message as it looks after the press (bots usually edit it into the next menu) and anything the bot sent after it, each with its own buttons and a message_id you can pass back to press the next one.
Refused: buttons that need the account's cloud password (the owner confirms those in the Telegram app), and buttons that share a phone number or location, pay, or open a web app. A link button is not pressed; its URL is returned. A label that matches no button returns the buttons that exist.
Get the message_id from messages.read_history. Telegram user accounts only; a bot account cannot press buttons.
| Name | Required | Description | Default |
|---|---|---|---|
| thread_id | Yes | Thread containing the message (numeric DB id or channel_ref like 'telegram:123456789'). | |
| message_id | Yes | Message the button is under: a DB id from messages.read_history, or a 'telegram:<id>' returned by a previous press. | |
| button_text | Yes | The button's label, exactly as shown (case is ignored if nothing matches exactly). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only give the generic destructive/non-idempotent/openWorld profile, but the description discloses the specific refusals (cloud-password buttons, phone/location sharing, pay, web apps) and why ('the owner confirms those in the Telegram app'). It also describes the return payload — the edited message plus subsequent bot messages with their own buttons and message_ids — which no annotation conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the action and what is returned, then the refusals and prerequisites. Four sentences for this much behavior is efficient, though a couple of return-value details pad it slightly beyond the minimum an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so fully: post-press message state, follow-on bot messages, per-message buttons and message_ids. Account-type restriction, refusal cases, and the id-hop pattern round out a complete picture for a 4-param, non-idempotent, destructive tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the meaning of thread_id, message_id, button_text, and in_workspace is already documented, making 3 the baseline. The description adds only marginal param-level value: it reiterates the exact-label requirement (case ignored) and the message_id chaining workflow, both of which the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb+resource+scope: 'Press a button under a message in a Telegram chat, as the connected Telegram account,' and immediately narrows it to the inline menus bots like @BotFather reply with. This is unmistakably distinct from siblings such as messages_send, messages_read_history, or browser_click.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the prerequisite source of message_id ('Get the message_id from messages.read_history'), the account constraint ('Telegram user accounts only; a bot account cannot press buttons'), and the conditions under which the tool will not act. The alternatives are handled explicitly: link buttons are not pressed but returned, and non-matching labels return the buttons that exist.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
messages_read_historyRead-onlyIdempotentInspect
Read messages from a conversation thread. Use text_contains to find specific messages by content. text_contains searches the loaded part of the conversation and, when that is not all of it, asks the channel's own server (Telegram) to search the rest and stores the hits with their neighbouring messages. search_coverage says which: complete = true with a source, or complete = false with the reason, in which case an empty list is not proof the text is absent. Returns the most recent messages, including sender info and timestamps.
Voice calls: each row carries a meta object with allowlisted keys (event_type ∈ 'call_started'|'call_ended'|null, source ∈ 'voice_transcript'|null, call_id, speaker_display_name, duration_seconds, outcome, direction) plus per-message channel. To find calls without scanning every row, use calls.list_history instead.
Usage:
Get thread_id from threads.list first, OR
Use contact_name to auto-resolve thread_id
Examples:
User: 'show me messages from chat with [contact]' → read_history(contact_name='[contact]', limit=10)
User: 'last 5 messages from thread 571' → read_history(thread_id=571, limit=5)
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of messages to return (default: 10, max: 100) | |
| offset | No | Number of messages to skip (for pagination, default: 0) | |
| thread_id | No | Thread ID from threads.list (e.g. '571'), or a channel_ref (e.g. 'telegram:1306644770'). Optional if contact_name provided. | |
| contact_name | No | Contact/thread name to search for (optional if thread_id provided). Example: 'Jane Smith', 'John Doe' | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| text_contains | No | Filter: only return messages containing this text (case-insensitive substring match) | |
| include_outgoing | No | Include messages sent by you (default: true) | |
| channel_account_id | No | Connected account the conversation must belong to, when you know it (from channels.list). Scopes the lookup to that account, so the same person reached on two connected accounts cannot be confused for one, and a chat that does not exist on this account is 'not found' instead of someone else's thread with a matching id. |
messages_sendADestructiveInspect
Send a message to a thread, channel, or contact. Supports Telegram, Email, LinkedIn, and other connected channels. For LinkedIn posts (comment_thread kind), this posts a comment on the post. Can automatically resolve recipients and channels when not specified. Can send files/images/documents as attachments — pass attachments=[file_id, ...] with integer file IDs obtained from collections.list_files, search.files, or files.search. text is optional when attachments are provided.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | Message text to send. Optional if attachments provided. | |
| format | No | Message format | text |
| silent | No | Send without notification | |
| buttons | No | Inline reply buttons shown under the message, one per row (max 10). Each item is {"label": "<visible button text>"} plus exactly one target: "value" (a string sent back as a normal incoming message when tapped, defaults to the label, max 64 bytes), "url" (opens a link) or "web_app" (opens an https page as a Telegram Mini App). Only on channels that support buttons (Telegram bot accounts); can accompany at most one attachment, and when the text is too long for a media caption the buttons arrive with the text as a second message. | |
| template | No | WhatsApp Business API only. An approved template to send instead of free text, in Meta's shape: {name, language, components:[{type:'body', parameters:[{type:'text', text:...}]}, ...]}; `language` may be the code as a string ('en_US') or Meta's {code:'en_US'}. Required when the 24-hour window of the conversation is closed (the send answers WINDOW_EXPIRED otherwise). Read the approved templates and their variables with channels.list(account_id=...). | |
| thread_id | No | Target thread. OMIT to reply in the same chat you received the triggering message from — the backend defaults to the current thread. Pass an explicit value ONLY to reply in a DIFFERENT thread, and only use: (a) a numeric DB thread id from search.threads, or (b) a channel_ref like 'telegram:-12345'. To text a phone number, pass 'sms:+<number in international format>': the message goes out from a workspace phone number that has SMS turned on (from_account_id picks which one) and lands in that number's conversation with the recipient, next to their calls. NEVER use a chat-type word (dm, group, channel, livechat) — those are category labels from the SITUATION block, not ids. | |
| attachments | No | Array of integer file IDs to send as attachments (images, documents, any files). Get file IDs from collections.list_files (field `file_id`), search.files (field `file_id`), or files.search — only ids returned by those calls in this workspace. The file must already exist in the workspace (status=ready) — no separate upload step needed. When attachments are provided, `text` becomes optional (a caption can be included alongside). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| private_reply | No | Comment threads only (Instagram, Facebook): answer the comment in the commenter's private messages instead of publicly under the post (Meta's Private Replies). One message per comment, within 7 days of the comment. The message is filed in the person's private conversation. To also answer publicly, make a second call without this flag. OMIT for a normal send. | |
| recipient_name | No | Name of person to send to (e.g., 'Jane', 'John'). Tool will auto-resolve channel. Optional if thread_id provided. | |
| from_account_id | No | Which of the workspace's accounts on this channel SENDS the message, i.e. the number or handle the recipient sees. Only meaningful when starting a NEW conversation — an existing thread already belongs to an account and that one is used. OMIT and the platform picks the most recently active account, which is a coin flip in a workspace with several numbers: pass it whenever one of them must not be used to open conversations (a personal line, or one under a spam restriction). An account that is not active in this workspace is refused, never silently swapped. | |
| recipient_username | No | Telegram @username to message (e.g. '@some_username'). Use this for a Telegram user NOT yet in contacts — it resolves the handle, adds the contact, and creates the thread. Telegram only; for existing contacts prefer thread_id or recipient_name. | |
| reply_to_message_id | No | ID of message to reply to (optional) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as non-readonly, destructive, non-idempotent, open-world, so the safety bar is covered. The description adds real behavior beyond them: automatic recipient/channel resolution, attachments requiring pre-existing ready file IDs with no upload step, and text becoming optional when attachments are present. It says nothing about rate limits or failure modes, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, and the attachment/comment specifics follow in a logical order. Slight redundancy: the attachments contract is explained in both the description and the schema's attachments entry, but overall it is tight and earns its sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 13 optional parameters, no required fields, no output schema, and full schema coverage plus rich annotations, the description supplies the key nuances an agent needs (auto-resolution, comment_thread posting, attachment sourcing, optional text). It omits return/response behavior, but for a send tool with these structured fields that is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the per-parameter descriptions already carry deep semantics (window expiry, private replies, account selection). The description does add useful cross-references for where file IDs come from, but it largely restates the attachments contract the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Send a message to a thread, channel, or contact') and clarifies channel coverage (Telegram, Email, LinkedIn) plus the special comment_thread behavior. It does not, however, differentiate itself from near-siblings like messages_send_email or messages_forward, and listing 'Email' as a supported target partially overlaps with messages_send_email.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete context: 'automatically resolve recipients and channels when not specified', and points to the sibling sources for attachment IDs (collections.list_files, search.files, files.search). It stops short of explicitly naming when to prefer alternative send/forward/edit tools, so it stops at clear context rather than full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
messages_send_emailADestructiveInspect
Compose and send an email — with subject, CC/BCC, and attachments. Use for email; for chat messages (Telegram/WhatsApp/livechat) use messages.send instead.
| Name | Required | Description | Default |
|---|---|---|---|
| cc | No | Email addresses to CC. OMIT to skip. | |
| bcc | No | Email addresses to BCC. OMIT to skip. | |
| text | No | Email body. | |
| subject | No | Email subject line. Required for new emails; for replies it auto-generates 'Re: ...' when omitted. | |
| attachments | No | Array of integer file IDs to attach. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| recipient_email | No | Recipient email address (e.g. 'john@example.com'). Provide to start a new email thread; OMIT to reply in the current email thread. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true, idempotentHint=false and readOnlyHint=false, so the agent knows this is an irreversible external write. The description adds no behavioral context beyond that (no irreversibility warning, no auth/recipient-verification note, no delivery semantics), so with annotations carrying the safety profile a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the capability statement front-loaded and the routing rule immediately after. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with full schema coverage and complete annotations, the description covers purpose and channel routing adequately; the reply-vs-new-thread and workspace-scoping behavior are fully specified in the schema. It stops just short of noting that sending is irreversible, but nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%; the description only echoes field names (subject, CC/BCC, attachments) already documented in the schema. It adds no format or reply-vs-new-thread semantics beyond what the schema's own parameter descriptions provide, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb pair ('Compose and send') and resource ('an email') plus the salient fields (subject, CC/BCC, attachments). It explicitly distinguishes itself from the chat-messaging sibling messages.send, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative explicitly ('for chat messages (Telegram/WhatsApp/livechat) use messages.send instead') and gives the condition that selects it (email vs chat channel). This is the when/when-not/alternative pattern in two sentences.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notes_deleteADestructiveIdempotentInspect
Delete a note by ID from the target notebook. Same identity rules as notes.save — agents can only delete from their own notebook.
| Name | Required | Description | Default |
|---|---|---|---|
| note_id | Yes | ID of the note to delete | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| target_agent_id | No | Target notebook. In agent mode optional (defaults to your own). From MCP, omit it to delete a workspace note; pass an agent id for that agent's own note. Agents cannot target other agents' notebooks. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description adds the important constraint that agents can only delete from their own notebook, which is genuine context beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and the identity constraint second. No wasted filler, though the 'notes.save' reference adds a small ambiguity rather than clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive delete with annotations covering destructiveness and idempotency, plus a fully documented schema, the description supplies the one extra thing needed: the identity restriction. Return-value detail is not required since no output schema exists but a simple delete is self-explanatory.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so note_id, in_workspace, and target_agent_id are already documented in the schema. The description only restates 'by ID' and 'target notebook', adding no syntax or behavior beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete a note by ID from the target notebook'), so the agent immediately knows what it does. It is unambiguous against siblings like notes_search and notes_recall, though it does not explicitly name a differentiating alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an identity-scoping rule ('agents can only delete from their own notebook'), which is useful context, but gives no explicit when-to-use/when-not guidance or named alternatives. The reference to 'notes.save' does not match any sibling (notes_save exists), which is a minor source of confusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notes_recallARead-onlyIdempotentInspect
Recall notes from your notebook. By default returns only your own notes (all scopes, newest first). Pass filter_agent_id= to read another agent's notebook, or filter_agent_id="all" (or "*") to read across every agent in the workspace. Pass scope to narrow to global/thread/person. Each result includes agent_id and agent_name of the author.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | Recall a specific note by key | |
| limit | No | Max notes (default 20, max 50). Newest first. | |
| scope | No | Optional filter: global | thread | person. Omit for all scopes. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| scope_ref_id | No | Filter by specific thread_id or person_id | |
| filter_agent_id | No | Optional. Omit to read your own notes (an agent: its notebook; a person over MCP: the whole workspace). Pass a numeric agent_id as a string (e.g. "57") to read another agent's notebook (read-only). Pass "all" or "*" to read across all agents in the workspace. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, idempotent, non-destructive, and closed-world. The description adds genuinely new behavioral context beyond them: the default ownership scope, newest-first ordering, that cross-agent access is read-only, and that results carry agent_id/agent_name of the author. It stops short of describing pagination or volume, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core purpose followed by default behavior and the two filtering modes, then result fields. Tight overall, though the filter_agent_id detail overlaps with the schema description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema absent, the description usefully notes the author fields returned, and annotations cover the safety profile. Ownership defaults, filtering modes, and ordering are all addressed, so an agent has what it needs to call it correctly; only return volume/pagination is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema. The description largely restates filter_agent_id and scope behavior rather than adding syntax, formats, or edge cases the schema lacks. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Recall notes from your notebook') and immediately defines scope behavior (own notes by default, all scopes, newest first). It does not explicitly distinguish itself from the closest sibling, notes_search, so an agent must infer the recall-vs-search boundary. Clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear operating context: default returns own notes, filter_agent_id=<int> reads another agent's notebook, "all"/"*" reads across the workspace, and scope narrows the result set. There is no explicit when-not guidance or naming of notes_search as the alternative for keyword queries, which keeps it at 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notes_saveAInspect
Save a fact or note into the agent's memory. Use scope to choose visibility: 'workspace' = visible to every agent in this workspace (use for shared facts, project conventions); 'agent' = private to this agent (use for personal working notes); 'thread' = scoped to one conversation (use for thread-specific reminders); 'person' = scoped to one contact (use for per-contact context). If a note with the same key+scope exists it will be updated. Do NOT use this tool for behavioral rules or corrections — use feedback.save for those.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Short identifier for this note (must not start with '__' — reserved) | |
| scope | Yes | Scope of the note. 'workspace' = shared across all agents; 'agent' = private to this agent (was 'global' pre-PR1); 'thread' = per-conversation; 'person' = per-contact. 'global' is accepted as a deprecation alias for 'agent'. | |
| value | Yes | The note content | |
| pinned | No | Pin this note so it's always loaded first. Default false. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| scope_ref_id | No | Reference ID — thread_id (for scope=thread) or person_id (for scope=person). Required for thread/person scope. In MCP mode (no thread context), must be passed explicitly. | |
| target_agent_id | No | Target notebook. In agent mode optional (defaults to your own); required from MCP. Agents cannot target other agents' notebooks. Ignored when scope='workspace' (workspace memory is shared). | |
| expires_in_hours | No | Auto-delete after N hours. Omit for permanent notes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-read-only, non-idempotent write with no destruction, and the description adds the crucial upsert semantic ('If a note with the same key+scope exists it will be updated') that the annotations alone do not convey. It stops short of describing return shape or failure behavior, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then scope semantics, then upsert behavior, then the negative routing in a single tight block. It is dense and enumerates four scopes, but every clause carries actionable information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and 100% schema coverage, the description covers scoping, upsert behavior, and tool selection. It omits return/failure behavior and the deprecation alias handling, which the schema does carry, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds usage-oriented meaning for the scope enum beyond the schema's definitions (which scenario each scope is for). It does not add syntax for the other seven parameters, which the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Save a fact or note into the agent's memory') and clearly separates itself from siblings notes_recall, notes_search, and notes_delete. It also names the tool it is not (feedback.save), so an agent can distinguish it without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives per-scope when-to-use guidance ('workspace' for shared facts/conventions, 'agent' for personal working notes, etc.) and an explicit exclusion: 'Do NOT use this tool for behavioral rules or corrections — use feedback.save for those.' Both the when and the when-not are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notes_searchARead-onlyIdempotentInspect
Full-text search in your notebook. By default searches only your own notes. Pass filter_agent_id= to search another agent's notebook, or "all" (or "*") for workspace-wide. Or list all notes for a person/thread by scope_ref_id.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results (default 10, max 50) | |
| query | No | Text to search for in note keys and values. Optional if scope_ref_id is provided. | |
| scope | No | Limit search to scope | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| scope_ref_id | No | Filter by specific thread_id or person_id. If provided without query, lists all notes for that ref. | |
| filter_agent_id | No | Optional. Omit to search your own notes (an agent: its notebook; a person over MCP: the whole workspace). Pass a numeric agent_id as a string (e.g. "57") to search another agent's notebook (read-only). Pass "all" or "*" to search across all agents in the workspace. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely useful behavior the annotations do not: the default is scoped to your own notes, other agents' notebooks are read-only, and the call can be retargeted to the workspace-wide corpus.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the core purpose front-loaded and the three search modes following in priority order; no filler. The final sentence is a fragment but remains readable and information-dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, 6-parameter, fully schema-documented tool, the description covers purpose, default scope, and mode switching adequately. It does not describe the return shape (note keys/values, result ordering), which is a minor gap given there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents filter_agent_id, scope_ref_id, scope, and limit in comparable detail, so the description is largely consolidating rather than adding. It also writes filter_agent_id=<int> while the schema specifies a numeric string (e.g. "57"), a small mismatch that could mislead on type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Full-text search in your notebook') and immediately clarifies scope behavior, so the agent knows it is a search/read tool. It does not, however, distinguish itself from close siblings such as notes_recall or notes_save, so the reader must infer the boundary from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete mode-selection guidance: default is own notes, filter_agent_id for another agent's notebook, 'all'/'*' for workspace-wide, and scope_ref_id for listing a person/thread. It never names an alternative tool (e.g. notes_recall) or states when not to use this one, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onboarding_start_signupAInspect
🚪 Start a person's own DialogBrain account without a password.
Creates the account if the email is new (it keeps no password, like a Google sign-in), mails a 6-digit sign-in code to that address, and returns the link to send them. Use it for a lead who ASKED to try the product themselves, or who ticked 'create my account' on a form — never for someone who did not ask.
Tell them the code is in the email and give them the returned link: they enter the code and land on the screen that connects their channel. Do NOT ask them for a password and do not send one; there is none. The code is not returned to you, so you cannot read it out — if they say it never arrived, call again after a minute (a second call inside 60s is refused with rate_limited).
| Name | Required | Description | Default |
|---|---|---|---|
| Yes | The address the person gave. The code goes here, nowhere else. | ||
| channel | No | The channel they want to run their agent on (e.g. 'whatsapp', 'telegram', 'instagram'), usually the one they picked on the form. It decides which connect dialog the returned link opens. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (which only declare a non-readonly, open-world write): it discloses the passwordless model, that the code is mailed but never returned to the caller, that a second call within 60s is refused with rate_limited, and that no password exists to ask for. These are exactly the behavioral traits an agent needs and cannot get from the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence and the subsequent sentences are dense operational guidance (code delivery, link handoff, retry behavior) rather than filler. It is somewhat long and the emoji is decorative, but almost every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Though there is no output schema, the description explains what comes back (the link to send) and how to handle the non-returned code, and it covers the auth model, rate-limit behavior, and the post-signup handoff screen. An agent has everything needed to call this correctly and instruct the user.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents email (code destination), channel (which connect dialog opens), and in_workspace (single-call workspace override). The description reinforces the email/code relationship but adds little parameter-level meaning the schema lacks, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first line states a specific verb+resource ('Start a person's own DialogBrain account without a password') and the body elaborates exactly what happens: creates the account if the email is new, mails a 6-digit code, and returns a link. This is a passwordless signup flow that is unmistakably distinct from sibling tools like workspace_invite or agents_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit when ('a lead who ASKED to try the product themselves, or who ticked create my account on a form') and a clear when-not ('never for someone who did not ask'), plus operational guidance for retries after 60s. It stops short of naming the alternative tool to use for people who did not ask, which is the only gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
present_statusARead-onlyIdempotentInspect
Ask what the room is ACTUALLY seeing on the screen-share right now. Returns presenting plus the page_id, url and title read live from the shared tab. Use it before telling participants what is on screen, and whenever you are not certain the share still points at the tab you meant — pressing keys or scrolling does NOT change which tab is shared, only present_tab does.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, but the description adds real behavioral context beyond them: it clarifies that pressing keys or scrolling does NOT change the shared tab, only present_tab does. It reads the state live from the shared tab, which explains why it is a safe, side-effect-free check.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences with zero waste: what it reads, what it returns, and the caveat about tab switching. Emphasis on 'ACTUALLY' signals the point-of-truth framing efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description enumerates the return fields, and annotations cover the safety profile. For a read-only status check with one optional param, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single optional parameter at 100% schema description coverage, the schema already documents in_workspace fully; the description adds nothing about it. Baseline 3 is appropriate since the schema carries the parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Ask what the room is ACTUALLY seeing on the screen-share') and enumerates the returned fields (presenting, page_id, url, title). It explicitly distinguishes itself from the sibling present_tab, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete triggers ('before telling participants what is on screen', 'whenever you are not certain the share still points at the tab you meant') and names the alternative that changes the shared tab. The when-to-use is explicit and tied to a named sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
present_stopARead-onlyIdempotentInspect
Take whatever is on the call's screen-share back down, so the room sees nobody presenting. Takes no arguments — it stops whatever THIS call is showing. Use it when you are done with a tab and the room should stop looking at it; present_tab with a different page_id switches the share instead, and browser.close on the shared tab also ends it. Stopping when nothing is being shared is a success, not an error — the room already sees nothing. Returns presenting=false once the share is down; present_status will then report nothing on screen. If it comes back ok=false the share may STILL BE UP — do not tell the room you stopped it; retry, or call present_status to see what the room can actually see.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (which already declare readOnly/idempotent/non-destructive): it discloses that stopping with nothing shared is a success not an error, that the return is presenting=false, and critically that ok=false means the share may STILL BE UP with retry/present_status guidance. This is exactly the failure-mode context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, and every subsequent sentence carries distinct value (alternatives, idempotent-success semantics, return value, failure handling). It runs long, but the density of unique information justifies the length rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by explaining both the success return (presenting=false) and the failure signal (ok=false, share still up) plus the follow-up tool. For a 1-param mutation-of-share-state tool, an agent has everything needed to call it and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so in_workspace is already fully documented in the schema, and the description adds no format or semantics beyond it. The phrase 'Takes no arguments' is about the absence of required targeting args but slightly understates the optional in_workspace override; baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('take whatever is on the call's screen-share back down, so the room sees nobody presenting') and scopes it to THIS call. It explicitly distinguishes itself from siblings present_tab and browser.close, so an agent can pick correctly without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('when you are done with a tab and the room should stop looking at it') plus named alternatives with the condition that selects them: present_tab with a different page_id switches the share, browser.close on the shared tab also ends it. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
present_tabARead-onlyIdempotentInspect
Put one of the agent's browser tabs on the live call's screen-share. Pass the page_id you got from browser.open. Only usable while the agent is in an active voice call. THIS IS THE ONLY WAY TO CHANGE WHAT THE ROOM SEES — scrolling or pressing keys changes the page, never which tab is shared. The shared tab stays the active share until you call present_tab with a different page_id, call present_stop to take it down, close the tab via browser.close, or the call ends. Returns the page_id, url and title actually on screen; check them before telling anyone what they are looking at, and use present_status to re-check later.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | page_id returned by browser.open for the tab you want to share. Must be a tab still open in the agent's browser context. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly, idempotent, non-destructive), and the description adds genuinely extra context: the share persists until a different page_id is passed, present_stop is called, the tab is closed, or the call ends. It also warns to verify the returned url/title before describing the screen to others. No contradiction with the readOnly/high-level hints, since this mutates the call presentation rather than stored data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and dense, with the critical constraint (only in an active call) and the anti-pattern (scrolling doesn't change the share) called out emphatically. Slightly long, but nearly every clause carries operational information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by describing what is returned (page_id, url, title actually on screen) and how to re-verify state via present_status. Combined with the lifecycle rules for the share, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so page_id's origin and the in_workspace override are already documented in the schema. The description restates the browser.open source but adds no syntax, format, or semantics beyond it, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a concrete verb+resource+destination: 'Put one of the agent's browser tabs on the live call's screen-share.' An agent can immediately distinguish this from siblings like present_stop, present_status, and the browser_* tab tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the prerequisite ('Only usable while the agent is in an active voice call'), the source of the required argument (page_id from browser.open), and the routing rule against the tempting alternatives: scrolling or key presses change the page, never the shared tab. It also names present_stop as the way to take the share down and present_status for re-checking.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prompts_blocks_getARead-onlyIdempotentInspect
Read the stored prompt-block overrides. Without agent_id: the workspace level (applies to every agent). With agent_id: that agent's own layer (beats the workspace per field) plus its legacy include_* toggles. Missing keys and null fields mean 'inherit'.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent id for the agent level; omit for the workspace level. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower. The description adds real behavioral context beyond annotations: field-level precedence (agent layer beats workspace) and that missing keys/null mean 'inherit'. It does not cover return shape or the legacy include_* toggle semantics in detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the operation, then the two agent_id modes, then the inheritance rule. No filler; every sentence adds semantic value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with no output schema, the description supplies the critical return semantics (inheritance, null/missing = inherit) that an agent needs to interpret results. It could say more about the block payload structure, but coverage is strong for the complexity involved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes further by explaining the layering and precedence semantics of agent_id (workspace vs agent level, per-field override, inheritance), which the schema does not capture. in_workspace is left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Read the stored prompt-block overrides.' It's immediately clear this is a read operation on prompt block overrides. It does not, however, name the distinguish siblings (prompts_blocks_registry, prompts_blocks_preview), so an agent must infer the difference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly explains the two call modes: omitting agent_id returns the workspace level, providing agent_id returns that agent's layer. This is essentially when-to-use guidance driven by intent. It does not name alternative tools or describe the registry/preview siblings, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prompts_blocks_previewARead-onlyIdempotentInspect
Assemble an agent's system prompt through the REAL builder and return it block by block with attribution (id, kind, cache phase, token estimate), the cached-prefix/dynamic-suffix split, and any ignored stale overrides or missing translations. Byte-equality with the production assembly is verified — the preview cannot lie. Use after prompts.blocks_update to see the effect.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Agent whose prompt to assemble. | |
| thread_id | No | Optional thread to load real situation/history context from. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior, so the bar is lower, yet the description adds real value: 'Byte-equality with the production assembly is verified — the preview cannot lie' declares fidelity guarantees beyond the annotations, and it notes ignored stale overrides/missing translations in the output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the core action and followed by the usage trigger. Every clause carries information, though the first sentence is packed with several enumerated items that could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only preview with full schema coverage and no output schema, the description compensates by describing the return structure (per-block attribution, cached/dynamic split, ignored overrides). The behavioral and output picture is essentially complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so agent_id, thread_id and in_workspace are already documented in the schema. The description adds no parameter-level syntax or format detail beyond what structured data provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (assemble/preview) and resource (an agent's system prompt) and details the exact return shape — block-by-block attribution, cache-phase split, ignored overrides. This is far more specific than the sibling prompts_blocks_get and lets an agent distinguish the preview function without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use after prompts.blocks_update to see the effect,' giving a concrete trigger and pointing at the related sibling workflow. It stops short of naming any when-not conditions or contrasting with prompts_blocks_get, so it is clear context without full exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prompts_blocks_registryARead-onlyIdempotentInspect
List every block the platform assembles into agent system prompts: id, kind (fact = platform data, guidance = platform text you can rewrite or disable, user = your own text), cache phase, order, the fact labels you can reference from your instructions (e.g. thread_id in [SITUATION]), which labels are locked by enabled tools, and each guidance text's platform default with its md5. Read this before prompts.blocks_update.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the read-only, idempotent, non-destructive safety profile, and the description adds genuinely useful context: the fact/guidance/user kind taxonomy, that guidance text can be rewritten or disabled, and that some labels are locked by enabled tools. It doesn't mention pagination or auth requirements, but the added domain context is substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the verb and resource, and every clause (kind taxonomy, cache phase, order, locks, md5) conveys information about the return payload rather than filler. It is a single dense run-on sentence, which slightly hurts readability but not wastefulness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and only one param, the description carries the burden of explaining return contents and does so thoroughly, enumerating id, kind, cache phase, order, labels, and md5 defaults. It leaves only minor gaps such as pagination, which are low-impact for a read-only registry listing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single optional in_workspace parameter whose meaning is fully documented in the schema (100% coverage), covering the workspace-override semantics. The description adds no parameter-level detail, so the baseline 3 for high schema coverage is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('List every block the platform assembles into agent system prompts') and even enumerates the returned fields, so the agent knows exactly what this produces. It doesn't, however, differentiate itself from the closest siblings prompts_blocks_get or prompts_blocks_preview, which an agent would need to tell apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit sequencing guidance ('Read this before prompts.blocks_update'), which clearly signals its role in the block-editing workflow. It does not name alternatives or state when NOT to use it versus get/preview, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prompts_blocks_updateAInspect
Replace the stored prompt-block overrides at one level (workspace when agent_id is omitted, else that agent). prompt_blocks is the FULL desired object for the level: {"": {"enabled": bool|null, "text": str|null (guidance blocks only), "order": int|null (reorders within the block's cache tier), "hidden_rows": [label, ...] (fact blocks only)}}. Pass null for an entry to remove it (inherit). Fails loudly: unknown block/label ids, editing a fact block's text, empty text, size caps, and newly hiding a fact row that an enabled tool consumes are all rejected with the full error list — read prompts.blocks_registry for valid ids. Pre-existing stored state is never re-validated (grandfathered).
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent id for the agent level; omit for the workspace level. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| prompt_blocks | Yes | Full desired overrides object for the level (see tool description for the entry shape). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing the full-object replacement semantics, null-to-inherit behavior, the exhaustive validation failure list (unknown ids, editing fact text, empty text, size caps, hiding a consumed fact row), and the grandfathered re-validation exemption. This is rich behavioral context an agent needs before mutating.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core operation and target level, then details the entry shape and failure modes. Dense and information-packed with little waste, though the inline JSON shape makes sentences long and harder to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param nested-object mutation with no output schema, the description covers shape, semantics, and failure modes thoroughly. It does not describe the successful return/response, which is the main remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning: the per-block entry shape (enabled/text/order/hidden_rows), which fields apply to guidance vs fact blocks, and null semantics. in_workspace's non-persistence is also reinforced.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Replace) and resource (stored prompt-block overrides at one level), and disambiguates the target level via agent_id presence. An agent can distinguish it from prompts_blocks_get/preview/registry without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the full-replace contract, that prompt_blocks is the FULL desired object, how to remove an entry (null → inherit), and points to prompts.blocks_registry for valid ids. It gives clear context but does not explicitly contrast with siblings like prompts_blocks_preview or prompts_update.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prompts_getARead-onlyIdempotentInspect
Get full content of a prompt template: system instructions (prompt_text) and auto-reply rules.
Run prompts.list first to find the prompt_id.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt_id | Yes | ID of the prompt template to fetch | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds genuinely useful disclosure about what the payload contains (prompt_text and auto-reply rules), which matters because there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the return content front-loaded and the prerequisite trailing. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description helpfully names the returned content, and the annotations carry the safety profile. It is nearly complete for a simple read tool, lacking only notes on error behavior for a missing/invalid prompt_id.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both prompt_id and in_workspace are already fully described in the schema. The description only reinforces prompt_id by referring to finding it via prompts.list and adds no syntax or semantics beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get full content of a prompt template') and enumerates exactly what is returned (system instructions / prompt_text and auto-reply rules). An agent can distinguish it from prompts_list, prompts_update, and prompts_blocks_get without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit prerequisite: 'Run prompts.list first to find the prompt_id.' That is clear operational context, though it names no when-not condition or alternative sibling (e.g., prompts_blocks_get for block content).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prompts_listARead-onlyIdempotentInspect
List all prompt templates in this workspace.
Returns id + name + description + category so you know which prompt_id to use in prompts.get or prompts.update.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so safety is covered; the description adds that the workspace scope comes from the session and that results include id/name/description/category. It does not mention pagination or result limits for a list call, which is the remaining behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first states the action, the second explains the return payload and its use. Every sentence earns its place and the core purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by enumerating the returned fields (id, name, description, category), which is what an agent needs to wire the result into prompts.get/update. Only pagination/ordering behavior is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single optional in_workspace parameter is fully documented in the schema (including its non-persistent, session-local semantics). The description adds nothing about parameters, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('all prompt templates in this workspace') and explicitly names the downstream tools (prompts.get, prompts.update) that consume the returned prompt_id. An agent can tell this apart from siblings like prompts_prompt_history or agents_list_drafts from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context: call it to obtain the prompt_id needed by prompts.get or prompts.update. It does not state when NOT to use it (e.g., to inspect a version history vs. current templates), so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prompts_prompt_historyARead-onlyIdempotentInspect
List past versions of a prompt template's prompt_text. Every edit is snapshotted to an append-only table — use this to browse history and find a version_number for prompts.prompt_restore.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max versions to return (1-200, default 50) | |
| prompt_id | Yes | ID of the prompt template | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| before_version | No | Cursor: return versions strictly below this version_number |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive, closed-world behavior, so the safety bar is met. The description adds real value beyond them by disclosing the append-only snapshot model ('every edit is snapshotted to an append-only table'), which tells the agent the history is immutable and complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action stated first and the follow-up use (finding a version_number for restore) second. No filler or restated title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by explaining what is returned (past versions of prompt_text) and how to use a returned version_number. Combined with annotations covering safety and a fully documented schema, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so limit, prompt_id, in_workspace, and the before_version cursor are already documented in the schema. The description references the concept of a version_number but adds no format or pagination guidance beyond the schema; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'List past versions of a prompt template's `prompt_text`'. The 'past versions' scope implicitly separates it from the current-state siblings like prompts_get, though no sibling is named as an exclusion. Clear enough for an agent to pick correctly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit usage context ('use this to browse history and find a version_number') and names the downstream tool `prompts.prompt_restore` that consumes the output. It lacks a when-not clause (e.g. use prompts_get for the current version), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prompts_prompt_restoreAInspect
Restore a past version of a prompt template by version_number. Creates a new version pointing at the restored content — history is preserved. Fans out to every agent using this template without a per-agent override; the response includes affected_agents as a receipt of the fan-out.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | Optional: why this restore is happening (shows up in history UI) | |
| prompt_id | Yes | ID of the prompt template | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| version_number | Yes | The version_number to restore (get it from prompts.prompt_history) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (which already say readOnlyHint=false, destructiveHint=false, idempotentHint=false), the description adds material behavioral context: history is preserved because restore creates a new version, and it fans out to every agent using the template without a per-agent override, returning affected_agents as a receipt. The fan-out side effect is exactly the kind of non-obvious behavior an agent must know before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences, front-loaded with the core action and parameter, then consequences, then the return receipt. No filler and no repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing the response's affected_agents field, and it covers the mutation's side effects and the history-preserving guarantee. An agent has everything it needs to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented, and the schema even explains that version_number comes from prompts.prompt_history. The description restates version_number's meaning but adds nothing for reason or in_workspace, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (restore), resource (a past version of a prompt template), and the driving parameter (version_number). It also discloses the resulting state (a new version pointing at restored content) which distinguishes it from a generic update tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It points the agent at prompts_prompt_history as the source of version_number, establishing the correct workflow. It gives clear context for when to use it but names no explicit alternative or 'when-not' condition (e.g., use prompts_update to make a fresh edit instead of restoring).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prompts_updateAInspect
Update a prompt template's name, system instructions, or auto-reply rules.
Changes affect every agent using this template, unless the agent has its own override (set via agents.update → prompt_text).
All parameters except prompt_id are optional — only provided fields are updated.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | New name for the prompt template | |
| prompt_id | Yes | ID of the prompt template to update | |
| description | No | New description for the prompt template | |
| prompt_text | No | The AI system prompt: persona, tone, rules, behavior. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| auto_reply_rules | No | Pre-classifier rules that run BEFORE the main AI. Format: bullet list of conditions → actions (SKIP / SIMPLE_REPLY / SEARCH / CALENDAR). Pass null to clear. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=false. The description adds a genuine blast-radius disclosure (affects every agent on the template) and partial-update semantics (only provided fields are updated), which is more than the annotations give. But it omits permissions requirements and whether changes take effect immediately, so it is a modest rather than rich addition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, purpose front-loaded, followed by scope impact and update semantics. No filler or restated boilerplate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-destructive update tool with no output schema and fully documented params, the definition covers purpose, scope, override interaction, and partial-update behavior. It stops short of noting required permissions or the effect-on-response, which are minor gaps rather than blocking omissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes slightly beyond by grouping the fields conceptually and stating that all parameters except prompt_id are optional with only supplied fields applied, adding PATCH-like semantics not spelled out in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Update) and resource (prompt template) and enumerates the mutable fields (name, system instructions, auto-reply rules). It is clearly distinguishable from read siblings like prompts_get/prompts_prompt_history, but it never names those siblings explicitly to sharpen the contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides real context: changes affect every agent on the template unless the agent carries its own override set via agents.update → prompt_text. That gives the agent a routing signal toward the per-agent alternative. However, it does not state a positive when-to-use condition or explicit exclusion territory against the sibling prompt tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reminder_cancelBInspect
Cancel an active reminder by its trigger ID.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent ID (required when calling from MCP; ignored in agentic mode). | |
| trigger_id | Yes | ID of the reminder to cancel | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, so the safety profile is largely covered. The description adds the meaningful constraint that only *active* reminders can be cancelled, but says nothing about error behavior, permissions, or whether a cancelled reminder can be restored.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single eight-word sentence with zero padding, front-loading the verb and resource immediately. Nothing in it is redundant or buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-purpose mutation tool with no output schema and annotations covering the safety profile, the description is minimally sufficient. It omits, however, how to obtain a valid trigger ID and what the caller should expect on a stale or non-existent ID, both of which matter for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all three parameters (agent_id, trigger_id, in_workspace) are documented in the schema itself. The description adds no syntax or semantic detail beyond repeating the trigger_id concept, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Cancel) and resource (an active reminder) plus the identifier used to select it (trigger ID). It is clearly distinguishable from the sibling reminder_set and reminder_list by verb, though the description never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of alternatives such as reminder_list (to discover a trigger ID) or reminder_set (to create one). The 'active' qualifier implies already-fired reminders are not cancellable, but this is inference, not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reminder_listBRead-onlyIdempotentInspect
List your active reminders (both one-time and recurring).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results (default 20) | |
| agent_id | No | Only this agent's reminders. Leave it out to list the whole workspace's; each row then names the agent it belongs to. Ignored in agentic mode. | |
| thread_id | No | Filter by thread | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| include_fired | No | Include already-fired one-time reminders (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so safety is covered. The description adds the useful scope detail that only 'active' reminders are returned and that both one-time and recurring types are included, which aligns with the include_fired default of false. It stops short of describing pagination or the returned row shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler, and the key scope qualifier ('active') is front-loaded. Nothing could be trimmed without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a read-only list tool whose annotations cover the safety profile and whose schema covers all parameters, but with no output schema the description could have sketched what a returned row contains (e.g. the owning agent) or the default exclusion of fired reminders.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the schema, including agent_id scoping and include_fired. The description only loosely gestures at one of them ('one-time and recurring'), adding no syntax or format detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List your active reminders') and scopes it with 'active' plus 'both one-time and recurring'. It does not explicitly distinguish itself from siblings reminder_set and reminder_cancel, but the read/list nature is unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance, and no mention of the sibling tools for creating or cancelling reminders. The agent must infer that this is the retrieval counterpart to reminder_set/reminder_cancel.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reminder_setAInspect
Schedule a reminder. One-time reminders fire at a specific datetime. Recurring reminders fire on a schedule (daily, weekly, every N days, or every N minutes). Optionally scope to a thread or target another agent.
| Name | Required | Description | Default |
|---|---|---|---|
| time | No | Time of day HH:MM for daily/weekly/every_n_days (e.g. '09:00'). Required for daily/weekly/every_n_days. | |
| reason | Yes | What this reminder is for (you'll see this when it fires) | |
| agent_id | No | Agent ID (required when calling from MCP; ignored in agentic mode). | |
| datetime | No | ISO datetime for one_time. Required for one_time. Either with an offset ('2026-04-01T09:00:00+03:00') or without one plus `timezone` ('2026-04-01T09:00:00' + timezone 'Europe/Moscow'): a datetime with no offset is read in `timezone`, and in UTC only when that is omitted. | |
| timezone | No | IANA timezone (e.g. 'Europe/Moscow'). Defaults to UTC. | |
| thread_id | No | Optional thread ID to scope the reminder to. Omit for workspace-level reminders. | |
| days_of_week | No | Days for weekly: 0=Mon, 1=Tue, 2=Wed, 3=Thu, 4=Fri, 5=Sat, 6=Sun. Required for weekly. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| interval_days | No | For every_n_days: fire every N days (min 2). | |
| schedule_type | Yes | one_time = fires once at datetime. daily = fires daily at time. weekly = fires on specific days_of_week at time. every_n_days = fires every N days at time. interval = fires every N minutes. | |
| interval_minutes | No | For interval: fire every N minutes (5-1440). | |
| target_agent_slug | No | Optional: activate a different staff member instead of yourself when the reminder fires. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the description only needs to add context. It usefully discloses that reminders fire on a schedule and can be scoped to a thread or another agent, but says nothing about duplicate-creation behavior (idempotentHint=false), permissions, or limits on how many reminders can be scheduled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the action front-loaded and no padding; the schedule taxonomy in sentence two is the only part that partially duplicates the schema's schedule_type enum. Efficient overall, with a minor redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter mutation tool with no output schema and safety covered by annotations, the description covers the conceptual model (one-time vs recurring, scoping) adequately, and the schema fills in every parameter's format. It omits only what happens at fire time and delivery semantics, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 12 parameters in detail. The description restates the schedule families and the optional thread/agent scoping at a high level but adds no syntax or format detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ("Schedule a reminder") and immediately enumerates the scheduling models it supports, so an agent knows exactly what the tool produces. It does not explicitly contrast itself with reminder_cancel or reminder_list, which keeps it at a 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives useful mode-selection guidance (one-time vs daily/weekly/every N days/every N minutes), which implies when each variant applies. However, it never states when not to use the tool, nor points to sibling tools like reminder_cancel or reminder_list, so usage guidance remains implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_filesARead-onlyIdempotentInspect
Search files and attachments across the workspace — by content, filename, document type, or origin. For message content use search.messages; for links use search.links.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum files to return. | |
| query | No | What to search for (content or filename). | |
| file_origin | No | File origin: 'generated' (created by tools), 'received' (from messages), 'uploaded' (manual). Use 'generated' for files the user created/sent. OMIT to include all origins. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| document_type | No | Filter by document category. OMIT unless the user explicitly mentions one — picking a value narrows the search and is a common cause of zero-result mistakes. | |
| attachment_name | No | Exact filename filter. OMIT to skip (do NOT pass an empty string). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered by structured data. The description adds that attachments are included and scoping is workspace-wide, but says nothing about result shape, pagination, or the effect of the session-workspace default used by in_workspace.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, zero filler, with the primary scope statement front-loaded ahead of the disambiguation clause. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter, zero-required search tool with a fully documented schema and complete safety annotations, the description supplies enough to select and invoke it correctly. The only gap is that with no output schema it never hints at what a result row contains, which would help an agent interpret the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds no parameter-specific behavior beyond the schema — it does not touch the limit ceiling, the OMIT-semantics for filters, or the caution around narrowing document_type that the schema itself documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and resource (files and attachments), scopes it ('across the workspace'), and enumerates the searchable axes (content, filename, document type, origin). It explicitly distinguishes itself from the nearest siblings by naming search.messages and search.links, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names two alternatives and the condition that selects each ('For message content use search.messages; for links use search.links'), which is strong routing. It stops short of addressing overlap with file-listing siblings (e.g. agents_list_files), so the 'when-not' for this tool itself is not fully closed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_linksARead-onlyIdempotentInspect
Search links/URLs shared across the workspace — by type, owner, or associated contact. For files use search.files; for message content use search.messages.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum links to return. | |
| owner | No | Link owner: 'self' (user's own) or 'contact' (from others). OMIT to include links regardless of owner. | |
| query | No | What to search for in shared links. | |
| link_kind | No | Filter links by type. OMIT to include all kinds — picking a value narrows the search and is a common cause of zero-result mistakes. | |
| contact_hint | No | Name hint to filter links for a specific contact. OMIT to skip. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds only the 'shared across the workspace' scope, and says nothing about rate limits, result volume beyond the schema's limit, or return format. A 3 is appropriate given annotations carry the behavioral load.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, zero waste. The primary capability is front-loaded and the disambiguation follows, so the most decision-relevant content comes first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with complete annotations and full schema coverage, the description covers what is searched and where the boundaries lie with siblings. It stops short of describing return shape or pagination, but with no output schema and a low-complexity read operation this is only a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters including enum meanings and OMIT behavior. The description's 'by type, owner, or associated contact' roughly maps to link_kind/owner/contact_hint but adds no syntax or format detail beyond the schema. Baseline 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Search links/URLs shared across the workspace') and immediately scopes it ('by type, owner, or associated contact'). It explicitly names the sibling tools to use for adjacent needs (files, message content), so an agent can route without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use (shared links/URLs) and when-not alternatives ('For files use search.files; for message content use search.messages'). This is a textbook alternative-routing statement, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_messagesARead-onlyIdempotentInspect
Search message content across all chats — semantic + keyword. Use to find what was said: quotes, topics, info exchanged. For chats/threads themselves use search.threads; for files use search.files; for links use search.links.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum messages to return. | |
| query | No | What to search for in message content. | |
| date_to | No | ISO8601 date (YYYY-MM-DD) upper bound. OMIT to skip. | |
| date_from | No | ISO8601 date (YYYY-MM-DD) lower bound. OMIT to skip. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| participant_name | No | Filter to messages involving this participant/contact name. OMIT to search across everyone (do NOT pass an empty string). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive, so the safety profile is covered. The description adds real behavioral context beyond that: it discloses that search is both semantic and keyword, and that it spans all chats. It omits ordering/pagination behavior, but that is a minor gap against the annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action and followed by the disambiguation. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with full annotation coverage and a fully documented schema, the description supplies enough to call it correctly. Return format is unspecified, but with no output schema that is a minor omission rather than a blocking gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters (query, limit, dates, workspace, participant) are already documented in the schema. The description adds no syntax or format detail beyond it, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (search message content) with scope (across all chats) and the search modality (semantic + keyword). It distinguishes itself from sibling searches by naming search.threads, search.files, and search.links explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use ('find what was said: quotes, topics, info exchanged') and explicit alternatives for adjacent needs (chats/threads, files, links). An agent can route correctly without opening any schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_threadsARead-onlyIdempotentInspect
Find or list chat threads/conversations — by topic, participant, unread/unanswered status, or recency. Omit query to list threads by filter. For message content use search.messages; for files use search.files. since filters by recency and pairs with only_unread / only_unanswered. Each thread names the LINE it speaks from (account: id, channel, name, phone) and whether the owner ever really wrote there (has_real_outgoing; imported history does not count). The same person can appear once per connected number — pick the thread by account, not by title, before you messages.send.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum threads to return. | |
| query | No | Topic/keyword to search threads for. OMIT to list threads by filter. | |
| since | No | ISO date (YYYY-MM-DD). Only threads with any message activity since this date (recency filter, not 'unanswered'). OMIT to skip. | |
| offset | No | Number of threads to skip (for pagination, default: 0) | |
| only_unread | No | Limit to threads with unread messages. OMIT to include read threads. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| only_unanswered | No | Limit to threads where the last message is incoming (you haven't replied). Covers 'threads I haven't replied to'. OMIT to include answered threads too. | |
| participant_name | No | Filter to threads with this participant/contact. OMIT to include everyone (do NOT pass an empty string). | |
| include_last_messages | No | With only_unanswered: inline the most recent N messages (0-2) per thread so one call is enough to draft replies. Text truncated to 150 chars; full body via messages.read_history. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, but the description adds real behavioral context: the `since`+filter pairing, that each thread carries an `account` (id, channel, name, phone), the meaning of `has_real_outgoing` (imported history excluded), and the warning that one person recurs once per number so selection must be by account.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence, then routing, then semantics, then a final selection warning. Dense but every clause carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-param read-only list tool with no output schema, the description supplies both the routing guidance and the return-shape essentials (account fields, has_real_outgoing). Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter meaning the schema lacks: `since` pairs with only_unread/only_unanswered and is a recency filter, not 'unanswered', and omitting `query` switches from search to list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Find or list chat threads/conversations') and enumerates the axes it searches by (topic, participant, unread/unanswered, recency). It cleanly disambiguates from siblings by naming search.messages, search.files, and messages.send as the tools for adjacent tasks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing: use search.messages for content, search.files for files, and omit `query` to list rather than search. It also sequences the tool before messages.send, telling the agent when to reach for it in a workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
system_sleepARead-onlyIdempotentInspect
Pause execution for a given number of seconds (max 30). Use when you need to wait for an external process to complete before retrying — e.g. message sync, backfill, or API propagation. Total sleep per run is capped at 60 seconds.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | Why you are waiting (logged for debugging) | |
| seconds | Yes | Number of seconds to sleep (1-30) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it a safe, idempotent, non-destructive read, so the safety profile is covered. The description adds useful behavior the annotations do not: the per-call cap of 30 seconds and the aggregate 60-second cap per run, which directly affects retry-loop design.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: the action and limit first, the selection condition second, the cumulative constraint third. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter tool with full schema coverage, no output schema, and annotations carrying the safety profile, the description covers everything an agent needs: what it does, hard limits, and when to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and every parameter (seconds, reason, in_workspace) is documented in the schema with its own description and range. The description restates the 30-second maximum already stated in the schema, adding no new semantics, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Pause execution for a given number of seconds (max 30).' The scope and limit are immediately clear, though it never distinguishes itself from sibling wait-style tools such as browser_wait_for or calls_wait, which an agent might otherwise consider.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger condition: 'Use when you need to wait for an external process to complete before retrying,' with concrete examples (message sync, backfill, API propagation). It lacks an explicit 'do not use this for X, use Y instead' exclusion, but the context is clear enough to select correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_commentAInspect
Post a comment on a task. Attach files (screenshots, documents) by uploading them with files.upload first and passing the file IDs.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | Comment text | |
| task_id | Yes | Task to comment on | |
| agent_id | No | Comment as this agent instead of as yourself. Optional. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| attachment_file_ids | No | File IDs to attach (from files.upload) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the write profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the agent knows this mutates state and is not idempotent. The description adds the cross-tool upload dependency, which is genuinely useful, but says nothing about notification side effects, permissions, or what a repeat call produces beyond what idempotentHint implies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both load-bearing: the first declares the action, the second handles the one non-obvious parameter. Purpose is front-loaded with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the agent is not told what a successful call returns (e.g. a comment ID) or how errors surface. For a write tool with annotation coverage and a fully described schema, the description is serviceable but leaves the response contract unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are documented in the schema itself. The description's note about attachment_file_ids largely restates the schema's '(from files.upload)' hint and says nothing about agent_id or in_workspace, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Post a comment on a task'), which is unambiguous. No sibling tool also posts comments on tasks, so disambiguation is implicit rather than stated, but the action is immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete prerequisite workflow for the optional attachment path: upload via files.upload first, then pass the returned IDs. It does not state when to prefer this over other task tools, but there is no competing sibling, so the guidance is adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_createInspect
Create a task in this workspace. Leave the assignee empty to put it in the backlog, or name someone to hand it over.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Task title | |
| due_at | No | ISO datetime when the task is due, e.g. '2026-03-31T15:00:00Z'. A time without a zone is read as UTC. | |
| tag_ids | No | Topics for the task: ids of workspace tags that apply to tasks (ai_tags.list with applies_to='task'). Replaces the current set; [] clears it. | |
| agent_id | No | Act as this agent instead of as yourself — the task is recorded as authored by it. Optional. | |
| assignee | No | Who should do this: a workspace member (email, username, or user:<id> as workspace.members lists them), an AI agent (its name), or a contact (display name). Use 'me' for yourself. Call workspace.members to see who can be assigned. Where this listing offers assignee_user_id / assignee_agent_id / assigned_to_contact_id, they name an exact target instead. | |
| due_date | No | Date when task is due (e.g. '2026-03-31'). Use with due_time or alone. | |
| due_time | No | Time when task is due (e.g. '15:00'). Used with due_date. | |
| priority | No | Task priority (default: medium) | |
| thread_id | No | Related thread ID | |
| description | No | Detailed description | |
| repeat_unit | No | Unit for repeat_every. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| repeat_every | No | Make the task repeat: completing it creates the next occurrence, due this many repeat_unit after the previous deadline. E.g. repeat_every=3, repeat_unit='month' for quarterly. | |
| target_session | No | Which desktop Claude session should run this, by the name shown in workspace.desktops. Only needed when the assignee is an AI agent that runs on a machine and more than one session is connected — otherwise the task is not dispatched at all, rather than sent to a machine you did not mean. | |
| assignee_user_id | No | Exact workspace member to assign to | |
| assignee_agent_id | No | Exact AI agent to assign to | |
| attachment_file_ids | No | File IDs to attach (from files.upload) | |
| assigned_to_contact_id | No | Exact contact to assign to |
tasks_deleteADestructiveIdempotentInspect
Delete a task. Only its creator or a workspace owner/admin may do this; to close a task instead, set status='cancelled'.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | ID of the task to delete | |
| agent_id | No | Act as this agent instead of as yourself. Optional. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the danger profile is covered. The description adds genuinely new behavioral context above that bar: who is authorized to delete. It stops short of describing side effects on related data (e.g., comments/subtasks) or the response, but for a covered mutation this is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the destructive action is front-loaded before the permission rule and the alternative. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param destructive delete with full schema coverage and no output schema, the definition covers purpose, permission, and the main alternative. It could add one clause on whether deletion is reversible or what happens to dependent records, but nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for all three parameters (task_id, agent_id, in_workspace), so the schema already carries the parameter burden and the baseline is 3. The only parameter-adjacent detail in the description (status='cancelled') actually belongs to a different tool, so it adds no semantics for this call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Delete a task") and immediately distinguishes itself from the sibling update path by noting that closing a task is done via status='cancelled' rather than deletion. An agent can tell this apart from tasks_update/tasks_get without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit when-not with a concrete alternative: if the intent is to close rather than remove the task, set status='cancelled'. It also states the authorization condition (creator or workspace owner/admin), so the agent knows whether the call is even permitted.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_getARead-onlyIdempotentInspect
Read one task in full: its fields, attachments, and the whole comment thread including recorded status and assignee changes. Pass agent_id when you are acting as an agent, and the reply says assigned_to_me so you need not infer it from matching ids.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | ID of the task to read | |
| agent_id | No | Read as this agent. Set it to your own agent_id to be told whether the task is assigned to you. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds real value beyond that by disclosing the return shape (whole comment thread with recorded status/assignee changes) and the assigned_to_me reply behavior, which the agent could not otherwise predict.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler: the read scope is front-loaded and the agent_id behavior follows. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with no output schema, the description covers the return contents (fields, attachments, comment thread) and the agent_id reply semantics, while annotations handle the safety profile and the schema handles all parameters. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents task_id, agent_id and in_workspace. The description largely restates the agent_id semantics ('the reply says assigned_to_me') rather than adding syntax or constraints beyond the schema text, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read one task in full') and enumerates what the read returns (fields, attachments, comment thread), which separates it from tasks_list and tasks_comment. It stops short of naming any sibling explicitly, so differentiation is implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only conditional guidance is about the agent_id parameter ('Pass agent_id when you are acting as an agent'), which is parameter-level rather than tool-selection guidance. There is no stated when-to-use vs. tasks_list/tasks_comment or any exclusion, leaving usage implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_listARead-onlyIdempotentInspect
List tasks in this workspace. Defaults to everything a person can see; an agent calling this sees its own tasks. Filter by assignee, status, or overdue to narrow it down.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results per page (default 20, max 100) | |
| offset | No | Skip this many results — use with limit to page through | |
| status | No | ||
| tag_id | No | Only tasks with this topic | |
| overdue | No | Only tasks past their due date that are not finished | |
| agent_id | No | Show one agent's own tasks instead of the whole workspace. Optional. | |
| assignee | No | Only tasks assigned to this member / agent / contact. Use 'me' for your own. | |
| thread_id | No | Filter by related thread | |
| unassigned | No | Only tasks with nobody assigned (the backlog) | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| created_by_me | No | Only tasks you created | |
| assignee_user_id | No | ||
| assignee_agent_id | No | ||
| assigned_to_contact_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower. The description adds real value beyond that by disclosing the visibility scoping rule — a person sees everything, an agent sees only its own tasks — which is non-obvious behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose, then scope behavior, then filter hints. No filler and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 14 parameters, no output schema, and rich annotations, the description gives the essential behavioral nuance (default visibility scoping) that the schema cannot convey. It falls short only on distinguishing the many overlapping identity filters, but is otherwise adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 71%, so the schema documents most parameters itself. The description names only three of the fourteen (assignee, status, overdue) as filters, adding marginal meaning but not clarifying the agent_id/assignee/created_by_me distinctions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("List tasks in this workspace") and clarifies scope with the default-visibility note. It is clear what the tool does, but it does not differentiate from siblings such as tasks_get or tasks_create, so an agent still infers the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the default behavior (everything a person can see; an agent sees its own tasks) and mentions filters to narrow down, which implies when to use it. However, no alternatives are named and there is no explicit when-not guidance versus tasks_get or other list tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_updateInspect
Update a task. Set status='done' to complete it, 'cancelled' to cancel. Reassign with assignee. Status and assignee changes are recorded in the task's comment thread.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | ||
| due_at | No | ISO datetime | |
| status | No | ||
| summary | No | Completion note (stored when marking done) | |
| tag_ids | No | Topics for the task: ids of workspace tags that apply to tasks (ai_tags.list with applies_to='task'). Replaces the current set; [] clears it. | |
| task_id | Yes | ID of the task to update | |
| agent_id | No | Act as this agent instead of as yourself. Optional. | |
| assignee | No | Who should do this: a workspace member (email, username, or user:<id> as workspace.members lists them), an AI agent (its name), or a contact (display name). Use 'me' for yourself. Call workspace.members to see who can be assigned. Where this listing offers assignee_user_id / assignee_agent_id / assigned_to_contact_id, they name an exact target instead. | |
| priority | No | ||
| description | No | ||
| repeat_unit | No | Unit for repeat_every. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| repeat_every | No | Make the task repeat: completing it creates the next occurrence, due this many repeat_unit after the previous deadline. E.g. repeat_every=3, repeat_unit='month' for quarterly. Pass repeat_every=0 to stop repeating. | |
| clear_assignee | No | Unassign the task, returning it to the backlog | |
| target_session | No | Which desktop Claude session should run this, by the name shown in workspace.desktops. Only needed when the assignee is an AI agent that runs on a machine and more than one session is connected — otherwise the task is not dispatched at all, rather than sent to a machine you did not mean. | |
| assignee_user_id | No | ||
| assignee_agent_id | No | ||
| attachment_file_ids | No | Replace the task's attachments with these file IDs | |
| assigned_to_contact_id | No |
threads_channel_account_statsBRead-onlyIdempotentInspect
Get the connected Meta Threads account's profile (username, name, avatar) and account-level insights (followers, views, likes, replies, reposts, quotes).
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds the useful context of exactly which metrics are returned, but says nothing about auth requirements, rate limits, or the absence of historical/date scoping.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence, front-loaded with the verb and resource, with the returned field list serving as immediate clarification. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the right thing by enumerating the returned profile fields and metrics. Combined with annotations covering safety and a fully documented single parameter, it is nearly complete; the only gap is any sense of scope/time-window for the insights.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single in_workspace parameter is fully documented in the schema, so the description carries no additional parameter burden. Baseline 3 applies since the schema does the heavy lifting and the description never references the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and a precise resource (the connected Meta Threads account's profile and account-level insights), then enumerates the exact fields returned. The 'account-level' scope implicitly distinguishes it from post-level siblings like threads_channel_post_insights, but no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, and no alternatives are named. An agent must infer from the 'account-level' phrasing that this differs from threads_channel_post_insights or threads_channel_list_posts rather than being told.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
threads_channel_hide_replyAInspect
Hide a reply on a Meta Threads post owned by the connected account. Also hides every nested reply beneath it.
| Name | Required | Description | Default |
|---|---|---|---|
| reply_id | Yes | The Threads reply id to hide. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false), and the description adds a genuine behavioral fact they don't convey: the hide cascades to every nested reply. It stops short of explaining the declared idempotentHint=false (what happens if the reply is already hidden) or reversibility via the unhide sibling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, and the cascade behavior is front-loaded right after the core action. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter mutation with full schema coverage and no output schema, the description covers the essential action and its side effect. It omits only the reversal path and the meaning of the declared non-idempotency, which are minor rather than blocking gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both reply_id and the session-vs-workspace semantics of in_workspace are already fully documented in the schema. The description adds no parameter-level meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (hide) and resource (a reply on a Meta Threads post), plus the ownership scope ('owned by the connected account'). An agent can distinguish it from the inverse sibling threads_channel_unhide_reply without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the action itself, but there is no explicit when-to-use statement, no pointer to threads_channel_unhide_reply as the reversal, and no prerequisite/eligibility conditions (e.g., what qualifies as an owned post). Adequate but with clear gaps.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
threads_channel_list_postsARead-onlyIdempotentInspect
List posts published by the connected Meta Threads account. Returns id, text, permalink, media_type and timestamp. Paginate with the returned after cursor.
| Name | Required | Description | Default |
|---|---|---|---|
| after | No | Pagination cursor from a previous call's `after`. | |
| limit | No | Page size, 1-100. Default 25. | |
| since | No | ISO date (YYYY-MM-DD) lower bound. | |
| until | No | ISO date (YYYY-MM-DD) upper bound. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful behavior beyond that: the exact returned fields (id, text, permalink, media_type, timestamp) and the cursor-based pagination pattern, which matters given there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: purpose, return shape, pagination. Front-loaded with the core action and free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, listing the return fields is a meaningful addition, and pagination behavior is covered. Minor gaps remain (e.g., default page size interaction, rate-limit or auth context), but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (after, limit, since, until, in_workspace) are already documented. The description reinforces pagination via the 'after' cursor but adds no syntax or format meaning beyond the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List posts published by the connected Meta Threads account'), which is unambiguous. It does not, however, differentiate itself from siblings such as threads_channel_post_insights, threads_channel_account_stats, or instagram_list_media, so an agent gets a clear purpose but no routing signal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or exclusions are given. The description never says when to prefer this over threads_channel_post_insights (per-post metrics) or the other posting siblings, leaving the agent to infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
threads_channel_post_insightsBRead-onlyIdempotentInspect
Get engagement metrics for one Meta Threads post: views, likes, replies, reposts and quotes.
| Name | Required | Description | Default |
|---|---|---|---|
| post_id | Yes | The Threads post id (from threads_channel.list_posts). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive, and closed-world, so the safety profile is fully covered. The description adds the list of metric fields returned, which is genuinely useful given there is no output schema, but it says nothing about auth requirements, freshness/delay of metrics, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler: verb, resource, scope, and returned fields all in one line. Nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only tool with full schema coverage, the definition covers purpose and return fields, and annotations carry the safety profile; with no output schema, naming the metrics is the key missing piece and it is present. It stops short of 5 only because no usage context or metric-freshness caveat is offered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both post_id and in_workspace are already documented in the schema (post_id even cites threads_channel.list_posts as the source). The description adds no syntax or format detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('engagement metrics for one Meta Threads post') and enumerates the exact metrics returned (views, likes, replies, reposts, quotes). The 'one post' scoping implicitly separates it from the account-level sibling threads_channel_account_stats, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description is purely a purpose statement with no when-to-use, when-not-to-use, or alternative routing. An agent must infer on its own that this is for per-post metrics versus threads_channel_account_stats or threads_channel_list_posts, and nothing is said about prerequisites such as owning or having access to the post.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
threads_channel_publish_postADestructiveInspect
Publish a post to the connected Meta Threads account. Supports TEXT, IMAGE, VIDEO and CAROUSEL (2-20 items). Text is limited to 500 characters and 5 links. Media must be given as a PUBLIC URL — Threads has no file upload endpoint. Returns the published post id and permalink.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | Post text. Required for TEXT, optional caption otherwise. Max 500 characters and 5 links. | |
| alt_text | No | Accessibility text for a single IMAGE or VIDEO post. | |
| image_url | No | Public URL of the image (JPEG/PNG, max 8MB). IMAGE only. | |
| video_url | No | Public URL of the video (MOV/MP4, max 1GB / 5 min). VIDEO only. | |
| media_type | Yes | TEXT, IMAGE, VIDEO or CAROUSEL. | |
| reply_to_id | No | Publish as a reply to this Threads post/reply id. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| reply_control | No | Who may reply to this post. OMIT to leave it at the account's default (everyone). | |
| carousel_items | No | CAROUSEL only: 2-20 items, each {media_type: IMAGE|VIDEO, image_url|video_url, alt_text?}. Order is the render order. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false and openWorldHint=true, so the mutation/safety profile is covered. The description adds real behavioral context beyond that: media must be PUBLIC URLs because Threads has no upload endpoint, plus text/char/limit constraints and the post-id/permalink return.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, zero filler, and the core capability plus key constraint (public URL only) are front-loaded before the return-value note. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter publish tool with no output schema, the description covers media modes, key constraints, the URL requirement, and even the return shape (post id and permalink). Only the absence of routing guidance against siblings keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all nine parameters including the enum media_type and reply_control. The description's constraint mentions (500 chars, 5 links, carousel 2-20) largely restate what the schema already says, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('Publish a post to the connected Meta Threads account') and enumerates supported content modes (TEXT, IMAGE, VIDEO, CAROUSEL). An agent can immediately distinguish it from sibling mutation tools like threads_update and threads_delete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives useful operational context (public URLs required, no file-upload endpoint) but never states when to choose this tool over alternatives such as threads_update or instagram_publish_media. Usage is implied by the resource, not routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
threads_channel_unhide_replyAInspect
Unhide a previously hidden reply on a Meta Threads post owned by the connected account. Also unhides its nested replies.
| Name | Required | Description | Default |
|---|---|---|---|
| reply_id | Yes | The Threads reply id to unhide. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations set readOnlyHint=false, destructiveHint=false and idempotentHint=false, so the mutation profile is already declared. The description adds genuinely new behavioral context not in the annotations: the operation cascades to nested replies, and it is bounded to content on a post owned by the connected account. It does not detail auth requirements or failure behavior, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences, with the core action front-loaded and the cascade detail following. No filler or redundancy; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and annotations already covering the safety profile, the description says what the tool does, what it affects, and the cascade side-effect, which is sufficient for correct invocation. It is slightly short on permission/ownership failure semantics, but nothing essential is missing for a simple toggle mutation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both reply_id and the in_workspace parameter are already documented in the schema. The description adds no syntax, format, or semantic detail beyond that, which is the expected baseline when the schema carries the full burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (unhide) and resource (reply), scopes it to a Meta Threads post owned by the connected account, and adds the cascading detail that nested replies are unhidden too. An agent can immediately distinguish this from its inverse sibling threads_channel_hide_reply without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'a previously hidden reply' implies the use case (reversing a prior hide), so usage is inferable. However, it never names the counterpart tool or states any condition for when this should or should not be invoked, leaving the when-to-use reasoning implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
threads_deleteADestructiveIdempotentInspect
🗑️ PERMANENTLY delete conversation thread(s) and their messages + user-facing state (tags, assignments, drafts, reminders, RAG chunks). Destructive and NOT undoable — requires confirm=true.
Pass thread_id for one, or thread_ids for a bulk delete. Kept intentionally: call history (voice_sessions), usage/billing, traces, and outreach dedup are NOT removed. Note: for a live synced channel (Telegram/WhatsApp) this clears the LOCAL copy; a new inbound can re-create the thread on next sync. For livechat/voice/test/duel threads it's effectively permanent.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | Yes | Must be true to actually delete — a safety gate for this destructive action. | |
| thread_id | No | A single thread ID to delete. Use this OR thread_ids. | |
| thread_ids | No | Multiple thread IDs to delete in one call (bulk). Use this OR thread_id. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, but the description goes well beyond them: it lists exactly what is destroyed, explicitly enumerates what is preserved (call history, billing, traces, outreach dedup), states the operation is NOT undoable, and discloses the live-sync caveat where a thread can be re-created.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the irreversible-destruction warning and the confirm requirement, then layers what is deleted, what is kept, and the channel nuance. Despite its length, nearly every sentence carries decision-relevant information and none is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, multi-target tool with no output schema, the description covers the confirm gate, the single/bulk parameter choice, the full set of side effects, preserved data, and the channel-dependent permanence distinction. Nothing an agent needs to invoke it safely is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so confirm, thread_id, thread_ids and in_workspace are already documented in the schema, and the description largely restates the OR-relationship and the confirm gate. With the schema doing the heavy lifting, a baseline 3 is appropriate; the description offers no additional syntax or format detail beyond what is already structured.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (permanently delete conversation thread(s)) plus the exact scope of collateral deletion and the required confirm gate. It is clearly distinguishable from siblings like threads_update or messages_delete because it enumerates what is removed (messages, tags, assignments, drafts, reminders, RAG chunks).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use guidance via the thread_id (single) vs thread_ids (bulk) selection and the confirm=true prerequisite. It also explains channel-dependent behavior (live synced channels only clear the local copy; livechat/voice/test/duel are permanent), which is strong context, though it never names an explicit alternative tool for narrower deletions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
threads_updateAInspect
✏️ Update a conversation thread: rename it, add notes/description, move it to a folder, or take it out of sync.
When to use:
User wants to rename a chat or group
User wants to add notes/context about a conversation
User wants to organize threads into folders
User wants a chat to stop showing up and stop syncing (
synced=false) — e.g. the personal chats on an account that is only in the workspace for its work groups. Nothing is deleted;synced=truebrings it back.
For DM threads, renaming also updates the linked contact's display name by default. Pass thread_id for one thread, or thread_ids to apply the same change to several. Ids come from threads.list / search.threads.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | New title for the thread (max 255 chars). For a DM this ALSO renames the linked contact, so an empty string is ignored, not applied — omit the parameter when you are not renaming. | |
| synced | No | false takes the chat out of sync: it disappears from the inbox, from knowledge/RAG search, and the sync service stops fetching it. Messages already stored are kept and nothing is deleted. true puts it back. The decision survives later syncs. OMIT to leave sync as it is. | |
| folder_id | No | Move thread to this folder (null removes from folder) | |
| thread_id | No | Thread ID from threads.list. Use this OR thread_ids. | |
| thread_ids | No | Several thread IDs to apply the same change to, e.g. taking a batch of personal chats out of sync in one call. Use this OR thread_id. | |
| description | No | AI context / notes for this thread. Empty string clears description. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| update_contact | No | For DM threads, also rename the linked contact (default: true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false; the description confirms and enriches this with 'Nothing is deleted; synced=true brings it back' and that the sync decision persists across later syncs. It also discloses the DM rename side-effect and the ignored-empty-string rule. Idempotency behavior (idempotentHint=false) is left unaddressed, so it falls just short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line summary followed by clean, scannable bullets; each sentence largely earns its place. There is mild redundancy with the schema (DM rename and synced behavior are restated), which keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with 8 optional parameters and no output schema, the description covers the key operations, side effects, and persistence rules. It omits any mention of return values or error behavior and does not address the in_workspace override, leaving minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description still adds value by telling the agent where ids come from (threads.list / search.threads) and reinforcing the thread_id/thread_ids mutual-exclusion semantics. It does not add much beyond the richly documented schema for the remaining parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Update a conversation thread') and enumerates the exact operations (rename, add notes/description, move to folder, take out of sync). An agent can immediately tell this is the mutation counterpart to threads_list/search_threads and distinct from threads_delete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit 'When to use' list with concrete scenarios, including the non-obvious case of taking personal chats out of sync on a work-only account. It also explains the thread_id vs thread_ids choice and clarifies that nothing is deleted, effectively steering away from delete semantics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tiktok_account_statsARead-onlyIdempotentInspect
Get the connected TikTok account's profile (username, display name, avatar) and stats (followers, following, likes, videos).
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety and determinism profile is covered. The description adds that it reads the 'connected' account (single-account scope) and lists returned fields, but says nothing about auth requirements, rate limits, or what happens when no account is connected. With annotations carrying the safety burden, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, well-formed sentence with zero waste, front-loading the verb and resource and parenthetically enumerating the returned fields. Nothing needs to be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by listing the exact fields returned, which is the key information an agent needs. It omits edge cases (no connected account, error behavior), which keeps it just short of fully complete for a stats tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter (in_workspace) is fully documented in the schema, so baseline 3 applies. The description adds nothing about the parameter, which is acceptable given the schema already explains its semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (connected TikTok account's profile and stats) and enumerates exactly what comes back (username, display name, avatar, followers, following, likes, videos). This clearly distinguishes it from the siblings tiktok_list_videos and tiktok_publish_video, which operate on videos rather than the account itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the name and description (fetch account-level metrics), but there is no explicit when-to-use, when-not-to-use, or routing to alternatives. An agent can infer the purpose but gets no guidance on conditions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tiktok_list_videosARead-onlyIdempotentInspect
List videos on the connected TikTok account. Returns id, title, share_url, view/like/comment/share counts. Paginate via cursor.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | Pagination cursor from a previous call's `cursor`. Omit for the first page. | |
| max_count | No | Page size, 1-20. Default 20. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so safety behavior is fully covered structurally. The description adds only that it returns specific fields and paginates via cursor, which is modest additional context rather than rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences: purpose, return shape, and pagination. No filler, and the primary purpose leads.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully compensates by enumerating the returned fields (id, title, share_url, counts) and explaining pagination, which is enough for correct invocation. Only minor gaps remain, such as ordering or rate/scope limits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (cursor, max_count, in_workspace) are already documented in the schema. The description's only parameter-related content is 'Paginate via cursor,' which adds nothing beyond the schema's cursor field description. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (videos) with the scope narrowed to 'the connected TikTok account' and the platform named, which cleanly separates it from youtube_list_videos, instagram_list_media, and tiktok_account_stats. An agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its listing role and gives one usage hint ('Paginate via cursor'), but never states when to prefer it over siblings or any precondition (e.g., connected account state). Usage is inferable rather than explicitly guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tiktok_publish_videoADestructiveInspect
Publish a workspace-owned video file (file_id) to the connected TikTok account via the Content Posting API. Returns publish_id + final status. Note: until the TikTok app audit passes, TikTok forces SELF_ONLY (private) visibility.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Post caption/title (max 2200 chars; hashtags allowed). | |
| file_id | Yes | Workspace `files.id` of the video to publish. Must be a video/* MIME type and status='ready'. | |
| disable_duet | No | Disable duets. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| privacy_level | No | Post visibility. Must be one the creator allows (checked against creator_info). Default SELF_ONLY — the only level TikTok accepts from unaudited apps. | SELF_ONLY |
| disable_stitch | No | Disable stitches. | |
| disable_comment | No | Disable comments on the post. | |
| channel_account_id | No | Connected TikTok channel_account.id. OMIT to auto-resolve the workspace's account. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true, idempotent=false, and openWorld=true, so the write/side-effect profile is covered. The description goes beyond that by disclosing the return payload (publish_id + final status) and a real behavioral constraint: TikTok forces SELF_ONLY until the app audit passes, which directly affects the effective visibility outcome.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences with zero filler: purpose first, return contract second, and the surprising SELF_ONLY constraint last. Every sentence earns its place and nothing is redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description supplies the return contract (publish_id + final status) and the key constraint on visibility, while annotations handle the safety profile. It is nearly complete; minor gaps remain around error/failure behavior, but nothing essential to calling the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all eight parameters, including the enum values for privacy_level and the auto-resolve behavior of channel_account_id. The description reinforces that file_id is a workspace-owned video but adds no syntax or format detail beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (publish), a specific resource (a workspace-owned video file via file_id), and a specific destination (the connected TikTok account via the Content Posting API). This clearly separates it from siblings like tiktok_list_videos, tiktok_account_stats, and the other publish tools (instagram_publish_media, youtube_upload_video) without requiring the schema to be opened.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case (publishing a video to TikTok) and notes the preconditions on file_id, but never explicitly states when to use this versus an alternative publish tool or when not to use it. There are no exclusions or routing guidance, leaving the agent to infer context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
videos_generateAInspect
Generate a short video (5-10s) from a text prompt using BytePlus Seedance. Optionally accepts up to 12 image file IDs from the user's attached files (visible in the [ATTACHMENTS] block) as reference_file_ids for style and composition. Returns immediately with a job_id; the video is delivered back via continuation when the job completes (~30-90s for fast model, ~2-5min for pro). Reference images are temporarily re-hosted on a third-party CDN (imgbb) for the duration of generation and deleted on completion — don't submit confidential references. Gated behind a workspace opt-in flag.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Random seed for reproducibility (0-2147483647). Omit for random. | |
| model | No | Video model. 'wan2.6-i2v-flash' (default, cheap, 720p/1080p, optional audio), 'wan2.6-i2v' (premium, always-on audio), 'wan2.6-t2v' (text-only input, 720p/1080p, no audio), 'wan2.2-i2v-flash' (cheapest, 480p/720p, no audio). Legacy BytePlus: 'seedance-2-fast', 'seedance-2-pro' (720p only). | wan2.6-i2v-flash |
| style | No | Style preset. Seedance models only. OMIT for no style preset. | |
| prompt | Yes | Text description of the video to generate (3-4000 chars). | |
| duration | No | Output video duration in seconds. Single-clip: 5 or 10. Long-form (chained, i2v models only): 15, 20, 30, 45, or 60. Long-form videos are silent (no audio in v1) and use only reference_file_ids[0] when refs are provided. | |
| shot_type | No | Shot mode: 'single' (continuous) or 'multi' (scene cuts). wan2.6-t2v only. OMIT to use the model default. | |
| resolution | No | Output resolution. '720p' is the safe default; '1080p' is wan2.6 only; '480p' is wan2.2-i2v-flash only. Per-model support enforced by validation. | 720p |
| aspect_ratio | No | Output aspect ratio. Wan supports '16:9', '9:16', '1:1'; Seedance also supports '4:3', '3:4', '21:9'. Per-model support enforced by validation. | 16:9 |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| camera_motion | No | Camera motion preset. Seedance models only. OMIT for no camera motion. | |
| generate_audio | No | Whether the model should produce native audio. For wan2.6-i2v-flash this doubles the per-second rate (e.g., 720p+audio is $0.05/s vs $0.025/s silent) — set False for cheaper silent clips. wan2.6-i2v always produces audio regardless of this flag. wan2.6-t2v / wan2.2-i2v-flash / seedance-2-fast never produce audio. | |
| negative_prompt | No | Optional text describing what to AVOID in the output. Honored by Wan and Seedance models. | |
| reference_file_ids | No | Optional list of up to 12 image file_ids to use as visual references (style, composition). Files must be image MIME types (image/png, image/jpeg, image/webp, image/gif). Get IDs from the [ATTACHMENTS] block, files.search, or search.files. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: it reveals the async contract (returns immediately with a job_id, delivered via continuation), expected timing for fast vs pro models (~30-90s vs ~2-5min), and a genuine privacy caveat that reference images are re-hosted on a third-party CDN (imgbb) and deleted on completion. The confidentiality warning and gating flag are operational details an agent could not infer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then progressively discloses optional inputs, the async return, and the privacy caveat. It is one dense paragraph but nearly every sentence carries decision-relevant information; only minor trimming of the timing detail would tighten it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an async generation tool with no output schema, the description supplies everything needed to call it correctly: return contract (job_id + continuation), latency expectations, the reference-image constraints, the CDN privacy caveat, and the workspace opt-in requirement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: the [ATTACHMENTS] block as the source of reference_file_ids, the number limit (12), and the confidentiality implication of submission. Other parameter nuance (models, duration, resolution) is fully documented in the schema itself.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Generate) and resource (short video 5-10s) with the upstream provider named (BytePlus Seedance) and the input modality (text prompt). An agent immediately understands what this produces, and no sibling tool competes for this purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains when reference images apply ('optionally accepts up to 12 image file IDs... for style and composition') and points to the [ATTACHMENTS] block as the source. It also discloses the workspace opt-in gating, which affects whether the call succeeds. It stops short of naming explicit when-not conditions or a close alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vision_queryARead-onlyIdempotentInspect
Look at the screen currently being shared in a meeting and answer a question about it. Returns a natural-language answer based on the visual content. Use ONLY when the user explicitly asks about the screen/slide/document being shown.
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | Question about the shared screen. | |
| image_b64 | No | Base64-encoded JPEG image of the screen-share frame. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the read-only, non-destructive, idempotent safety profile, so the description's job is supplemental. It adds two useful behavioral facts: it operates on the *live* shared frame rather than arbitrary images, and it returns a natural-language answer (relevant because there is no output schema). It does not mention latency, cost, or what happens if no share is active, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: purpose, return value, then the usage restriction. The most decision-relevant constraint is front-loaded alongside the purpose rather than buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly explains the return value, and the annotations cover safety, so the tool is callable as described. The only gap is the slight tension between 'look at the screen currently being shared' and the optional image_b64 frame parameter, which the description never reconciles for the caller.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters and the baseline of 3 applies. The description only restates the question semantics already present in the schema and adds no format, syntax, or constraint detail for question, image_b64, or in_workspace.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Look at') and resource ('the screen currently being shared in a meeting') plus the action taken ('answer a question about it'). This clearly separates it from the screenshot/vision siblings (android_screenshot, browser_take_screenshot) which capture rather than interpret a live screen-share. An agent can identify it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit positive trigger ('explicitly asks about the screen/slide/document being shown') and an explicit exclusion ('ONLY when'), which is the strongest form of routing guidance in the corpus. Nothing about when to prefer this over a plain screenshot or web_fetch is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_fetchAInspect
Fetches a single URL and returns its content. Use this when you have a specific URL in mind — for example, after web.search returns a link you want to read, or when the user pastes a URL.
Modes (extract):
'auto' (default): picks the right mode based on response content type.
'markdown': for HTML pages; returns cleaned markdown plus the page .
'text': for JSON/XML/plaintext APIs; returns the raw decoded body.
'file': for images, PDFs, audio, video, archives, or any binary — ingests the bytes into the user's file storage and returns a file_id you can pass to messages.send (to send as an attachment), agents.add_file (to add to agent knowledge), or files.read.
Use web.fetch (not files.upload) when you need the file_id immediately for the next tool call — files.upload(source_url=…) is async and won't have the file ready in the same turn.
Use web.search (not web.fetch) when you don't have a specific URL yet and need to find one.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to fetch (http or https). Must be publicly reachable. | |
| extract | No | How to handle the response: 'auto' (default), 'markdown' (HTML → markdown), 'text' (raw body), or 'file' (ingest as binary, return file_id). | auto |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false/openWorldHint=false/idempotentHint=false but add little; the description carries the real burden, explaining that 'file' mode ingests bytes into the user's file storage and returns a file_id suitable for messages.send, agents.add_file, or files.read. It also discloses the async behavior of files.upload, a side-effect/comparison detail not captured by any structured field.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded verb+resource sentence, then a tight bulleted mode list, then two routing rules. Efficient overall, though the mode semantics partially restate the enum descriptions already in the input schema, which is mild redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, so the description appropriately describes return content per mode (cleaned markdown plus title, raw decoded body, file_id) and names the downstream tools that consume the file_id. For a 3-parameter fetch tool this is complete enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three parameters are already documented, but the description adds meaning beyond the schema: which mode fits which content type (markdown for HTML, text for JSON/XML APIs, file for binaries) and that markdown mode also returns the page <title>. in_workspace gets no narrative treatment, so it is not a full 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Fetches a single URL and returns its content') and immediately disambiguates from siblings web.search and files.upload. An agent can tell exactly what this does and what it is not without inspecting the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('after web.search returns a link', 'when the user pastes a URL') plus two named alternatives with the selecting condition: use web.search when you lack a URL, use files.upload when the file_id isn't needed this turn. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
webhooks_configure_enrichmentAInspect
Enable enrichment on an agent's webhook trigger so reference-only webhook payloads are resolved to full data before the agent runs. One tool for every provider — choose provider and pass its config.
Providers & required config:
meta_leadgen (Meta/Facebook Lead Ads): config = {app_id, app_secret, user_access_token}. The user token needs pages_show_list + leads_retrieval (+ pages_manage_metadata; add business_management to auto-discover Business-Manager pages that /me/accounts omits). Stores all page tokens encrypted per workspace and resolves each lead's name/email/phone. Reuses the existing webhook URL (no token rotation); creates a webhook trigger if the agent has none.
tiktok_leadgen (TikTok Instant Form and direct-message leads): config = {access_token, advertiser_ids?}. access_token is a TikTok Marketing API advertiser token issued to DialogBrain's Business app; without advertiser_ids every ad account the token reaches is subscribed. Adds a separate webhook trigger on the agent (a Meta-wired trigger is left alone) and subscribes it to LEAD events; each lead arrives with its answers, signature-verified.
| Name | Required | Description | Default |
|---|---|---|---|
| config | Yes | Provider-specific configuration/credentials. For meta_leadgen: {app_id, app_secret, user_access_token}. For tiktok_leadgen: {access_token, advertiser_ids?}. | |
| agent_id | Yes | ID of the agent whose webhook trigger should be enriched. | |
| provider | Yes | Enrichment provider to configure. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavioral detail beyond the annotations: page tokens stored encrypted per workspace, no token rotation, trigger creation when absent, a separate trigger for TikTok that leaves a Meta-wired trigger alone, LEAD-event subscriptions, and signature verification. This is exactly the side-effect and credential context the flat annotation hints cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then uses a per-provider bullet structure that keeps the dense scope/behavior details scannable. It is long, but nearly every clause carries actionable information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers config requirements, credential handling, and side effects thoroughly. It omits what the call returns and how to later disable enrichment, which are minor gaps against an otherwise complete picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description goes further — it spells out which scopes the Meta user_access_token requires (pages_show_list, leads_retrieval, business_management) and that omitting advertiser_ids subscribes every reachable ad account. It adds real meaning over the generic per-field schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Enable enrichment on an agent's webhook trigger' — and clarifies scope with 'One tool for every provider,' which distinguishes it from the provider-specific or trigger-management siblings. An agent knows exactly what this does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when this applies (reference-only payloads needing resolution before the agent runs) and how behavior differs by situation (reuses existing URL, creates a trigger if none exists, adds a separate trigger for TikTok). It does not explicitly name alternatives such as agents_trigger_create/update, so the 'vs sibling' routing is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web__local_searchARead-onlyIdempotentInspect
Multi-source web research with citations. Returns a synthesized answer with numbered [^1] markers and a citations array of {url, title, snippet, index}. Use for evidence-backed synthesis (competitive analysis, regulatory summary, whitepaper section). For quick fact lookups use web.search instead.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Research question. Specific scoped questions outperform vague keywords. | |
| language | No | Search language hint (BCP-47, e.g. 'en', 'ru'). Defaults to 'en'. The synthesis output language matches the query language regardless. | en |
| num_sources | No | How many top search results to fetch and synthesize (1-4, default 4). Lower = faster + cheaper, higher = more comprehensive. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the description is free to add the return contract and the cost/latency trade-off of num_sources, plus the reassurance that in_workspace stores nothing and affects no other session. Minor tension: 'web research' sits awkwardly against openWorldHint=false, though this reads as a curated/local index rather than a hard contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose, then return shape, then usage routing. Every sentence carries distinct information and none restates the name or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly supplies the return shape (markers plus citations array of url/title/snippet/index). Safety is covered by annotations, all four parameters are documented, and the sibling routing is present, so nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so query, language, num_sources, and in_workspace are already fully documented with types, defaults, and bounds. The description adds no parameter-level syntax or semantics beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Multi-source web research with citations') and describes the concrete artifact returned (synthesized answer with [^1] markers plus a citations array). It distinguishes itself from web.search, but leaves its boundary against the close sibling web_research unexplained, so it is clear rather than fully disambiguated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use cases ('evidence-backed synthesis: competitive analysis, regulatory summary, whitepaper section') and an explicit alternative for the opposite case ('For quick fact lookups use web.search instead'). The selection rule is stated outright rather than left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_researchARead-onlyIdempotentInspect
Answer a research question from live web sources in one call — returns a synthesized answer with numbered [N] citation markers and a citations array of {url, title, index}. Supports recency and domain filters. Use for questions needing current, sourced information (news about a company, market state, comparisons). For raw search result links use web.search; mode='deep' runs minutes-long exhaustive research — only when explicitly requested.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Research depth: 'quick' (default, seconds, cheapest), 'pro' (harder questions, better sourcing), 'deep' (autonomous multi-step research, takes minutes — only when the user explicitly wants an exhaustive report). | quick |
| query | Yes | Research question. Specific scoped questions outperform vague keywords. | |
| domains | No | Restrict search to these domains (max 10), e.g. ['coindesk.com', 'cointelegraph.com']. Prefix with '-' to exclude a domain. | |
| recency | No | Only use sources from this window. Omit for no limit. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, openWorld, non-destructive), so the bar is lower; the description adds genuinely useful context beyond them, namely the return format (numbered citation markers plus citations array) and the cost/latency tradeoff of deep mode. It does not mention rate limits, caching, or source-quality caveats, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single dense paragraph that front-loads what the tool returns before covering filters, alternatives, and the deep-mode caveat. Every clause carries information, though the paragraph is packed enough that a short bulleted split would scan faster.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully describes the return shape itself, and it covers routing versus the web_search sibling. It omits anything about result freshness guarantees, source quality, or failure modes, so it is strong but not exhaustive for a research tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters including mode, recency, and domain filtering. The description only restates that recency and domain filters exist and echoes the deep-mode cost warning, adding little beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (answer a research question) plus resource (live web sources) and describes the output shape (synthesized answer with [N] markers and a citations array). It also explicitly differentiates itself from the sibling web_search tool for raw links.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the alternative (web.search for raw result links) and gives explicit conditions for each mode, including the warning that mode='deep' runs minutes-long research and should only be used when explicitly requested. It also lists concrete use cases (news about a company, market state, comparisons).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchARead-onlyIdempotentInspect
Search the web for current information, news, facts, prices, or events. Use this when the user asks about something that requires up-to-date information from the internet, or when internal knowledge base doesn't have the answer. Examples: recent news, stock prices, weather, product information, current events.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query - what to search for on the web. | |
| time_range | No | Only return results published within this window. The provider applies it, so stale articles never reach you — prefer this over judging freshness from a snippet. Applies to search_type='news'. OMIT for no age limit. | |
| num_results | No | Number of results to return (1-10). | |
| search_type | No | Type of search: 'search' for general web, 'news' for news articles. | search |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is fully covered by structured data. The description adds nothing behavioral beyond that — no rate limits, result shape, or latency notes — so it neither adds value nor contradicts anything.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose in the first clause and no wasted preamble. The trailing examples sentence partially repeats the category list already named in the first sentence ('news, facts, prices, or events' vs 'recent news, stock prices, weather, product information'), which is mild redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a fully annotated, fully schema-documented search tool with no output schema, the description covers purpose and invocation conditions adequately. Nothing essential to selecting or calling the tool is missing, though sibling disambiguation would be a nice refinement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the enum parameters (time_range, search_type) are documented in detail within the schema itself, so the baseline is 3. The description's example list ('recent news, stock prices, weather') gives mild intuition about what the query parameter accepts but adds no syntax or constraint detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search the web for current information, news, facts, prices, or events'), which clearly separates it from internal-retrieval siblings like knowledge_query or notes_search. It does not, however, explicitly differentiate itself from adjacent web siblings such as web_fetch, web_research, or web__local_search, so an agent still has to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete trigger ('when the user asks about something that requires up-to-date information from the internet, or when internal knowledge base doesn't have the answer') plus illustrative examples. There is no explicit when-not guidance and no named alternative tool to route to, which keeps it short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
widgets_createBInspect
Create a new livechat widget for your website.
The widget will be created with default settings. You can customize theme, auto-reply mode, and more.
Use this when user wants to add a chat widget to their site.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Name for the widget (e.g., 'Website Chat', 'Support Widget') | |
| position | No | Widget position on screen | bottom-right |
| allow_voice | No | Master switch for voice — set true to show the mic so visitors can talk to the agent. Everything else voice-related is inert until this is on. The mic also needs a voice-capable agent in the workspace: one with an enabled incoming_call trigger. OMIT to leave voice off (the default). | |
| display_mode | No | Visual mode of the widget. Pick exactly one: - 'chat' (default): full chat panel + voice mic — use for support / sales / general. - 'voice_only': mic-only bubble that launches a voice call directly — pick only when the user explicitly asks for a voice-only widget (e.g. 'just a voice button', 'no chat, just call'). - 'headless': no UI; customer drives via window.DialogBrain JS API — pick only when the user explicitly says 'embed in our own design' / 'no widget chrome'. | chat |
| header_title | No | Title shown in chat header | Chat with us |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| primary_color | No | Primary color for widget theme (hex, e.g., '#2563eb') | #2563eb |
| auto_reply_mode | No | Auto-reply mode: 'draft' (review before sending) or 'auto' (send immediately) | draft |
| voice_button_label | No | Localized aria-label and hover tooltip for the voice-only mic bubble (only used when display_mode='voice_only'). ≤ 100 chars. Defaults to 'Talk to agent' if omitted. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=false, giving a clear safety profile. The description adds that the widget is created with default settings and can be customized later, which is useful context about post-creation behavior. However, it does not disclose any additional traits beyond what annotations provide (e.g., no mention of required permissions or side effects), so it meets but does not exceed expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three brief sentences, front-loading the core action ('Create a new livechat widget') and then adding contextual details. It is well-structured and wastes no words, though the third sentence about customization is somewhat redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (9 parameters, most optional, no nested objects, no output schema), the description covers the essential purpose and a basic usage scenario. However, it lacks details on what happens on failure, whether the widget is immediately live, or any versioning/undo considerations. It is adequate but not rich enough to be fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, meaning all parameters are well-documented in the schema itself. The description adds no extra parameter semantics; it only vaguely mentions customizable aspects like theme and auto-reply mode, which are already covered in the schema. A baseline score of 3 is appropriate when the schema carries the full semantic load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Create a new livechat widget for your website'), making the purpose immediately clear. It is distinguishable from sibling tools like widgets_list, widgets_update, and widgets_get by the create verb, but it adds no explicit differentiation in the text itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a usage scenario ('Use this when user wants to add a chat widget to their site'), which gives context for when to invoke it. However, it does not mention alternatives (e.g., widgets_update for modifying existing widgets) or prerequisites, so it is only minimally adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
widgets_deleteADestructiveIdempotentInspect
Delete a livechat widget permanently.
This will remove the widget and its embed code will stop working. Its conversations and their messages are deleted with it: the widget IS the channel, and thread.channel_account_id cascades.
Use this when user wants to remove a chat widget.
| Name | Required | Description | Default |
|---|---|---|---|
| widget_id | Yes | ID of the widget to delete | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, but the description adds substantial behavior the annotations cannot convey: the embed code stops working, and conversations plus their messages are deleted because the widget IS the channel with thread.channel_account_id cascading. This is exactly the blast-radius detail an agent needs before invoking a destructive tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the destructive verb and permanence, then the consequences, then the usage trigger. No filler and no repetition of the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter destructive delete with no output schema, the description covers permanence, collateral destruction (conversations/messages), and the cascade mechanism. There is nothing else an agent needs in order to call this correctly and safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so widget_id and in_workspace are already documented in the schema ('ID of the widget to delete', the workspace override semantics). The description adds no syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+scope: 'Delete a livechat widget permanently.' This is unambiguously distinguishable from siblings like widgets_update, widgets_get, and widgets_get_embed_code without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the triggering condition: 'Use this when user wants to remove a chat widget.' It gives clear context for when to invoke, but names no alternatives or exclusions (e.g., when to prefer widgets_update for non-destructive changes), which is the only thing separating it from a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
widgets_getARead-onlyIdempotentInspect
Get full configuration of a single livechat widget.
Returns all settings including theme, identification, actions, and more.
Use this when user wants to see or verify a specific widget's settings.
| Name | Required | Description | Default |
|---|---|---|---|
| widget_id | Yes | ID of the widget to retrieve | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds useful return content ('theme, identification, actions') but says nothing about failure modes such as an unknown widget_id or whether 'and more' hides fields. Adequate, not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, no filler. The trailing 'and more' is a minor hedge but overall the text is tight and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only getter with full schema coverage and complete annotations, the description covers what it does, what it returns, and when to call it. No output schema exists, and the summary of returned settings partially compensates; only error/edge behavior is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (widget_id, in_workspace) are documented in the schema, including the non-obvious in_workspace session semantics. The description adds no parameter-level meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get full configuration of a single livechat widget') with scope made explicit by 'single', which implicitly separates it from the plural widgets_list and from widgets_get_embed_code. It stops short of naming those siblings, so it is clear but not maximally differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this when user wants to see or verify a specific widget's settings' gives a concrete triggering condition. It does not mention alternatives (e.g. widgets_list for enumerating widgets, widgets_get_embed_code for embed markup), so there is no exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
widgets_get_embed_codeARead-onlyIdempotentInspect
Get the embed code snippet for a livechat widget.
Returns HTML/JavaScript code to add to your website. The code should be placed before the closing tag.
Use this when user wants to install the chat widget on their site.
| Name | Required | Description | Default |
|---|---|---|---|
| widget_id | Yes | ID of the widget to get embed code for | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so safety is covered. The description adds useful behavioral context by stating the return format (HTML/JavaScript code) and where to place it (before </body>), which goes beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core purpose, then return format, placement guidance, and usage context. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity, full schema coverage, and absence of an output schema, the description is complete enough: it explains what the tool returns and when to use it. Nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are fully documented in the input schema. The description does not add any new parameter-level meaning beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Get the embed code snippet for a livechat widget.' It clearly distinguishes this tool from sibling widgets_get/list/update by focusing on embed code retrieval, and it specifies the exact output format.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear condition for use: 'Use this when user wants to install the chat widget on their site.' However, it does not name alternatives or explicitly state when not to use it, so it falls short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
widgets_listARead-onlyIdempotentInspect
List all livechat widgets.
Returns widgets with their configuration, embed code, and status.
Use this when user wants to see their widgets or chat widgets.
| Name | Required | Description | Default |
|---|---|---|---|
| active_only | No | Only return active widgets. OMIT to include inactive widgets too. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint false, and openWorldHint false, covering the safety profile. The description adds return content ('configuration, embed code, and status'), which is useful because there is no output schema, but it discloses no further behavioral traits such as permissions, rate limits, or default inactive inclusion (the schema handles active_only).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description uses three short sentences with no filler: purpose first, then return content, then usage context. Every sentence earns its place and the key information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description usefully summarizes return fields, and the annotations plus schema cover safety and parameters. It is mostly complete for a simple list tool, though it omits any mention of pagination, sorting, or explicit sibling routing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both optional parameters fully documented in the schema. The description adds no extra meaning or syntax beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List all livechat widgets.' It is clearly a list operation, but it does not differentiate itself from sibling tools such as widgets_get or widgets_get_embed_code, leaving sibling selection to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit usage context: 'Use this when user wants to see their widgets or chat widgets.' However, it offers no when-not guidance or direct alternatives, so an agent must infer when to prefer a different widget tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
widgets_updateAInspect
Update an existing livechat widget configuration.
You can change name, theme, auto-reply mode, and other settings. Only provided fields will be updated.
Use this when user wants to modify their chat widget settings.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | New name for the widget | |
| position | No | Widget position on screen. OMIT to leave the position unchanged. | |
| is_active | No | Enable or disable the widget. OMIT to leave the active flag unchanged. | |
| widget_id | Yes | ID of the widget to update | |
| allow_voice | No | Master switch for voice on this widget — set true to show the mic and let visitors talk to the agent. Everything else voice-related (greeting, button label, STT/TTS from the agent's own config) is inert until this is on. The mic also needs a voice-capable agent in the workspace: one with an enabled incoming_call trigger. OMIT to leave the setting unchanged. | |
| website_url | No | Website URL for product/site search integration | |
| calendly_url | No | Booking URL for calendar action (e.g., 'https://calendly.com/yourname') | |
| color_scheme | No | Widget color scheme. 'auto' follows the visitor's OS dark/light mode preference. OMIT to leave the color scheme unchanged. | |
| display_mode | No | Visual mode of the widget. Pick exactly one: - 'chat': full chat panel + voice mic — default for support / sales / general. - 'voice_only': mic-only bubble that launches a voice call directly — pick only when the user explicitly asks for a voice-only widget. - 'headless': no UI; customer drives via window.DialogBrain JS API — pick only when the user explicitly says 'embed in our own design'. OMIT to leave the display mode unchanged. | |
| header_title | No | Title shown in chat header | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| greeting_text | No | Custom greeting message shown when visitor opens the chat (e.g., 'Hello! How can I help you today?') | |
| primary_color | No | Primary color for widget theme (hex, e.g., '#2563eb'). Paints the header, the visitor's message bubbles and the send button — and the launcher bubble too unless launcher_color overrides it. | |
| launcher_color | No | Color of the closed launcher bubble ONLY (hex, e.g., '#ffffff'). Use when the site pairs a light button with a dark panel and one colour cannot express both. Pass an empty string to clear it and let the launcher follow primary_color. Text and glyphs pick themselves from the background, so a light value stays readable. | |
| voice_greeting | No | Spoken opening line when a visitor starts a voice call through this widget. Played via TTS before the AI model runs. Empty string disables the greeting. Requires allow_voice=true to be audible. | |
| allowed_domains | No | List of allowed domains for the widget | |
| auto_reply_mode | No | Auto-reply mode: 'draft' or 'auto'. OMIT to leave the auto-reply mode unchanged. | |
| header_subtitle | No | Subtitle shown in chat header | |
| greeting_enabled | No | Enable or disable the proactive greeting. OMIT to leave this flag unchanged. | |
| greeting_behavior | No | notification = show badge after delay; auto_open = open widget automatically after delay; on_open = greet only when visitor manually opens. OMIT to leave the greeting behavior unchanged. | |
| enable_form_action | No | Enable or disable the contact form action button. OMIT to leave this flag unchanged. | |
| voice_button_label | No | Localized aria-label and hover tooltip for the voice-only mic bubble (only used when display_mode='voice_only'). ≤ 100 chars. Defaults to 'Talk to agent' if not set. | |
| contact_form_fields | No | Fields to collect in contact form (e.g., ['name', 'email', 'phone']) | |
| enable_search_action | No | Enable or disable the search action button. OMIT to leave this flag unchanged. | |
| show_visitor_history | No | Show full chat history to returning visitors. OMIT to leave this flag unchanged. | |
| identification_fields | No | Fields to require for visitor identification (e.g., ['name', 'email']) | |
| enable_calendar_action | No | Enable or disable the calendar booking action button. OMIT to leave this flag unchanged. | |
| greeting_delay_seconds | No | Delay in seconds before the proactive greeting appears (0–300). 0 = send immediately on page load. Default: 30. | |
| require_identification | No | Require visitor to identify before chatting. OMIT to leave the identification policy unchanged. | |
| returning_greeting_text | No | Greeting for returning visitors who already have chat history (e.g., 'Welcome back! How can I help you today?'). Falls back to greeting_text if not set. | |
| max_voice_duration_seconds | No | Hard cap on a single voice call from this widget, in seconds (default 300). OMIT to leave the cap unchanged. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate a non-destructive, non-idempotent, non-read-only update. The description adds a useful contextual detail: 'Only provided fields will be updated' (partial update semantics). However, it omits other relevant behavioral aspects such as required permissions, side effects, or response format. Since annotations cover the safety profile, this partial disclosure is adequate but not thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loading the core action (update widget configuration) in the first sentence, followed by scope and usage context. No superfluous information, though it could be slightly more structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 31-parameter update tool with full schema documentation, no output schema, and clear annotations, the description provides the essential high-level purpose and partial-update behavior. It is largely complete given the rich structured data, though it could include more on permissions or side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully documents all 31 parameters with comprehensive descriptions, enums, and OMIT semantics. The description only lists a few example fields ('name, theme, auto-reply mode, and other settings') without adding any new semantic insight. With 100% schema coverage, the description meets the baseline of 3 by not contradicting or degrading parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states a specific verb (update) and resource (livechat widget configuration) and directly names which fields can be modified. Does not explicitly differentiate itself from sibling tools like widgets_create or widgets_delete, but the purpose is otherwise unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a single usage cue ('Use this when user wants to modify their chat widget settings'), which is helpful but vague. It does not specify when to choose this tool over alternatives such as widgets_create or when not to use it. There is no mention of prerequisites like widget ownership.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_run_pythonARead-onlyIdempotentInspect
Run Python in an isolated sandbox to process LARGE or paginated tool results without pulling every row into the conversation. Inside the code, call your connected integration tools with call_tool('ext<id>_<name>', {..}), or this agent's own platform tools by their dotted id (e.g. call_tool('db.query', {'sql': 'SELECT ...'})) — a platform tool must be in the agent's allowed_tools, and calling one requires an agent context. RETURN SHAPE: call_tool ALWAYS returns a dict with a boolean r['success']. The payload key DIFFERS by tool kind: integration (ext*) results are under r['body'] (e.g. r['body']['results']), platform tools are under r['data'] (e.g. r['data']['rows'] for db.query). Reading the wrong key returns nothing even though the call SUCCEEDED — so when in doubt print(r) once and inspect before extracting. On FAILURE r['success'] is False and r['error'] explains. Aggregate/filter/paginate in the sandbox, then assign ONLY the small summary you want back to a variable named result. FIRST discover exact tool slugs with integrations_search_tools, THEN write code that calls them. pandas/numpy available.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | Python source to execute. call_tool(slug, {..}) returns a dict: for integrations ('ext<id>_<name>') the HTTP payload is under r['body'] (e.g. r['body']['results']); for platform tools (dotted ids like 'db.query') it is under r['data'] (e.g. r['data']['rows']). Failure is {'success': False, 'error': ...}. Assign the small summary to `result`. pandas/numpy available. | |
| agent_id | No | Which agent's tool policy the sandbox runs under — this scopes which ext* integrations call_tool may reach (enabled + denied_tools for that agent). Only needed when calling this tool OUTSIDE a normal agent run (e.g. directly from an external MCP client); during an agent run the running agent is used and this is ignored. Without it, call_tool can reach no integrations. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover only the safety profile (readOnly/idempotent/openWorld); the description goes well beyond by disclosing the call_tool return contract (r['success'], r['body'] vs r['data'] per tool kind), the failure shape (r['error']), the allowed_tools + agent-context requirement for platform tools, and the 'assign to `result`' output convention. These are non-obvious behaviors an agent would otherwise get wrong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then mechanics, then the discovery step. Dense and mostly earned, though the return-shape detail is repeated almost verbatim in the schema's `code` description, so some redundancy exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, no-output-schema execution tool it covers return shape, failure mode, output convention, and tool-discovery workflow. It omits practical runtime details an agent might want (execution timeout, resource limits, whether sandbox files/state persist across calls), which keeps it short of 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three params are already documented in the schema, and the description largely restates the `code` contract. It adds the behavioral note that a platform tool must be in allowed_tools and that without agent_id no integrations are reachable, but that is more behavioral than parameter-level. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+mechanism ('Run Python in an isolated sandbox') and immediately scopes the use case ('process LARGE or paginated tool results without pulling every row into the conversation'). No sibling in the list does code execution, so the agent can distinguish it instantly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to reach for it (large/paginated results that shouldn't enter the conversation) and an explicit prerequisite workflow ('FIRST discover exact tool slugs with integrations_search_tools, THEN write code'). It lacks an explicit when-NOT (e.g. use db_query directly for a small result), which keeps it from a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_createAInspect
Create a new team workspace owned by the calling user. Returns its id and slug — pass either to workspace.switch to start working inside it. slug is derived from name when omitted. NOT idempotent: calling twice with the same name yields two workspaces (the second gets a -1 slug suffix), so check workspace.list first if the workspace may already exist. A user may own at most 25 team workspaces.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Workspace display name, e.g. 'TUI BLUE'. | |
| slug | No | URL-friendly slug: lowercase letters, digits and single dashes between them (e.g. 'tui-blue'), max 100 chars. Omit to derive one from `name` (lowercased, spaces to dashes, uniqueness suffix added when taken). An explicit slug is NOT auto-uniquified — a collision is an error. | |
| description | No | Optional workspace description. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=false, but the description goes well beyond them: it explains *why* (a second call yields a duplicate workspace with a `-1` slug suffix), states a hard quota (max 25 owned team workspaces), and describes return values in the absence of an output schema. This is exactly the added context the annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and return values, then the operational caveats (non-idempotency, quota). Every sentence carries distinct, decision-relevant information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a creation tool with no output schema and no annotations on side effects, the description covers the return payload, the non-idempotent duplicate behavior, the slug derivation rule, and the ownership quota — everything an agent needs to call it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters including slug derivation and the collision-error behavior. The description's 'slug is derived from name when omitted' largely restates the schema rather than adding new syntax or format meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Create a new team workspace') with scope ('owned by the calling user') and an immediate statement of what it returns. It is clearly separable from siblings like workspace_list, workspace_switch, and workspace_update.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: check `workspace.list` first if the workspace may already exist, and pass the returned id/slug to `workspace.switch` to work inside it. Both the when-to-use and the alternative-tool path are stated, not implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_currentARead-onlyIdempotentInspect
Return the workspace this MCP API key is currently routed to, with the caller's role inside it. Use this to confirm context before/after workspace.switch.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so safety is covered. The description adds value by disclosing the resolution basis (the API key's current routing) and what the response contains (workspace + caller's role), which matters since there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the purpose front-loaded and the usage trigger second.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description names the returned fields, and the parameter is fully documented in-schema. Only minor gaps remain, such as whether the result reflects session vs. key routing in edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the sole parameter's non-persistent, single-call semantics are documented there. The description says nothing about `in_workspace`, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Return) and resource (the workspace this API key is routed to) plus the payload detail (caller's role inside it). This clearly separates it from workspace_list and workspace_switch without needing to open any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use this to confirm context before/after `workspace.switch`" gives an explicit trigger condition tied to a named sibling. It stops short of stating when NOT to use it or offering workspace_list as the alternative for enumerating workspaces.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_desktopsARead-onlyIdempotentInspect
List the Claude Code sessions currently connected to this workspace over a channel. Each one can be handed a task by name. Use this before assigning work to a claude_channels agent when more than one session may be open, so the task goes to the machine you mean.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive/open-world status, so the safety profile is covered. The description adds real value beyond that: it explains that sessions are connected 'over a channel' and each can be handed a task by name, which is the operative behavioral trait for the calling agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the purpose before the routing guidance. Reasonably tight, though the trailing 'so the task goes to the machine you mean' slightly restates the preceding when-to-use clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-required-param, read-only list tool with annotations carrying the safety profile, the description covers purpose, target resource, and the decision context. It never describes the shape of the returned list (names/identifiers used to assign tasks), a minor gap for a tool whose output drives a follow-up call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single in_workspace parameter is fully documented in the schema, including its non-persistent, non-affecting-others behavior. The description adds nothing about parameters, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'List the Claude Code sessions currently connected to this workspace over a channel.' This clearly identifies what is returned and clarifies the ambiguous name 'desktops'. It doesn't explicitly distinguish itself from any sibling tool by name, so it stops short of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this before assigning work to a claude_channels agent when more than one session may be open' gives a concrete condition for when to call it. It lacks a when-not statement or a named alternative tool, but the trigger condition is explicit and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_inviteAInspect
Invite a person to the active workspace (or one named by workspace_id) by email address. Requires owner or admin role. The backend emails them a join link and the invite expires in 7 days; the same link is returned as invite_url so you can pass it on yourself. Inviting an address that already has an account works too — they see the invitation in the app once logged in. Use workspace.invites to see what is still pending.
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | Role to grant on acceptance. OMIT to invite as a regular member. 'admin' additionally lets them manage members and workspace settings. | |
| Yes | Email address to invite. | ||
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| workspace_id | No | Numeric workspace id to invite into. Omit to target the currently-routed workspace (workspace.current). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavior beyond the annotations (which only declare non-readonly, non-idempotent, non-destructive): backend emails a join link, the invite expires in 7 days, the link is echoed as invite_url for manual forwarding, and existing accounts see it in-app. These are exactly the operational facts an agent needs and cannot get from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded: the action, then the authorization requirement, then the side effects, then the sibling pointer. Every sentence carries information, though the parenthetical redundantly restates the schema's workspace_id default.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly discloses the return value (invite_url) and covers authorization, expiry, and the already-registered edge case. Nothing material is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each parameter already documents itself, including the OMIT-to-default behavior for role and workspace_id. The description restates the workspace targeting and role outcome without adding syntax or format detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Invite) plus resource (a person to the workspace) and the targeting mechanism (active workspace or workspace_id). It is clearly distinguishable from the sibling workspace_invites and workspace_member_update tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the prerequisite (owner or admin role) and routes the agent to the sibling tool for the adjacent need: 'Use workspace.invites to see what is still pending.' It also flags the non-obvious case where the invitee already has an account.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_invite_revokeADestructiveIdempotentInspect
Revoke a pending invitation to the active workspace (or one named by workspace_id) so its link stops working. Requires owner or admin role. Get the invite id from workspace.invites. An invite that is already accepted, expired or revoked answers not_found.
| Name | Required | Description | Default |
|---|---|---|---|
| invite_id | Yes | Id of the pending invite, from workspace.invites. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| workspace_id | No | Numeric workspace the invite belongs to. Omit to target the currently-routed workspace (workspace.current). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the mutation profile is covered. The description adds meaningful context beyond them: the owner/admin authorization requirement, the concrete effect (link stops working), and the not_found outcome for invites that are already accepted, expired, or revoked.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the primary action and effect, then prerequisites, then edge-case outcome. No sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still covers what an agent needs: required role, effect of the call, id provenance, and the failure mode. Nothing material for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three parameters are already documented, including the workspace_id fallback and in_workspace semantics. The description only adds the source of invite_id (workspace.invites), which largely duplicates the schema. Baseline 3 is correct when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (revoke) and resource (pending invitation), with the exact scope (active workspace or one named by workspace_id) and the effect (link stops working). An agent can distinguish it from sibling workspace_invite (create) and workspace_invites (list) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the prerequisite (owner or admin role), the condition (pending invites only), and where to source the id from workspace.invites. It stops short of explicitly naming alternatives, but the 'pending / not_found for accepted, expired, revoked' clause effectively routes the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_invitesARead-onlyIdempotentInspect
List the invitations to the active workspace (or one named by workspace_id) that are still pending — who was invited, as what role, and when the invite expires. Requires owner or admin role. Join links are not returned; revoke with workspace.invite_revoke or re-send with workspace.invite.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| workspace_id | No | Numeric workspace id to read. Omit to target the currently-routed workspace (workspace.current). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds real value beyond that: the owner/admin authorization requirement and the pending-only filter, plus the explicit exclusion of join links. It does not mention pagination or result limits, which is the only meaningful gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and scope, followed by prerequisites and sibling routing. Every clause earns its place — the field list, the role requirement, and the exclusion all carry information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by naming the returned fields (invitee, role, expiry). Combined with the auth requirement and the pending-only scope, an agent has enough to call it correctly. Only pagination/result-size behavior is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already fully documented in the schema, including the in_workspace vs workspace_id distinction. The description restates the workspace_id default ('or one named by workspace_id') but adds no syntax or behavior beyond the schema. Baseline 3 applies when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource+scope: 'List the invitations ... that are still pending' on the active or named workspace. It pre-empts confusion with sibling tools by stating join links are not returned and pointing to workspace_invite_revoke / workspace_invite for those actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the prerequisite ('Requires owner or admin role'), the default target ('active workspace (or one named by workspace_id)'), and routes follow-up actions to the correct siblings for revoking or re-sending invites. An agent knows both when to call this and what to call instead for related operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_listARead-onlyIdempotentInspect
List every workspace the caller is a member of, with is_current marking the workspace this MCP key is currently routed to. Pair with workspace.switch to change the active workspace without reconnecting.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world). The description adds value beyond them by defining the is_current field and noting that switching workspaces via workspace.switch does not require reconnecting. It does not discuss pagination or ordering, but for a simple read tool this is solid added context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action and scope, followed by the workflow hint. No redundant restatement of the name or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only list tool with no output schema, the description supplies the one piece of return-value meaning an agent actually needs (is_current) plus the follow-up action. Minor gaps around ordering or pagination remain, but nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single in_workspace parameter is fully documented in the schema. The description says nothing about it, so baseline 3 applies: the schema carries the semantics and the description neither compensates nor detracts.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'List every workspace the caller is a member of.' That scope naturally distinguishes it from workspace_current (single workspace) and workspace_members (people in a workspace), though it doesn't explicitly name those siblings. It also explains the is_current marker, so the agent knows what the payload means.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives a concrete workflow: pair with workspace.switch to change the active workspace without reconnecting. That tells the agent when this tool fits into a sequence, but it does not state when to prefer this over workspace_current or list explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_member_removeADestructiveIdempotentInspect
Remove a person from the active workspace (or one named by workspace_id), and revoke the channel shares they had granted into it. Requires owner or admin role. Name them by email, username or numeric user id — workspace.members lists them. The workspace owner cannot be removed.
| Name | Required | Description | Default |
|---|---|---|---|
| member | Yes | Who to remove: a member's email, username, or numeric user id. Call workspace.members to see who is in the workspace. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| workspace_id | No | Numeric workspace id. Omit to target the currently-routed workspace (workspace.current). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and idempotentHint=true, and the description still adds substantive behavior beyond them: the removal also revokes channel shares the member granted, it is permission-gated to owner/admin, and the owner is an unremovable exception. That is exactly the cascade and authorization context an agent needs before firing a destructive call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences: action and effect first, then the permission prerequisite, then the identifier guidance and the owner exception. No filler, nothing buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, idempotent membership mutation with a fully documented 3-param schema and no output schema, the description covers effect, cascade, authorization, target identification, and the one unremovable case. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents `member`, `in_workspace`, and `workspace_id` in similar terms. The description restates the accepted identifier forms (email, username, numeric id) without adding syntax or format detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb + resource ('Remove a person from the active workspace') plus scope clarifier (default workspace or one named by `workspace_id`) and the cascading effect on channel shares. An agent can distinguish this from workspace_member_update and workspace_invite_revoke without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real preconditions: requires owner or admin role, the target must be named by email/username/numeric id, workspace.members is the discovery surface, and the workspace owner cannot be removed. It lacks an explicit 'to change role or modify a member instead, use workspace_member_update' routing statement, which is the only remaining gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_membersARead-onlyIdempotentInspect
List the members of the current workspace with their roles, so you know who can be assigned a task. Set include_agents to also list the workspace's AI agents, which can hold tasks too.
| Name | Required | Description | Default |
|---|---|---|---|
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| include_agents | No | Also list this workspace's AI agents (default: true). Agents can be task assignees just like people. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so the safety profile is covered. The description adds useful scope context (defaults to the current workspace, agents are assignable like people), but says nothing about return shape or ordering, so it is solid rather than rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the core purpose is front-loaded ahead of the optional parameter note. Every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with full annotation coverage, a fully described two-parameter schema, and no output schema, the description covers scope, default workspace behavior, and the agent-inclusion nuance. Only the returned fields beyond 'roles' remain unstated, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema, including the include_agents default of true and the non-persistent nature of in_workspace. The description restates include_agents behavior and never mentions in_workspace, so it adds little beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('List the members of the current workspace with their roles') and adds intent ('so you know who can be assigned a task'), which clearly separates it from workspace_list or workspace_member_update. It does not, however, name any sibling tool explicitly to route the agent, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied through the purpose clause ('so you know who can be assigned a task') and the include_agents note; there is no explicit statement of when to prefer this tool over workspace_list or workspace_member_remove. The agent must infer the context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_member_updateAInspect
Change a member's role in the active workspace (or one named by workspace_id). Requires owner or admin role. Name the person by email, username or numeric user id — workspace.members lists them. Promoting to 'admin' also grants member management; ownership is moved separately and cannot be set here.
| Name | Required | Description | Default |
|---|---|---|---|
| role | Yes | New role for this member. | |
| member | Yes | Who to change: a member's email, username, or numeric user id. Call workspace.members to see who is in the workspace. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| workspace_id | No | Numeric workspace id. Omit to target the currently-routed workspace (workspace.current). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), but the description adds real context beyond them: the owner/admin authorization requirement and the side effect that promoting to 'admin' grants member management. It does not say whether a change can be reverted or what happens on failure, so it falls just short of the top.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each load-bearing: action, precondition plus identification, then the role-semantics caveat. The core action is front-loaded and no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the important ground: authorization, who to target, how to find members, and what promotion does and does not grant. Only minor gaps remain (reversibility, error behavior) that an agent could infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so role, member, in_workspace, and workspace_id are all already documented in the schema. The description restates the member identification formats and the workspace_id targeting, adding little new semantics, and never addresses the in_workspace override. Baseline 3 applies when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Change a member's role') plus the scope of the target workspace. It is immediately distinguishable from siblings like workspace_member_remove, workspace_invite, and workspace_members without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the precondition ('Requires owner or admin role'), the identification options for the member, and a pointer to workspace.members for discovery. It also states an explicit exclusion — ownership is moved separately and cannot be set here — which is exactly the when-not guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_switchAInspect
Switch the workspace THIS MCP session works in. Pass exactly one of workspace_id or slug (find them via workspace.list). Takes effect on the very next tool call — no MCP reconnect, no new API key. The switch is scoped to the current MCP session (the response says scope: session), or, for a client that sends no MCP session id, to that client (scope: client: every conversation of that client on this key moves together). Other sessions using the same API key are not affected, and the key's own workspace stays the default for new sessions. Sequential checkpoint: do not parallelize tool calls across a switch — calls already in flight when the switch commits will run against the previous workspace. Keys created with workspace_locked=true refuse to switch in either scope (they are pinned to their workspace); use a key belonging to the target workspace instead.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | No | Workspace slug to switch to. Resolved within the caller's memberships, so cross-tenant slug collisions are not possible. Mutually exclusive with `workspace_id`. | |
| workspace_id | No | Numeric workspace id to switch to. Mutually exclusive with `slug`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (which only say non-readonly, non-idempotent, non-destructive). It discloses that the switch takes effect on the next call with no reconnect, that other sessions on the same key are unaffected, that the key's default workspace is preserved, that in-flight calls run against the previous workspace, and that locked keys are refused.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded and each sentence carries information (scoping, effect timing, key behavior, sequencing, locked keys). It is dense and fairly long, but the length is justified by the intricacy of session/client scope semantics rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and minimal annotations, the description carries the full burden and does so: it covers effect timing, scope, side effects on other sessions, error conditions (locked keys), and even mentions the response's scope field. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both slug and workspace_id fully documented including mutual exclusivity and slug resolution. The description reinforces the 'exactly one' requirement, which is mildly useful, but adds no syntax or format detail beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (switch the workspace this MCP session works in) and scopes it precisely to the session. The name alone could be ambiguous against siblings like workspace_current or workspace_list, but the description immediately clarifies what it changes and at what scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to workspace.list to obtain ids, states the exactly-one-of constraint, and gives concrete when-not guidance: workspace_locked keys refuse to switch, so use a key belonging to the target workspace instead. It also warns not to parallelize calls across a switch.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_updateAInspect
Edit the active workspace (or one named by workspace_id): its name, description, logo_url, and/or settings. Requires ADMIN role in that workspace. settings is MERGED, not replaced — send only the keys you want to change (e.g. {"translation_glossary": "..."} or {"translation_tts_provider": "cartesia"}); other keys are left intact. Omit workspace_id to target the workspace this MCP key is currently routed to (see workspace.current).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | New workspace name. Omit to leave unchanged. | |
| logo_url | No | New workspace logo URL. Omit to leave unchanged. | |
| settings | No | Partial settings object, MERGED into the existing workspace.settings JSONB (PATCH semantics). Only the keys you send change. Common keys: translation_glossary, translation_mt_provider, translation_tts_provider, translation_tts_voice, translation_languages, translation_default_lang. | |
| description | No | New workspace description. Omit to leave unchanged. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| workspace_id | No | Numeric workspace id to update. Omit to target the currently-routed workspace (workspace.current). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare write/non-destructive/idempotent=false; the description adds two high-value traits: an ADMIN role requirement and, crucially, that `settings` is MERGED (PATCH) rather than replaced, with concrete examples. These are exactly the behavioral facts an agent needs before mutating a workspace.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and target, then layers auth, merge semantics, and default targeting in descending priority. Dense but every clause carries actionable information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the essentials for a 6-param mutation with no output schema: target resolution, permissions, and merge behavior. Minor gaps remain (in_workspace override behavior and expected response), but with annotations carrying the safety profile these are not blocking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning: it explains the workspace_id omission default and illustrates the merge/patch semantics of the nested settings object with example payloads. It leaves in_workspace unexplained, so it is not a full 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ("Edit") plus resource ("workspace") and an explicit enumeration of the editable fields (name, description, logo_url, settings). An agent can distinguish this from workspace_create (new) and workspace_current (read) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly states the context (edit the active workspace or one addressed by workspace_id) and the default targeting behavior when workspace_id is omitted, referencing workspace.current. It does not explicitly name when-not or route to siblings like workspace_member_update, so it stops short of full alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
x_publish_postADestructiveInspect
Publish a post to the connected X (Twitter) account with optional media. Text can include mentions, hashtags, and URLs. Media files are uploaded via v2 chunked upload. Returns tweet_id and URL.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Post text (required). X enforces per-account character limits. | |
| account_id | No | X account: numeric channel account ID or @handle (optional; defaults to workspace's sole X account). | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| media_file_ids | No | DialogBrain file IDs to attach (optional, max 4 per post). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false and openWorldHint=true, so the safety profile is covered. The description adds genuinely new behavior: the v2 chunked upload path for media, the fact that text supports mentions/hashtags/URLs, and the return payload (tweet_id and URL). It does not explain irreversibility or failure modes, but it goes beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and scoped by media. The return-value sentence earns its place because there is no output schema. Minor redundancy in listing content types, but no real waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly states what comes back (tweet_id and URL), and the media upload mechanism is disclosed. Parameters are fully covered by the schema and safety hints by annotations. It stops short of stating auth/permission requirements or failure behavior for a public, destructive publish.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented, including the media cap and account_id defaulting. The description only adds 'Text can include mentions, hashtags, and URLs,' which is minor. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+target ('Publish a post to the connected X (Twitter) account') and immediately scopes it with 'with optional media.' This cleanly separates it from siblings like instagram_publish_media or threads_channel_publish_post without the agent needing to open a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance and no named alternative. The description implies the precondition of a 'connected X account' but never says what to do if none is connected, nor when to prefer this over other publishing tools. Only implicit usage context is present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_delete_commentADestructiveIdempotentInspect
Permanently delete a YouTube comment by id (or 'youtube:comment:'). Cannot be undone. Costs 50 quota units.
| Name | Required | Description | Default |
|---|---|---|---|
| comment_id | Yes | Bare commentId OR 'youtube:comment:<id>'. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the irreversibility signal partly overlaps. The description still adds real value beyond them: explicit 'Cannot be undone' confirmation and the 50-quota-unit cost, which no annotation conveys. It does not describe failure modes or what the response returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short fragments, front-loaded with the action, then the irreversibility warning, then the cost. Each sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter destructive tool with no output schema, the description covers the essential agent decisions: what it does, that it is irreversible, and its quota cost. It stops short of noting permission requirements or what happens when the comment id is invalid.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already fully documented in the schema, and the description's id-format note merely repeats the comment_id schema text. Nothing is said about in_workspace, but the schema handles it. Baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Permanently delete a YouTube comment'), plus the accepted id forms. There is no sibling differentiator against near-neighbours such as youtube_moderate_comment, which could also remove/hide a comment, so it stops short of the 5 tier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the verb; the description adds no explicit 'use this when' or 'use X instead' guidance. The quota note ('Costs 50 quota units') is useful selection context but is not an alternative-selection rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_delete_videoADestructiveIdempotentInspect
Permanently delete a YouTube video by id (or 'youtube:video:'). Cannot be undone. Costs 50 quota units. Caller must own the channel.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | Bare videoId OR 'youtube:video:<id>'. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so safety is covered. The description adds genuine behavioral context beyond that: irreversibility ('Cannot be undone'), a quota cost ('50 quota units') for budgeting, and an ownership/auth requirement — none of which appear in the annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the action and scope, then appends irreversibility, cost, and auth as short clauses. Every clause carries operational value with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive two-parameter tool with no output schema, the description covers the key decision factors: what is destroyed, that it is permanent, quota cost, and permission requirements. It does not explain the return value or error behavior on a failed deletion, but nothing critical to invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both video_id and in_workspace are already self-documenting; baseline is 3. The description echoes the accepted 'youtube:video:<id>' format but adds no syntax or constraint detail beyond the schema, and says nothing about the in_workspace parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Permanently delete') and resource ('a YouTube video by id'), with the permanent/irreversible scope made explicit. It is clearly distinguishable from siblings like youtube_update_video and youtube_upload_video without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides useful preconditions ('Caller must own the channel') and warns of irreversibility, which helps the agent decide whether to proceed. However, it names no alternative (e.g., use youtube_update_video to privatize instead of deleting, or youtube_video_query to confirm the target first), so the when-vs-which guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_list_commentsBRead-onlyIdempotentInspect
List comment threads on a YouTube video. Pass video_id (e.g. 'dQw4w9WgXcQ') or channel_ref ('youtube:video:'). Returns top-level comments with inline replies.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | YouTube videoId — bare 11-char form OR full 'youtube:video:<id>'. | |
| page_token | No | Pagination cursor from a previous call's `next_page_token`. | |
| max_results | No | Page size, 1-100. Default 25. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this a read-only, idempotent, non-destructive operation, so the safety profile is covered without the description. The description usefully adds the return shape ('top-level comments with inline replies'), which matters since no output schema exists, but says nothing about pagination limits, auth requirements, or rate behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero padding, with the core purpose and the return behavior front-loaded. It could be slightly tighter given the identifier formatting is redundant with the schema, but it is well-sized and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with full schema coverage and no output schema, the definition is adequate: purpose, identifier format, and return shape are all present. It stops short of completeness because pagination behavior and the total absence of usage context leave an agent to infer how and when to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented in the schema, making 3 the correct baseline. The description's mention of the video_id example and the 'youtube:video:<id>' form only restates what the schema already says, adding no new semantics for page_token, max_results, or in_workspace.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('List comment threads on a YouTube video') that is unambiguous and clearly distinct from write-side siblings like youtube_post_comment_reply or youtube_moderate_comment. It does not, however, explicitly contrast itself with the closest read-side siblings such as youtube_video_query or youtube_list_videos, so it stops at clear-but-undifferentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains how to format the video identifier but gives no when-to-use guidance, no conditions, and no routing to alternatives (e.g., youtube_video_query for metadata vs this tool for comments). An agent gets no help deciding which sibling to reach for.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_list_videosARead-onlyIdempotentInspect
List videos on the connected YouTube channel. Returns id, title, published_at, view_count. Paginate via page_token.
| Name | Required | Description | Default |
|---|---|---|---|
| page_token | No | Pagination cursor returned in a previous call's `next_page_token`. Omit for the first page. | |
| max_results | No | Page size, 1-50. Default 25. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so the safety profile is covered. The description adds genuinely new behavioral context by naming the returned fields (id, title, published_at, view_count) and the pagination mechanism, which matters because no output schema exists. It stops short of stating default page size, ordering, or result caps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and followed by return shape and pagination. Every sentence carries information the agent needs; there is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, zero-required-parameter list tool, the description covers what it does, what comes back, and how to page, which compensates for the absence of an output schema. Missing only secondary detail such as result ordering and the default page size behavior at the boundary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so page_token, max_results and in_workspace are already fully documented in the schema, including the default of 25 and the 1-50 range. The description only repeats the pagination concept without adding format or edge-case detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List videos on the connected YouTube channel'), which is clearly distinguishable from mutation siblings like youtube_upload_video or youtube_delete_video. It does not, however, distinguish itself from the close sibling youtube_video_query or tiktok_list_videos, leaving the agent to infer which lister applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the agent can infer this is the browse/list operation for the channel's videos, and the pagination sentence hints at multi-page iteration. There is no explicit when-to-use or when-not-to-use guidance, and no pointer to youtube_video_query for targeted lookups or to tiktok_list_videos for the other platform.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_moderate_commentAInspect
Apply a moderation status to a YouTube comment. Allowed status values: heldForReview, published, rejected, spam. Costs 50 quota units.
| Name | Required | Description | Default |
|---|---|---|---|
| status | Yes | One of: heldForReview, published, rejected, spam. | |
| comment_id | Yes | Bare commentId OR 'youtube:comment:<id>'. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the mutation profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=false), so the description's burden is lower. It earns credit by adding a genuinely useful operational detail not in the structured data: the 50-unit quota cost. It does not state reversibility, permission scope, or when a moderated comment becomes visible, but the annotation coverage makes that a modest gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the operation front-loaded and the two most decision-relevant facts (valid statuses, quota cost) packed in. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-destructive mutation with full schema coverage, annotations covering the safety profile, and no output schema to explain, the description covers the essentials plus cost. It is only missing permission requirements and the effect of each status, which is a minor omission at this complexity level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both required fields are self-documented, so the baseline is 3. The description's list of status values merely restates the schema's status description, adding no new meaning (no ordering, side effects, or default behavior for heldForReview).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (apply a moderation status) and resource (a YouTube comment), which is enough to distinguish it from youtube_delete_comment, youtube_list_comments, and youtube_post_comment_reply. It stops short of explicitly naming those siblings, so an agent must infer the boundary rather than being told it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description enumerates allowed status values but gives no guidance on when to choose heldForReview versus rejected versus spam, no prerequisites, and no pointer to related tools like the delete or reply endpoints. Usage must be inferred from the verb alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_post_comment_replyADestructiveInspect
Post a comment on a YouTube video, or reply to an existing comment. Pass video_id for a top-level comment, OR parent_comment_id to reply. AI-disclosure suffix appended automatically when configured.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Comment body. 1-10000 chars. AI-disclosure suffix may be auto-appended. | |
| video_id | No | Bare videoId or 'youtube:video:<id>' — for a top-level comment. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| parent_comment_id | No | Bare commentId or 'youtube:comment:<id>' — for a reply. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the full safety profile (readOnlyHint=false, destructiveHint=true, idempotentHint=false, openWorldHint=true), so the bar is lower. The description still adds non-obvious behavioral context beyond annotations: an AI-disclosure suffix is auto-appended 'when configured', meaning the submitted text may be modified server-side. It does not mention auth requirements, rate limits, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, and the core action is front-loaded. Each sentence carries distinct information: the action, the mode routing, and the side effect on the text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation tool with no output schema and rich annotations, the description covers the two invocation paths and the content-mutation side effect adequately. It leaves unstated whether both id fields can be supplied together, what the returned comment looks like, and any posting limits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds real value on top of the schema by stating the mutual-exclusivity/OR relationship between video_id and parent_comment_id, which the per-property schema descriptions state only in isolation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Post a comment on a YouTube video, or reply to an existing comment') and cleanly separates the two operating modes. It does not explicitly name the sibling it is not (e.g., youtube_moderate_comment or youtube_list_comments), but the post-vs-read/moderate distinction is unambiguous from the wording.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit branching guidance for invocation: pass video_id for a top-level comment OR parent_comment_id to reply. What is missing is any when-not-to-use guidance or comparison against sibling tools like youtube_moderate_comment or youtube_list_comments, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_update_videoADestructiveInspect
Update title, description, privacy, or tags on a YouTube video. Costs 1600 quota units. Only fields provided are changed.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | New tags list. Omit to keep current. | |
| title | No | New title (max 100 chars). Omit to keep current. | |
| privacy | No | 'private', 'unlisted', or 'public'. Omit to keep current. | |
| video_id | Yes | Bare videoId OR 'youtube:video:<id>'. | |
| description | No | New description (max 5000 chars). Omit to keep current. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (destructiveHint=true, idempotentHint=false, openWorldHint=true), so the bar is lower. The description still adds non-annotation value: a concrete quota cost (1600 units) and patch semantics ('Only fields provided are changed'), which tells the agent this is not a full-replace operation. It omits auth scope and whether prior values can be restored.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the operation and fields, followed by cost and patch semantics. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param mutation tool with no output schema, the description plus annotations cover the essentials: what changes, that it is a patch, and the quota price. Missing only peripheral details such as required permission scope or confirmation/rollback behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each property already documents 'Omit to keep current', so the description's field list largely duplicates the schema. The one added nuance, that only supplied fields change, is genuinely parameter-relevant, but it does not compensate beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (YouTube video) plus the exact mutable fields (title, description, privacy, tags). This clearly separates it from siblings like youtube_delete_video, youtube_upload_video, and youtube_video_query without needing to open any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the verb and the patch note, but the description never states when to prefer this over alternatives (e.g. youtube_upload_video for new content) or any prerequisites such as ownership/auth. The quota-cost line is useful decision input but is not a when-to-use rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_upload_videoADestructiveInspect
Upload a workspace-owned video file (file_id) to the connected YouTube channel. Returns video_id + thread_id. Costs 1600 quota units. Default privacy is 'private' — pass privacy='public' to publish.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Optional list of tag strings (max ~500 chars total). | |
| title | Yes | Video title (max 100 chars). | |
| file_id | Yes | Workspace `files.id` of the video to upload. Must be a video/* MIME type and `status='ready'`. Get IDs from the [ATTACHMENTS] block, files.search, or search.files. | |
| privacy | No | Privacy status. 'private' (default), 'unlisted', or 'public'. | private |
| category_id | No | YouTube category ID (default '22' = People & Blogs). See https://developers.google.com/youtube/v3/docs/videoCategories/list. | 22 |
| description | No | Video description (max 5000 chars). OMIT to upload without a description. | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. | |
| made_for_kids | No | COPPA flag. OMIT for the standard (non-kids) default. | |
| channel_account_id | No | The connected YouTube channel_account.id. OMIT to auto-resolve the workspace's YouTube account. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (non-readOnly, open-world, non-idempotent, destructive). The description adds genuine value beyond them: quota cost (1600 units), the default privacy behavior, and the return shape (video_id + thread_id). It does not disclose that re-uploading creates a new video rather than updating, but the added operational context is substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four terse sentences, front-loaded with purpose, then returns, then cost, then the critical privacy default. Every sentence carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter mutation tool with no output schema, the description supplies the return values, the quota cost, and the privacy default. Combined with full schema coverage and informative annotations, an agent has enough to call it correctly; only the lack of explicit alternative-tool routing keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all nine parameters, including file_id requirements, title/description limits, and category defaults. The description only restates the privacy default, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: upload a workspace-owned video file to the connected YouTube channel. The scoping ('workspace-owned', 'connected channel') clearly distinguishes it from siblings like youtube_update_video, youtube_list_videos, and tiktok_publish_video without requiring schema inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives actionable operational context: it costs 1600 quota units and defaults to private, with the explicit instruction to 'pass privacy=public to publish'. It lacks explicit when-not guidance or named alternatives (e.g., vs youtube_update_video), but the context is clear enough for correct usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_video_queryARead-onlyIdempotentInspect
Ask Gemini about a YouTube video. Pass a video URL and any prompt — verbatim transcript with timestamps, summary, targeted Q&A about content or visuals, translation, etc. Works on any public/unlisted video.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | YouTube video URL. Supported forms: youtube.com/watch?v=…, youtu.be/…, youtube.com/shorts/…, m.youtube.com/watch?v=…. Pass-through to Gemini verbatim. | |
| prompt | Yes | What to ask Gemini about the video. Examples: 'Provide a verbatim transcript with [HH:MM:SS] timestamps.' / 'What is the main claim made in the first 30 seconds?' / 'Describe what's shown on screen at 0:30.' / 'Translate the spoken Spanish to English.' | |
| in_workspace | No | Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly, idempotent, and non-destructive behavior, so the safety profile is covered. The description adds genuinely useful scope information — it works on public/unlisted videos (implying private videos are unsupported) and that the URL is passed through verbatim to Gemini, hinting at a delegated external call. Latency, quota, or failure modes are not disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, tightly written. The core action is front-loaded and the capability enumeration is compact and earns its place by telling the agent what kinds of prompts are viable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Read-only tool with no output schema and full parameter coverage; the description supplies the essential capability and access-scope information an agent needs. It could be slightly richer about response shape or limitations, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters — including the URL formats, prompt examples, and the in_workspace scoping note — are already documented in the schema. The description offers 'any prompt' with the same example categories the schema lists, so it adds no meaning beyond the structured fields. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (ask Gemini about) and resource (a YouTube video), then enumerates the concrete output kinds — verbatim transcript with timestamps, summary, Q&A about content or visuals, translation. This is clearly distinct from the other youtube_* siblings, which are upload/manage/comment tools, not content-query tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context via the prompt examples (transcript, summary, targeted Q&A, visual description, translation) and the scope 'works on any public/unlisted video'. It does not name an alternative tool or state when not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- Changed
tasks_create1 field changed- changed
Input schema / properties / assignee / descriptionPrevious value: -"Who should do this: a workspace member (email or username), an AI agent (its name), or a contact (display name). Use 'me' for yourself. Call workspace.members to see who can be assigned. For an exact target, pass assignee_user_id / assignee_agent_id / assigned_to_contact_id instead."New value: +"Who should do this: a workspace member (email, username, or user:<id> as workspace.members lists them), an AI agent (its name), or a contact (display name). Use 'me' for yourself. Call workspace.members to see who can be assigned. Where this listing offers assignee_user_id / assignee_agent_id / assigned_to_contact_id, they name an exact target instead."
- Changed
tasks_update1 field changed- changed
Input schema / properties / assignee / descriptionPrevious value: -"Who should do this: a workspace member (email or username), an AI agent (its name), or a contact (display name). Use 'me' for yourself. Call workspace.members to see who can be assigned. For an exact target, pass assignee_user_id / assignee_agent_id / assigned_to_contact_id instead."New value: +"Who should do this: a workspace member (email, username, or user:<id> as workspace.members lists them), an AI agent (its name), or a contact (display name). Use 'me' for yourself. Call workspace.members to see who can be assigned. Where this listing offers assignee_user_id / assignee_agent_id / assigned_to_contact_id, they name an exact target instead."
1 tool update
- Changed
messages_read_history1 field changed- added
Input schema / properties / channel_account_idAdded value: +{ + "description": "Connected account the conversation must belong to, when you know it (from channels.list). Scopes the lookup to that account, so the same person reached on two connected accounts cannot be confused for one, and a chat that does not exist on this account is 'not found' instead of someone else's thread with a matching id.", + "type": "string" +}
1 tool update
- Changed
calls_make3 fields changed- changed
Input schema / properties / channel / descriptionPrevious value: -"Voice transport: 'twilio' or 'telnyx' (phone via PSTN — both require phone_number in E.164; pick the carrier the workspace has connected), 'telegram' (MTProto 1:1 call — requires telegram_user_id, NOT a phone number or thread_id), 'maxru' (Max.ru voice call — requires maxru_user_id), 'android' (Android device voice call — requires phone_number in E.164), 'whatsapp' (WhatsApp voice call via the workspace's connected WhatsApp account — requires phone_number in E.164). OMIT to auto-select based on the current thread (e.g. inside a Telegram DM → uses 'telegram')."New value: +"Voice transport: 'twilio' or 'telnyx' (phone via PSTN — both require phone_number in E.164; pick the carrier the workspace has connected), 'telegram' (MTProto 1:1 call — requires telegram_user_id, NOT a phone number or thread_id), 'maxru' (Max.ru voice call — requires maxru_user_id), 'android' (Android device voice call — requires phone_number in E.164), 'whatsapp' (WhatsApp voice call via the workspace's connected WhatsApp account — requires phone_number in E.164), 'whatsapp_business' (a Meta Cloud API number; requires phone_number in E.164; the contact must have allowed calls, otherwise a permission request is sent instead). OMIT to auto-select based on the current thread (e.g. inside a Telegram DM → uses 'telegram')." - changed
Input schema / properties / channel / enumPrevious value: -[ - "twilio", - "telnyx", - "telegram", - "maxru", - "android", - "whatsapp", - "viber" -]New value: +[ + "twilio", + "telnyx", + "telegram", + "maxru", + "android", + "whatsapp", + "whatsapp_business", + "viber" +] - added
Input schema / properties / permission_request_textAdded value: +{ + "description": "Only for channel='whatsapp_business': the text of the call-permission request WhatsApp sends when the contact has not yet allowed calls. Omit for a generic default.", + "maxLength": 1024, + "type": "string" +}
1 tool update
- Changed
messages_send1 field changed- added
Input schema / properties / templateAdded value: +{ + "description": "WhatsApp Business API only. An approved template to send instead of free text, in Meta's shape: {name, language, components:[{type:'body', parameters:[{type:'text', text:...}]}, ...]}; `language` may be the code as a string ('en_US') or Meta's {code:'en_US'}. Required when the 24-hour window of the conversation is closed (the send answers WINDOW_EXPIRED otherwise). Read the approved templates and their variables with channels.list(account_id=...).", + "type": "object" +}
1 tool update
- Changed
instagram_publish_media1 field changed- added
Input schema / properties / alt_textAdded value: +{ + "description": "Alternative text for a photo (accessibility; read by screen readers), up to 1000 characters. Photos only: Instagram ignores it for Reels.", + "type": "string" +}
3 tool updates
- Changed
instagram_list_media1 field changed- added
Input schema / properties / account_idAdded value: +{ + "description": "channel_account id of the Instagram account to use when the workspace has several (see channels.list). OMIT to use the workspace's connected Instagram account.", + "type": "integer" +}
- Changed
instagram_publish_media1 field changed- added
Input schema / properties / account_idAdded value: +{ + "description": "channel_account id of the Instagram account to use when the workspace has several (see channels.list). OMIT to use the workspace's connected Instagram account.", + "type": "integer" +}
- Changed
instagram_update_media1 field changed- added
Input schema / properties / account_idAdded value: +{ + "description": "channel_account id of the Instagram account to use when the workspace has several (see channels.list). OMIT to use the workspace's connected Instagram account.", + "type": "integer" +}
3 tool updates
- Added
channels_contacts - Changed
channels_list1 field changed- added
Input schema / properties / account_idAdded value: +{ + "description": "One account in detail. For a custom channel this adds whether it is listening or polling, why inbound is down, how many contacts it has, the progress of a history import, a health check and the agents that would answer on it. OMIT to list the workspace's channels.", + "type": "integer" +}
- Added
integrations_set_realtime
274 tool updates
- Changed
agent_handoff1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agent_silence1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_activity1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_add_file1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_approve_draft1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_ask1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_delete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_get1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_list_drafts1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_list_files1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_list_integrations1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_prompt_history1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_prompt_restore1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_reject_draft1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_remove_file1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_set_integration1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_simulate_inbound1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_task_complete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_trace_get1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_traces_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_traces_stats1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_trigger_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_trigger_delete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_trigger_update1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_update1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
agents_update_from_template1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
ai_filters_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
ai_filters_delete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
ai_filters_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
ai_filters_test1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
ai_filters_update1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
ai_tags_add_to_thread1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
ai_tags_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
ai_tags_delete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
ai_tags_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
ai_tags_remove_from_thread1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
ai_tags_update1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
analytics_query1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
android_close_app1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
android_current_app1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
android_install_apk1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
android_key1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
android_launch_app1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
android_list_apps1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
android_screenshot1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
android_shell1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
android_swipe1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
android_tap1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
android_type1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
android_ui_dump1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
artifacts_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
artifacts_export_pdf1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
artifacts_get1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
artifacts_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
artifacts_refresh1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
artifacts_set_access1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
artifacts_update1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
background_run1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_add_init_script1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_attach_identity1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_attach_meet1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_click1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_close1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_console_messages1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_drag1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_evaluate1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_file_upload1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_fill1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_fill_form1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_handle_dialog1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_hover1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_navigate_back1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_network_requests1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_open1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_press_key1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_resize1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_select_option1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_snapshot1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_tabs1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_take_screenshot1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_type1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
browser_wait_for1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calendar_check_availability1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calendar_create_event1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calendar_delete_event1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calendar_list_events1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calendar_update_event1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_agent_duel1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_dispatch_agent1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_dispatch_translator1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_get_transcript1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_hangup1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_list_active1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_list_history1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_make1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_meet_browser1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_mute_translation_tts1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_set_translation_language1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_set_translation_languages1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_transfer1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
calls_wait1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
campaigns_cancel1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
campaigns_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
campaigns_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
campaigns_pause1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
campaigns_resume1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
campaigns_status1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
channels_connect1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
channels_connect_telegram_bot1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
channels_get_profile1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
channels_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
channels_update_profile1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
collections_add_file1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
collections_add_website1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
collections_assign_agent1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
collections_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
collections_delete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
collections_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
collections_list_files1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
collections_list_websites1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
collections_refresh_website1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
collections_remove_file1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
collections_remove_website1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
collections_unassign_agent1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
contacts_add_channel1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
contacts_capture_lead1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
contacts_discover1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
contacts_find1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
contacts_merge1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
contacts_profile1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
contacts_research1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
contacts_sync1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
contacts_update1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
db_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
db_execute1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
db_query1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
db_schema1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
documents_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
feedback_save1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
files_complete_upload1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
files_create_upload_url1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
files_delete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
files_get_base641 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
files_info1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
files_ingest1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
files_read1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
files_upload1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
folders_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
folders_delete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
group_add1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
group_add_member1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
group_admins1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
group_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
group_get_invite_link1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
group_join1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
group_leave1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
group_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
group_moderate1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
group_preview_messages1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
group_promote_admin1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
group_scan1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
group_search1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
images_generate1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
images_search1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
instagram_list_media1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
instagram_publish_media1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
instagram_update_media1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
integrations_add_endpoints1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
integrations_capture_session1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
integrations_execute_tool1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
integrations_get_endpoints1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
integrations_get_tool_schema1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
integrations_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
integrations_remove_endpoints1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
integrations_search_tools1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
integrations_set_auth1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
integrations_sync_knowledge1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
job_complete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
job_escalate1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
job_read_context1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
job_update_context1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
kg_find_entity1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
kg_get_relationships1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
knowledge_gaps_answer1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
knowledge_gaps_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
knowledge_query1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
linkedin_raw_request1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
linkedin_update_profile1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
messages_delete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
messages_edit1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
messages_forward1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
messages_press_button1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
messages_read_history1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
messages_send1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
messages_send_email1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
notes_delete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
notes_recall1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
notes_save1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
notes_search1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
onboarding_start_signup1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
present_status1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
present_stop1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
present_tab1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
prompts_blocks_get1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
prompts_blocks_preview1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
prompts_blocks_registry1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
prompts_blocks_update1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
prompts_get1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
prompts_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
prompts_prompt_history1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
prompts_prompt_restore1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
prompts_update1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
reminder_cancel1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
reminder_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
reminder_set1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
search_files1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
search_links1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
search_messages1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
search_threads1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
system_sleep1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
tasks_comment1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
tasks_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
tasks_delete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
tasks_get1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
tasks_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
tasks_update1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
threads_channel_account_stats1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
threads_channel_hide_reply1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
threads_channel_list_posts1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
threads_channel_post_insights1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
threads_channel_publish_post1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
threads_channel_unhide_reply1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
threads_delete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
threads_update1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
tiktok_account_stats1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
tiktok_list_videos1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
tiktok_publish_video1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
videos_generate1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
vision_query1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
web__local_search1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
web_fetch1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
web_research1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
web_search1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
webhooks_configure_enrichment1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
widgets_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
widgets_delete1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
widgets_get1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
widgets_get_embed_code1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
widgets_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
widgets_update1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
workbench_run_python1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
workspace_create1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
workspace_current1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
workspace_desktops1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
workspace_invite1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
workspace_invite_revoke1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
workspace_invites1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
workspace_list1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
workspace_member_remove1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
workspace_member_update1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
workspace_members1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
workspace_update1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
x_publish_post1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
youtube_delete_comment1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
youtube_delete_video1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
youtube_list_comments1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
youtube_list_videos1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
youtube_moderate_comment1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
youtube_post_comment_reply1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
youtube_update_video1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
youtube_upload_video1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
- Changed
youtube_video_query1 field changed- added
Input schema / properties / in_workspaceAdded value: +{ + "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.", + "type": "integer" +}
2 tool updates
- Changed
agents_create3 fields changed- added
Input schema / properties / remote_toolAdded value: +{ + "description": "Only for text_engine='external_agent': the workspace integration tool that starts the remote agent, e.g. 'ext42_run_routine' (list them with integrations.search_tools). On each trigger DialogBrain calls it once with the event, the agent's instructions and the ids to answer with; the remote agent replies through the DialogBrain MCP tools (messages.send, tasks.comment, agents.task_complete). The endpoint and its secret belong to the integration, not to the agent.", + "type": "string" +} - changed
Input schema / properties / text_engine / descriptionPrevious value: -"Text-execution engine: 'rule_based', 'ai_assisted', 'agentic' (default), or 'claude_channels'. Voice is derived from triggers, not engine. OMIT to use the default ('agentic')."New value: +"Text-execution engine: 'rule_based', 'ai_assisted', 'agentic' (default), 'claude_channels', or 'external_agent' (the run is handed to an agent outside DialogBrain; needs remote_tool). Voice is derived from triggers, not engine. OMIT to use the default ('agentic')." - changed
Input schema / properties / text_engine / enumPrevious value: -[ - "rule_based", - "agentic", - "claude_channels" -]New value: +[ + "rule_based", + "agentic", + "claude_channels", + "external_agent" +]
- Changed
agents_update2 fields changed- added
Input schema / properties / remote_toolAdded value: +{ + "description": "Only for text_engine='external_agent': the workspace integration tool that starts the remote agent, e.g. 'ext42_run_routine' (list them with integrations.search_tools). On each trigger DialogBrain calls it once with the event, the agent's instructions and the ids to answer with; the remote agent replies through the DialogBrain MCP tools (messages.send, tasks.comment, agents.task_complete). The endpoint and its secret belong to the integration, not to the agent.", + "type": "string" +} - changed
Input schema / properties / text_engine / enumPrevious value: -[ - "rule_based", - "agentic", - "claude_channels" -]New value: +[ + "rule_based", + "agentic", + "claude_channels", + "external_agent" +]
1 tool update
- Changed
messages_send1 field changed- changed
Input schema / properties / thread_id / descriptionPrevious value: -"Target thread. OMIT to reply in the same chat you received the triggering message from — the backend defaults to the current thread. Pass an explicit value ONLY to reply in a DIFFERENT thread, and only use: (a) a numeric DB thread id from search.threads, or (b) a channel_ref like 'telegram:-12345'. To text a phone number that has no thread yet, pass 'sms:+<number in international format>' (the workspace needs a phone number with SMS turned on; from_account_id picks which one). NEVER use a chat-type word (dm, group, channel, livechat) — those are category labels from the SITUATION block, not ids."New value: +"Target thread. OMIT to reply in the same chat you received the triggering message from — the backend defaults to the current thread. Pass an explicit value ONLY to reply in a DIFFERENT thread, and only use: (a) a numeric DB thread id from search.threads, or (b) a channel_ref like 'telegram:-12345'. To text a phone number, pass 'sms:+<number in international format>': the message goes out from a workspace phone number that has SMS turned on (from_account_id picks which one) and lands in that number's conversation with the recipient, next to their calls. NEVER use a chat-type word (dm, group, channel, livechat) — those are category labels from the SITUATION block, not ids."
1 tool update
- Changed
messages_send1 field changed- changed
Input schema / properties / thread_id / descriptionPrevious value: -"Target thread. OMIT to reply in the same chat you received the triggering message from — the backend defaults to the current thread. Pass an explicit value ONLY to reply in a DIFFERENT thread, and only use: (a) a numeric DB thread id from search.threads, or (b) a channel_ref like 'telegram:-12345'. NEVER use a chat-type word (dm, group, channel, livechat) — those are category labels from the SITUATION block, not ids."New value: +"Target thread. OMIT to reply in the same chat you received the triggering message from — the backend defaults to the current thread. Pass an explicit value ONLY to reply in a DIFFERENT thread, and only use: (a) a numeric DB thread id from search.threads, or (b) a channel_ref like 'telegram:-12345'. To text a phone number that has no thread yet, pass 'sms:+<number in international format>' (the workspace needs a phone number with SMS turned on; from_account_id picks which one). NEVER use a chat-type word (dm, group, channel, livechat) — those are category labels from the SITUATION block, not ids."
1 tool update
- Changed
integrations_add_endpoints1 field changed- changed
Input schema / properties / name / descriptionPrevious value: -"Display name when creating a new integration. OMIT to default to the host."New value: +"Display name. OMIT to keep an existing integration's name, or to name a NEW one after its host."
2 tool updates
- Changed
calendar_check_availability1 field changed- added
Input schema / properties / timezoneAdded value: +{ + "description": "IANA timezone (e.g. 'Asia/Bishkek') that start_time/end_time WITHOUT an offset are in, and that the free slots are reported in. OMIT for UTC.", + "type": "string" +}
- Changed
reminder_set1 field changed- changed
Input schema / properties / datetime / descriptionPrevious value: -"ISO datetime for one_time (e.g. '2026-04-01T09:00:00+03:00'). Required for one_time."New value: +"ISO datetime for one_time. Required for one_time. Either with an offset ('2026-04-01T09:00:00+03:00') or without one plus `timezone` ('2026-04-01T09:00:00' + timezone 'Europe/Moscow'): a datetime with no offset is read in `timezone`, and in UTC only when that is omitted."
1 tool update
- Changed
agents_update2 fields changed- changed
Input schema / properties / voice_realtime_model / descriptionPrevious value: -"Realtime (v2v) model tier for voice_engine=openai_realtime: 'gpt-realtime' (~18c/min on short calls) or 'gpt-realtime-mini' (~75% cheaper). OMIT to keep the plugin default."New value: +"Realtime (v2v) model for the agent's voice_engine. openai_realtime: 'gpt-realtime' (~18c/min on short calls) or 'gpt-realtime-mini' (~75% cheaper). gemini_realtime: a Gemini Live model, newest first 'gemini-3.8-live', 'gemini-3.1-flash-live-preview', 'gemini-2.5-flash-native-audio-latest', 'gemini-2.5-flash-native-audio-preview-12-2025' (today's default). 'default' = back to the engine's default. OMIT to leave unchanged." - changed
Input schema / properties / voice_realtime_model / enumPrevious value: -[ - "gpt-realtime", - "gpt-realtime-mini" -]New value: +[ + "gpt-realtime", + "gpt-realtime-mini", + "gemini-3.8-live", + "gemini-3.1-flash-live-preview", + "gemini-2.5-flash-native-audio-latest", + "gemini-2.5-flash-native-audio-preview-12-2025", + "default" +]
1 tool update
- Changed
agents_update1 field changed- changed
Input schema / properties / voice_engine / descriptionPrevious value: -"Voice execution engine: 'pipeline' (default — Deepgram/Gladia STT + LLM + TTS), 'openai_realtime' (OpenAI Realtime API v2v; requires the workspace to have a BYOK OpenAI key connected — the worker falls back to 'pipeline' and logs why if the key is missing), or 'gemini_realtime' (Gemini 2.0 Flash with real-time API). OMIT to leave unchanged."New value: +"Voice execution engine: 'pipeline' (default — Deepgram/Gladia STT + LLM + TTS), 'openai_realtime' (OpenAI Realtime API v2v; requires the workspace to have a BYOK OpenAI key connected — the worker falls back to 'pipeline' and logs why if the key is missing), or 'gemini_realtime' (Gemini Live native-audio v2v; runs on the workspace owner's Gemini key, else on the platform key unless platform keys are switched off). If the chosen realtime engine has no usable key, the result carries `warnings` and calls use the default pipeline. OMIT to leave unchanged."
1 tool update
- Changed
reminder_list1 field changed- changed
Input schema / properties / agent_id / descriptionPrevious value: -"Agent ID (required when calling from MCP; ignored in agentic mode)."New value: +"Only this agent's reminders. Leave it out to list the whole workspace's; each row then names the agent it belongs to. Ignored in agentic mode."
275 tool updates
- Added
agent_handoff - Added
agent_silence - Added
agents_activity - Added
agents_add_file - Added
agents_approve_draft - Added
agents_ask - Added
agents_create - Added
agents_delete - Added
agents_get - Added
agents_list - Added
agents_list_drafts - Added
agents_list_files - Added
agents_list_integrations - Added
agents_prompt_history - Added
agents_prompt_restore - Added
agents_reject_draft - Added
agents_remove_file - Added
agents_set_integration - Added
agents_simulate_inbound - Added
agents_task_complete - Added
agents_trace_get - Added
agents_traces_list - Added
agents_traces_stats - Added
agents_trigger_create - Added
agents_trigger_delete - Added
agents_trigger_update - Added
agents_update - Added
agents_update_from_template - Added
ai_filters_create - Added
ai_filters_delete - Added
ai_filters_list - Added
ai_filters_test - Added
ai_filters_update - Added
ai_tags_add_to_thread - Added
ai_tags_create - Added
ai_tags_delete - Added
ai_tags_list - Added
ai_tags_remove_from_thread - Added
ai_tags_update - Added
analytics_query - Added
android_close_app - Added
android_current_app - Added
android_install_apk - Added
android_key - Added
android_launch_app - Added
android_list_apps - Added
android_screenshot - Added
android_shell - Added
android_swipe - Added
android_tap - Added
android_type - Added
android_ui_dump - Added
artifacts_create - Added
artifacts_export_pdf - Added
artifacts_get - Added
artifacts_list - Added
artifacts_refresh - Added
artifacts_set_access - Added
artifacts_update - Added
background_run - Added
browser_add_init_script - Added
browser_attach_identity - Added
browser_attach_meet - Added
browser_click - Added
browser_close - Added
browser_console_messages - Added
browser_drag - Added
browser_evaluate - Added
browser_file_upload - Added
browser_fill - Added
browser_fill_form - Added
browser_handle_dialog - Added
browser_hover - Added
browser_navigate_back - Added
browser_network_requests - Added
browser_open - Added
browser_press_key - Added
browser_resize - Added
browser_select_option - Added
browser_snapshot - Added
browser_tabs - Added
browser_take_screenshot - Added
browser_type - Added
browser_wait_for - Added
calendar_check_availability - Added
calendar_create_event - Added
calendar_delete_event - Added
calendar_list_events - Added
calendar_update_event - Added
calls_agent_duel - Added
calls_dispatch_agent - Added
calls_dispatch_translator - Added
calls_get_transcript - Added
calls_hangup - Added
calls_list_active - Added
calls_list_history - Added
calls_make - Added
calls_meet_browser - Added
calls_mute_translation_tts - Added
calls_set_translation_language - Added
calls_set_translation_languages - Added
calls_transfer - Added
calls_wait - Added
campaigns_cancel - Added
campaigns_create - Added
campaigns_list - Added
campaigns_pause - Added
campaigns_resume - Added
campaigns_status - Added
channels_connect - Added
channels_connect_telegram_bot - Added
channels_get_profile - Added
channels_list - Added
channels_update_profile - Added
collections_add_file - Added
collections_add_website - Added
collections_assign_agent - Added
collections_create - Added
collections_delete - Added
collections_list - Added
collections_list_files - Added
collections_list_websites - Added
collections_refresh_website - Added
collections_remove_file - Added
collections_remove_website - Added
collections_unassign_agent - Added
contacts_add_channel - Added
contacts_capture_lead - Added
contacts_discover - Added
contacts_find - Added
contacts_merge - Added
contacts_profile - Added
contacts_research - Added
contacts_sync - Added
contacts_update - Added
db_create - Added
db_execute - Added
db_query - Added
db_schema - Added
documents_create - Added
feedback_save - Added
files_complete_upload - Added
files_create_upload_url - Added
files_delete - Added
files_get_base64 - Added
files_info - Added
files_ingest - Added
files_read - Added
files_upload - Added
folders_create - Added
folders_delete - Added
group_add - Added
group_add_member - Added
group_admins - Added
group_create - Added
group_get_invite_link - Added
group_join - Added
group_leave - Added
group_list - Added
group_moderate - Added
group_preview_messages - Added
group_promote_admin - Added
group_scan - Added
group_search - Added
images_generate - Added
images_search - Added
instagram_list_media - Added
instagram_publish_media - Added
instagram_update_media - Added
integrations_add_endpoints - Added
integrations_capture_session - Added
integrations_execute_tool - Added
integrations_get_endpoints - Added
integrations_get_tool_schema - Added
integrations_list - Added
integrations_remove_endpoints - Added
integrations_search_tools - Added
integrations_set_auth - Added
integrations_sync_knowledge - Added
job_complete - Added
job_escalate - Added
job_read_context - Added
job_update_context - Added
kg_find_entity - Added
kg_get_relationships - Added
knowledge_gaps_answer - Added
knowledge_gaps_list - Added
knowledge_query - Added
linkedin_raw_request - Added
linkedin_update_profile - Added
messages_delete - Added
messages_edit - Added
messages_forward - Added
messages_press_button - Added
messages_read_history - Added
messages_send - Added
messages_send_email - Added
notes_delete - Added
notes_recall - Added
notes_save - Added
notes_search - Added
onboarding_start_signup - Added
present_status - Added
present_stop - Added
present_tab - Added
prompts_blocks_get - Added
prompts_blocks_preview - Added
prompts_blocks_registry - Added
prompts_blocks_update - Added
prompts_get - Added
prompts_list - Added
prompts_prompt_history - Added
prompts_prompt_restore - Added
prompts_update - Added
reminder_cancel - Added
reminder_list - Added
reminder_set - Added
search_files - Added
search_links - Added
search_messages - Added
search_threads - Added
system_sleep - Added
tasks_comment - Added
tasks_create - Added
tasks_delete - Added
tasks_get - Added
tasks_list - Added
tasks_update - Added
threads_channel_account_stats - Added
threads_channel_hide_reply - Added
threads_channel_list_posts - Added
threads_channel_post_insights - Added
threads_channel_publish_post - Added
threads_channel_unhide_reply - Added
threads_delete - Added
threads_update - Added
tiktok_account_stats - Added
tiktok_list_videos - Added
tiktok_publish_video - Added
videos_generate - Added
vision_query - Added
web__local_search - Added
web_fetch - Added
web_research - Added
web_search - Added
webhooks_configure_enrichment - Added
widgets_create - Added
widgets_delete - Added
widgets_get - Added
widgets_get_embed_code - Added
widgets_list - Added
widgets_update - Added
workbench_run_python - Added
workspace_create - Added
workspace_current - Added
workspace_desktops - Added
workspace_invite - Added
workspace_invite_revoke - Added
workspace_invites - Added
workspace_list - Added
workspace_member_remove - Added
workspace_member_update - Added
workspace_members - Added
workspace_switch - Added
workspace_update - Added
x_publish_post - Added
youtube_delete_comment - Added
youtube_delete_video - Added
youtube_list_comments - Added
youtube_list_videos - Added
youtube_moderate_comment - Added
youtube_post_comment_reply - Added
youtube_update_video - Added
youtube_upload_video - Added
youtube_video_query
232 tool updates
- Removed
agent_handoff - Removed
agent_silence - Removed
agents_activity - Removed
agents_add_file - Removed
agents_approve_draft - Removed
agents_ask - Removed
agents_create - Removed
agents_delete - Removed
agents_get - Removed
agents_list - Removed
agents_list_drafts - Removed
agents_list_files - Removed
agents_list_integrations - Removed
agents_prompt_history - Removed
agents_prompt_restore - Removed
agents_reject_draft - Removed
agents_remove_file - Removed
agents_set_integration - Removed
agents_simulate_inbound - Removed
agents_task_complete - Removed
agents_trace_get - Removed
agents_traces_list - Removed
agents_traces_stats - Removed
agents_trigger_create - Removed
agents_trigger_delete - Removed
agents_trigger_update - Removed
agents_update - Removed
agents_update_from_template - Removed
ai_filters_create - Removed
ai_filters_delete - Removed
ai_filters_list - Removed
ai_filters_test - Removed
ai_filters_update - Removed
ai_tags_add_to_thread - Removed
ai_tags_create - Removed
ai_tags_delete - Removed
ai_tags_list - Removed
ai_tags_remove_from_thread - Removed
ai_tags_update - Removed
analytics_query - Removed
android_close_app - Removed
android_current_app - Removed
android_install_apk - Removed
android_key - Removed
android_launch_app - Removed
android_list_apps - Removed
android_screenshot - Removed
android_shell - Removed
android_swipe - Removed
android_tap - Removed
android_type - Removed
android_ui_dump - Removed
artifacts_create - Removed
artifacts_export_pdf - Removed
artifacts_get - Removed
artifacts_list - Removed
artifacts_refresh - Removed
artifacts_set_access - Removed
artifacts_update - Removed
background_run - Removed
browser_attach_identity - Removed
browser_attach_meet - Removed
browser_click - Removed
browser_close - Removed
browser_console_messages - Removed
browser_drag - Removed
browser_evaluate - Removed
browser_file_upload - Removed
browser_fill - Removed
browser_fill_form - Removed
browser_handle_dialog - Removed
browser_hover - Removed
browser_navigate_back - Removed
browser_network_requests - Removed
browser_open - Removed
browser_press_key - Removed
browser_resize - Removed
browser_select_option - Removed
browser_snapshot - Removed
browser_tabs - Removed
browser_take_screenshot - Removed
browser_type - Removed
browser_wait_for - Removed
calendar_check_availability - Removed
calendar_create_event - Removed
calendar_delete_event - Removed
calendar_list_events - Removed
calendar_update_event - Removed
calls_agent_duel - Removed
calls_dispatch_translator - Removed
calls_get_transcript - Removed
calls_hangup - Removed
calls_list_active - Removed
calls_list_history - Removed
calls_make - Removed
calls_meet_browser - Removed
calls_mute_translation_tts - Removed
calls_send_to_meet - Removed
calls_send_to_telegram_call - Removed
calls_set_translation_language - Removed
calls_set_translation_languages - Removed
calls_wait - Removed
channels_connect_telegram_bot - Removed
collections_add_file - Removed
collections_assign_agent - Removed
collections_create - Removed
collections_delete - Removed
collections_list - Removed
collections_list_files - Removed
collections_remove_file - Removed
collections_unassign_agent - Removed
contacts_add_channel - Removed
contacts_capture_lead - Removed
contacts_discover - Removed
contacts_find - Removed
contacts_profile - Removed
contacts_research - Removed
contacts_sync - Removed
contacts_update - Removed
documents_create - Removed
feedback_save - Removed
files_delete - Removed
files_get_base64 - Removed
files_info - Removed
files_ingest - Removed
files_read - Removed
files_upload - Removed
folders_create - Removed
folders_delete - Removed
group_add - Removed
group_add_member - Removed
group_admins - Removed
group_create - Removed
group_get_invite_link - Removed
group_join - Removed
group_leave - Removed
group_list - Removed
group_preview_messages - Removed
group_promote_admin - Removed
group_scan - Removed
group_search - Removed
images_generate - Removed
images_search - Removed
instagram_list_media - Removed
instagram_publish_media - Removed
instagram_update_media - Removed
integrations_add_endpoints - Removed
integrations_capture_session - Removed
integrations_get_endpoints - Removed
integrations_list - Removed
integrations_remove_endpoints - Removed
integrations_set_auth - Removed
job_complete - Removed
job_escalate - Removed
job_read_context - Removed
job_update_context - Removed
kg_find_entity - Removed
kg_get_relationships - Removed
knowledge_query - Removed
linkedin_add_comment - Removed
linkedin_get_company - Removed
linkedin_get_profile - Removed
linkedin_invite - Removed
linkedin_list_connections - Removed
linkedin_list_invitations_sent - Removed
linkedin_list_reactions - Removed
linkedin_raw_request - Removed
linkedin_search - Removed
linkedin_search_filters - Removed
linkedin_update_profile - Removed
messages_delete - Removed
messages_edit - Removed
messages_forward - Removed
messages_read_history - Removed
messages_send - Removed
messages_send_email - Removed
notes_delete - Removed
notes_recall - Removed
notes_save - Removed
notes_search - Removed
owners_add - Removed
owners_list - Removed
owners_remove - Removed
present_status - Removed
present_tab - Removed
prompts_get - Removed
prompts_list - Removed
prompts_prompt_history - Removed
prompts_prompt_restore - Removed
prompts_update - Removed
reminder_cancel - Removed
reminder_list - Removed
reminder_set - Removed
search_files - Removed
search_links - Removed
search_messages - Removed
search_threads - Removed
system_sleep - Removed
tasks_create - Removed
tasks_delete - Removed
tasks_list - Removed
tasks_update - Removed
threads_delete - Removed
threads_update - Removed
videos_generate - Removed
vision_query - Removed
web__local_search - Removed
web_fetch - Removed
web_research - Removed
web_search - Removed
webhooks_configure_enrichment - Removed
widgets_create - Removed
widgets_delete - Removed
widgets_get - Removed
widgets_get_embed_code - Removed
widgets_list - Removed
widgets_update - Removed
workbench_run_python - Removed
workspace_current - Removed
workspace_list - Removed
workspace_switch - Removed
workspace_update - Removed
x_publish_post - Removed
youtube_delete_comment - Removed
youtube_delete_video - Removed
youtube_list_comments - Removed
youtube_list_videos - Removed
youtube_moderate_comment - Removed
youtube_post_comment_reply - Removed
youtube_update_video - Removed
youtube_upload_video - Removed
youtube_video_query
Related MCP Connectors
Unified messaging MCP server: WhatsApp, Instagram, Telegram, SMS, Messenger & email support inbox
Your customer inbox as tools: read and answer WhatsApp, Telegram, Instagram and e-mail.
Ask questions across your WhatsApp inbox from Claude, ChatGPT, Cursor or any MCP client.
Email, WhatsApp and Telegram for AI agents: send, campaigns, automations, contacts, agent inboxes.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceWhatsApp, Instagram, Telegram, SMS, Messenger & email in one AI-ready inbox. Manage conversations, send omnichannel messages, run campaigns, sync contacts, and query your knowledge base — 42 tools over a hosted, OAuth-secured MCP server.1MIT
- AlicenseBqualityCmaintenanceMCP server for Gmail shared inboxes and WhatsApp — read, reply, search, and organize your team's email from Claude, ChatGPT, or Cursor. 47 tools covering email threads, WhatsApp messaging, boards, contacts, knowledge base, analytics, and automations.4763 npm1MIT

SendAPI MCP Serverofficial
AlicenseAqualityDmaintenanceEnables any MCP-compatible AI agent to send WhatsApp messages, SMS, OTP codes, and email through a single REST API.18MIT- AlicenseNot gradedqualityAmaintenanceLets AI agents search and read your personal Telegram and WhatsApp conversations through a self-hosted, read-only MCP server backed by a single SQLite file.2AGPL 3.0
Glama MCP Gateway
Add one secure layer between your agents and this server.