Skip to main content
Glama
AIWerk

@aiwerk/mcp-server-elevenlabs

by AIWerk

@aiwerk/mcp-server-elevenlabs

An MCP server for the ElevenLabs API. 390 tools, generated from the live OpenAPI document, covering every documented endpoint.

Why this exists

ElevenLabs ships its own MCP server. It is hosted, it authenticates with OAuth, and it exposes about a dozen high-level tools for managing ElevenAgents plus one text to speech call. That is a good fit for "create a support agent and set its voice".

This one is for the rest of the API. The whole of it:

Tools

ElevenLabs REST API (OpenAPI, live)

390 operations

Official hosted MCP server

~12

Official local server (archived 2026-08)

27

This server

390

It also handles the two things a generated client usually gets wrong: 31 endpoints take file uploads (speech to text, voice cloning, dubbing, audio isolation, knowledge base) and 21 return raw bytes (all the text to speech variants). Parsing audio as JSON produces a plausible-looking string of mojibake, so those paths are written by hand and tested.

What the official hosted server does better: OAuth means no API key is copied into the client, and its agent tools are composed product logic rather than raw endpoints ("what would this agent cost per conversation on a different model" is not one API call). Use both if that is what you need. They do not conflict.

Related MCP server: ElevenLabs MCP Server

Install

npm install -g @aiwerk/mcp-server-elevenlabs

Or run it straight from npx in a client config:

{
  "mcpServers": {
    "elevenlabs": {
      "command": "npx",
      "args": ["-y", "@aiwerk/mcp-server-elevenlabs"],
      "env": {
        "ELEVENLABS_API_KEY": "your-key",
        "ELEVENLABS_OUTPUT_DIR": "/where/audio/should/land"
      }
    }
  }
}

Get a key at https://elevenlabs.io/app/settings/api-keys. The free tier includes 10k credits a month.

Keeping the key out of the config

An MCP client config is a plain file that tends to live in a repo or a dotfile, so a key written into its env block is a key in cleartext. If you keep secrets in a password manager, start the server through a small wrapper instead:

#!/usr/bin/env bash
set -euo pipefail
ELEVENLABS_API_KEY="$(pass show api/elevenlabs | head -1)"   # or your own manager
export ELEVENLABS_API_KEY
exec npx -y @aiwerk/mcp-server-elevenlabs@0.1.1 "$@"
{
  "mcpServers": {
    "elevenlabs": {
      "command": "/path/to/the/wrapper",
      "env": { "ELEVENLABS_OUTPUT_DIR": "/where/audio/should/land" }
    }
  }
}

The secret is read at start-up and handed to the process as its own environment variable, so it never appears in argv where other users on the machine could read it. Note the pinned version: a bare npx -y <package> resolves to whatever is newest at that moment, which is how an update lands in the middle of a production run.

Configuration

Variable

Default

Purpose

ELEVENLABS_API_KEY

required

Sent as the xi-api-key header.

ELEVENLABS_OUTPUT_DIR

unset

Where generated audio lands. Without it, small results come back inline as base64 and large ones error.

ELEVENLABS_ENABLED_DOMAINS

all

Comma-separated domains, e.g. text-to-speech,voices. An unknown name is reported, not ignored.

ELEVENLABS_HIDE_DEPRECATED

0

1 drops the 21 operations upstream marks deprecated.

ELEVENLABS_DRY_RUN

0

1 blocks every non-GET call and returns what would have been sent.

ELEVENLABS_API_TIMEOUT_MS

120000

Generation is slow; a dubbing job outlives a CRUD timeout.

ELEVENLABS_MAX_INLINE_BYTES

4194304

Above this, a binary result needs somewhere to be written.

ELEVENLABS_MAX_UPLOAD_BYTES

536870912

Guards against reading an enormous file into memory.

ELEVENLABS_MAX_RATE_LIMIT_WAIT_MS

10000

Longest 429 backoff to sit through before failing.

ELEVENLABS_API_BASE_URL

https://api.elevenlabs.io

Point at a data-residency region if your workspace is in one.

Files in and out

Uploads. Every binary field is offered two ways:

// local install: the server can read your disk
{ "file_path": "/home/me/interview.mp3", "model_id": "scribe_v1" }

// containerised or remote: send the bytes
{ "file_base64": "SUQzB...", "file_filename": "interview.mp3", "model_id": "scribe_v1" }

Fields that accept several files (add_voice, create_finetune, add_pvc_voice_samples) use files_paths / files_base64_list / files_filenames.

Downloads. Anything returning audio, video or a zip takes output_path:

{ "voice_id": "...", "text": "Guten Tag", "output_path": "greeting.mp3" }
// → { "contentType": "audio/mpeg", "bytes": 26375, "path": "/output/dir/greeting.mp3" }

A relative path resolves against ELEVENLABS_OUTPUT_DIR. With no path and no output dir, the audio comes back as an MCP audio block, as long as it is under the inline limit. Base64 inflates by a third and every byte crosses the model's context, so prefer a file for anything longer than a sentence.

Credits and safety

66 operations spend credits, and each one says so in its description. Nothing here guesses on your behalf:

  • ELEVENLABS_DRY_RUN=1 blocks every write and generation call.

  • Only GET is marked read-only. Several POSTs merely query, but every one of them also bills, so they are gated with the writes.

  • DELETE operations carry destructiveHint.

An ElevenLabs key can be restricted per endpoint group, given its own credit quota and locked to an IP range. All three failures arrive as HTTP 401, and this server tells them apart, so "out of credits", "this key may not touch this endpoint" and "this host is not on the allowlist" do not all read as "check your credentials".

Regenerating from the spec

The spec is committed, not fetched at build time, so a vendor edit cannot land in a release nobody reviewed.

npm run fetch-spec   # writes spec/elevenlabs-openapi.json, reports whether it changed
npm run regen        # naming table + generated tools
npm test

Development

npm install
npm run build
npm test             # 47 unit tests
npm run smoke        # live test against the real API, spends a few credits

The smoke test does a round trip that the unit tests cannot: it generates speech to a file, then uploads that same file back to speech-to-text and checks the words come out again.

License

MIT. Not affiliated with ElevenLabs.

Available Tools

390 tools
add_chapterC

Create Chapter Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe name of the chapter, used for identification only.
from_urlNo
project_idYesThe ID of the Studio project.

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly=false, destructive=false, idempotent=false, and openWorld=true. The description does add one genuinely useful trait beyond them: that the call consumes ElevenLabs credits, which matters for a billable creation call. It stops short of explaining required permissions, side effects, or what happens on duplicate names.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is very short and front-loads the action, but the single run-on fragment is grammatically broken, which costs clarity rather than saving space. The terseness here reflects under-specification, not disciplined concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a billable creation tool with no output schema, the description should at minimum state where the chapter is created and what the response contains. Instead it omits when-to-use, the undocumented from_url parameter, and any notion of the created chapter's identity or subsequent workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 67%, and 'from_url' is undocumented in both the schema and the description. The description contributes nothing about project_id, name, or how from_url alters behavior, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'Create Chapter' does name a specific verb and resource, but the sentence is malformed ('Create Chapter Spends ElevenLabs credits'), blurring whether it is one statement or two. It makes no attempt to distinguish this from siblings like edit_chapter, delete_chapter_endpoint, or add_project.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use add_chapter versus edit_chapter, add_project, or convert_chapter_endpoint, and no prerequisites or exclusions are stated. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_documentation_to_knowledge_baseC

Add To Knowledge Base Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoURL to a page of documentation that the agent will have access to in order to interact with users.
nameNoA custom, human-readable name for the document.
agent_idNo
file_pathNoDocumentation that the agent will have access to in order to interact with users. Local path.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.

TDQS

C2.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is covered structurally. The description's only added behavioral fact is that the tool is deprecated upstream, which is genuinely useful context not present in the annotations, but it reveals nothing about mutation effects, permissions, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short and front-loaded, but it is a sentence fragment rather than a complete, appropriately sized description. It avoids waste but errs toward under-specification for a 6-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutation tool with no output schema, the description omits what gets added, how the url vs file_path vs file_base64 inputs differ in use, and what a caller should do instead given the deprecation. The deprecation warning is the sole piece of usable context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 83%, so the schema already documents url, name, file_path, file_base64, and file_filename. The description adds no parameter meaning whatsoever (agent_id is undocumented in both places), so the baseline 3 applies with no uplift.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The text "Add To Knowledge Base" essentially restates the tool name without adding a specific verb+resource elaboration or differentiating it from siblings like add_from_file or create_file_document_route. The only extra information is the deprecation status, which is a status token rather than a purpose statement, so this reads as near-tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Deprecated upstream" implies the agent should prefer alternatives, which is weakly useful guidance, but the description names no replacement tool and states no conditions for when (if ever) this should still be invoked. No when-to-use, when-not-to-use, or alternative is explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_from_fileC

Add A Pronunciation Dictionary

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe name of the pronunciation dictionary, used for identification only.
file_pathNoA lexicon .pls file which we will use to initialize the project with. Local path.
descriptionNoA description of the pronunciation dictionary, used for identification only.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.
workspace_accessNoShould be one of 'admin', 'editor' or 'viewer'. If not provided, defaults to no access.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that — no note on what is created, whether duplicates are allowed, or how file_path vs file_base64 behave. It repeats the title rather than enriching behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short phrase with zero padding, so it is concise and front-loaded. But the brevity reflects under-specification rather than efficiency, giving the agent little to act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutating creation tool with no output schema, a bare resource label is not complete. It omits which file source to prefer, default workspace access behavior, and any outcome information, leaving meaningful gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (name, file_path, description, file_base64, file_filename, workspace_access) is already documented in the schema. The description adds no additional meaning, which is the expected baseline 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource ('pronunciation dictionary') and a verb ('Add'), which clarifies the family of the tool. However, it does not mention the 'from file' mechanism implied by the name, nor does it distinguish this tool from close siblings like add_from_rules or update_pronunciation_dictionaries. Purpose is understandable but shallow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus add_from_rules, add_rules, or update_pronunciation_dictionaries, no prerequisites, and no mention of the required 'name' field or the file source options. The agent must infer usage entirely from the name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_from_rulesC

Add A Pronunciation Dictionary

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe name of the pronunciation dictionary, used for identification only.
rulesYesList of pronunciation rules. Rule can be either: an alias rule: {'string_to_replace': 'a', 'type': 'alias', 'alias': 'b', } or a phoneme rule: {'string_to_replace': 'a', 'type': 'phoneme', 'phoneme': 'b', 'alphabet': 'ipa' }
descriptionNo
workspace_accessNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond that, not explaining side effects, permissions, or how the created dictionary behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is short and not bloated, but it is under-specified for a four-parameter creation tool. It does not front-load the distinguishing detail that the dictionary is created from a rules list.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that accepts a structured array of pronunciation rules plus optional metadata and workspace access, the description is far too thin. It omits the rules-based nature of the operation, optional parameters, and any indication of the creation result or output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: name and rules are documented in the schema, but description and workspace_access are not. The description itself adds no meaning for any parameter, so it fails to compensate for the undocumented optional fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Add A Pronunciation Dictionary'), but it does not mention that the dictionary is built from rules, despite the tool name add_from_rules. It also fails to distinguish the tool from siblings like add_rules, add_from_file, or set_rules.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives. The description gives no prerequisites, context, or exclusions, leaving the agent to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_languageC

Add A Language To The Resource Spends ElevenLabs credits. Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
languageYes
dubbing_idYesID of the dubbing project.

TDQS

C2.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-read-only, non-idempotent, non-destructive mutation, so the bar for additional disclosure is lower. The description adds two meaningful behavioral facts not in the annotations: it spends ElevenLabs credits and is deprecated upstream. It does not cover auth or return behavior, but cost and lifecycle warnings are valuable here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is short and front-loads the action, cost, and deprecation note. However, the first sentence reads awkwardly ('resource spends ElevenLabs credits') and lacks punctuation, so structure is not fully clean.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation tool with no output schema and only 50% schema description coverage, the description omits what resource is being modified, the expected language format, and whether it duplicates dubbing_language_create. The cost/deprecation notes help, but key invocation details are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%; dubbing_id is documented, but language has no schema description and the description provides no format, allowed values, or null semantics for it. The phrase 'Add A Language' merely restates the parameter name and does not compensate for the undocumented language parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb and action ('Add A Language') but the resource is vague ('the Resource'), and it does not distinguish itself from sibling tool dubbing_language_create. An agent cannot tell from the description alone which dubbing resource or endpoint this targets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance or alternative tool is named. The deprecation and credit-cost notes are useful caveats, but they do not tell the agent what to use instead or under what conditions this tool remains appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_mcp_server_tool_approval_routeC

Create Mcp Server Tool Approval

ParametersJSON Schema
NameRequiredDescriptionDefault
tool_nameYesThe name of the MCP tool
input_schemaNoThe input schema of the MCP tool (the schema defined on the MCP server before ElevenLabs does any extra processing)
mcp_server_idYesID of the MCP Server.
approval_policyNoDefines the tool-level approval policy.
tool_descriptionYesThe description of the MCP tool

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false and openWorldHint=true, so the safety profile is partly covered. The description adds nothing on top of that: it doesn't say whether re-adding an existing tool approval fails or overwrites (relevant given idempotentHint=false), what the approval_policy values actually gate, or what permission is needed to register an MCP tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but it is an under-specified fragment that lacks a terminating sentence and conveys no actionable content. Brevity here is a symptom of omission rather than of efficient writing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a write operation with five parameters including a nested object and no output schema, and the description says nothing about required inputs, the meaning of the approval policy choices, or the relationship to the paired remove_/update_ MCP approval tools. An agent would have to reconstruct intent entirely from the identifier and schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters including the approval_policy enum are documented in the schema itself; per the calibration rule this sets a baseline of 3. The description contributes no additional meaning about required fields or the semantics of the nested input_schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is essentially the tool title with capitalization stripped ('Create Mcp Server Tool Approval' vs. title 'Add Mcp Server Tool Approval Route'), so it reads as a tautological restatement of the name rather than an explanation. It does name a verb and a resource, but an agent gets no more information than it already had from the identifier, and nothing distinguishes it from close siblings like update_mcp_server_approval_policy_route or add_mcp_tool_config_override_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling handles the adjacent cases (removing an approval route, listing MCP tools, updating an approval policy). The only guidance available is the implied 'this creates an approval route' inference from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_mcp_tool_config_override_routeD

Create Mcp Tool Configuration Override

ParametersJSON Schema
NameRequiredDescriptionDefault
tool_nameYesThe name of the MCP tool
assignmentsNo
environmentNoEnvironment whose values are used when the MCP server URL, headers, or auth connection reference environment variables. Mirrors the environment a conversation would run in; defaults to production.
mcp_server_idYesID of the MCP Server.
execution_modeNo
response_mocksNo
input_overridesNo
pre_tool_speechNo
tool_call_soundNoOverrides the server's tool_call_sound setting for this tool. A sound name plays that sound; 'off' overrides to no sound (silence); null means do not override (inherit the server default).
interruption_modeNo
disable_interruptionsNo
force_pre_tool_speechNo
response_timeout_secsNo
tool_call_sound_behaviorNoDetermines how the tool call sound should be played.

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, but the description adds no behavioral context such as what an override changes, duplicate handling, permissions, or side effects. It neither contradicts nor supplements the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but its brevity reflects under-specification for a complex 14-parameter creation tool rather than useful concision. It is front-loaded only as a bare label.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 14 parameters, low schema description coverage, no output schema, and only sparse annotations, the description is inadequate. It omits required inputs, override semantics, and operational constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 36% across 14 parameters, so the description should compensate, but it mentions no parameters, formats, or semantics. It adds zero meaning beyond what the sparse schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name/title as 'Create Mcp Tool Configuration Override' without adding scope, target, or distinguishing details from siblings like update_mcp_tool_config_override_route. It names a verb and resource, but is essentially tautological.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no when-to-use guidance, prerequisites, or alternatives. An agent cannot tell when to add an override versus updating, reading, or removing one from this text alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_memberC

Add Member To User Group

ParametersJSON Schema
NameRequiredDescriptionDefault
emailYesThe email of the target workspace member.
group_idYesThe ID of the target group.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds nothing on top of that: it does not say what happens on a duplicate add (non-idempotency matters here), what permissions are required, or what errors can occur.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase, so it is front-loaded and free of bloat. It is arguably too terse for its job, but conciseness itself is not the problem.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a mutating membership operation with no output schema and only 2 parameters; a correct invocation depends on knowing prerequisites (member exists in workspace, caller has group-admin rights), duplicate behavior, and failure modes, none of which are stated. For a non-idempotent mutation the description is under-specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both group_id and email documented in the schema, so the baseline is 3. The description contributes no additional meaning about the parameters (e.g., whether email must be an existing workspace member), but nothing is missing from the structured data.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a verb+resource ('Add Member To User Group'), which clarifies that the resource is user-group membership rather than, say, a project or workspace invite. However, it is nearly a verbatim restatement of the tool name and title, adding little beyond the group scoping detail, and it does not distinguish itself from nearby siblings such as invite_user or add_project.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as invite_user, add_member, or update_workspace_member, nor any stated prerequisites (e.g., the member must already exist in the workspace). The one-line description leaves all routing decisions to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_projectB

Create Studio Project Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe name of the Studio project, used for identification only.
titleNoAn optional name of the author of the Studio project, this will be added as metadata to the mp3 file on Studio project or chapter download.
authorNoAn optional name of the author of the Studio project, this will be added as metadata to the mp3 file on Studio project or chapter download.
genresNoAn optional list of genres associated with the Studio project.
fictionNoAn optional specification of whether the content of this Studio project is fiction.
from_urlNoAn optional URL from which we will extract content to initialize the Studio project. If this is set, 'from_url' and 'from_content' must be null. If neither 'from_url', 'from_document', 'from_content' are provided we will initialize the Studio project as blank.
languageNoAn optional language of the Studio project. Two-letter language code (ISO 639-1).
descriptionNoAn optional description of the Studio project.
isbn_numberNoAn optional ISBN number of the Studio project you want to create, this will be added as metadata to the mp3 file on Studio project or chapter download.
source_typeNoThe type of Studio project to create.
auto_convertNoWhether to auto convert the Studio project to audio or not.
callback_urlNoA url that will be called by our service when the Studio project is converted. Request will contain a json blob containing the status of the conversion Messages: 1. When project was converted successfully: { type: "project_conversion_status", event_timestamp: 1234567890, data: { request_id: "1234567
content_typeNoAn optional content type of the Studio project.
mature_contentNoAn optional specification of whether this Studio project contains mature content.
quality_presetNoOutput quality of the generated audio. Must be one of: 'standard' - standard output format, 128kbps with 44.1kHz sample rate. 'high' - high quality output format, 192kbps with 44.1kHz sample rate and major improvements on our side. 'ultra' - ultra quality output format, 192kbps with 44.1kHz sample r
voice_settingsNoOptional voice settings overrides for the project, encoded as a list of JSON strings. Example: ["{\"voice_id\": \"21m00Tcm4TlvDq8ikWAM\", \"stability\": 0.7, \"similarity_boost\": 0.8, \"style\": 0.5, \"speed\": 1.0, \"use_speaker_boost\": true}"]
target_audienceNoAn optional target audience of the Studio project.
default_model_idNoThe ID of the model to be used for this Studio project, you can query GET /v1/models to list all available models.
from_content_jsonNoAn optional content to initialize the Studio project with. If this is set, 'from_url' and 'from_document' must be null. If neither 'from_url', 'from_document', 'from_content' are provided we will initialize the Studio project as blank. Example: [{"name": "Chapter A", "blocks": [{"sub_type": "p", "no
auto_assign_voicesNo[Alpha Feature] Whether automatically assign voices to phrases in the create Project.
from_document_pathNoAn optional .epub, .pdf, .txt or similar file can be provided. If provided, we will initialize the Studio project with its content. If this is set, 'from_url' and 'from_content' must be null. If neither 'from_url', 'from_document', 'from_content' are provided we will initialize the Studio project as
from_document_base64NoBase64 contents for "from_document". Use this when the server cannot read your local disk.
volume_normalizationNoWhen the Studio project is downloaded, should the returned audio have postprocessing in order to make it compliant with audiobook normalized volume requirements
create_publishing_readNoIf true, creates a corresponding read for direct publishing in draft state
default_title_voice_idNoThe voice_id that corresponds to the default voice used for new titles.
from_document_filenameNoFilename to send for "from_document". Some endpoints infer the audio format from it.
acx_volume_normalizationNo[Deprecated] When the Studio project is downloaded, should the returned audio have postprocessing in order to make it compliant with audiobook normalized volume requirements
apply_text_normalizationNoThis parameter controls text normalization with four modes: 'auto', 'on', 'apply_english' and 'off'. When set to 'auto', the system will automatically decide whether to apply text normalization (e.g., spelling out numbers). With 'on', text normalization will always be applied, while with 'off', it w
original_publication_dateNoAn optional original publication date of the Studio project, in the format YYYY-MM-DD or YYYY.
default_paragraph_voice_idNoThe voice_id that corresponds to the default voice used for new paragraphs.
pronunciation_dictionary_locatorsNoA list of pronunciation dictionary locators (pronunciation_dictionary_id, version_id) encoded as a list of JSON strings for pronunciation dictionaries to be applied to the text. A list of json encoded strings is required as adding projects may occur through formData as opposed to jsonBody. To specif

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds one important behavioral fact beyond annotations: the operation spends ElevenLabs credits. It does not mention required permissions, rate limits, or output behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise and front-loaded, with no wasted words. However, the phrasing 'Create Studio Project Spends ElevenLabs credits' is grammatically awkward and reads as two clauses fused together, slightly reducing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 31 parameters, no output schema, and rich annotations plus full schema descriptions, the description is minimally adequate. It omits when-to-use context and does not clarify what a Studio Project is or how initialization sources work, but the schema supplies those details; the main gaps are usage guidance and return expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all 31 parameters are fully documented in the input schema, including optional initialization sources and metadata fields. The description adds no additional parameter-level meaning beyond noting credit consumption, which is the baseline level when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Create') and resource ('Studio Project'), distinguishing it from broader sibling creation tools like create_audio_native_project or create_podcast. It does not explicitly contrast with the nearest sibling, edit_project, but the resource is clear enough to identify the intended action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives no indication of when to use this tool versus alternatives such as create_audio_native_project, create_podcast, or add_chapter. The only guidance is implicit in the tool name; no prerequisites, exclusions, or selection criteria are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_pvc_voice_samplesC

Add Samples To Pvc Voice

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
files_pathsNoAudio files used to create the voice. Local paths. Required for this call.
files_filenamesNoFilenames to send for "files". Some endpoints infer the audio format from them.
files_base64_listNoBase64 contents for "files", one entry per file. Use this when the server cannot read your local disk.
remove_background_noiseNoIf set will remove background noise for voice samples using our audio isolation model. If the samples do not include background noise, it can make the quality worse.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=false, openWorldHint=true, and destructiveHint=false. The description adds no behavioral context—it does not mention permissions, what happens to existing samples, whether training is required, or mutation impact. Given annotations already carry the safety profile, the description contributes nothing beyond them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single five-word phrase, front-loaded and with no wasted words. However, it is too terse to be useful, essentially duplicating the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with five parameters, open-world external call, and no output schema, the description is severely lacking. It provides no information about operation, return, or side effects, leaving the agent with only annotations and schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has 100% description coverage, documenting all five parameters (voice_id, files_paths, files_filenames, files_base64_list, remove_background_noise). The description provides no additional parameter semantics or format details, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Add Samples To Pvc Voice' simply restates the tool name and title. It states a verb and resource, but adds no scope, no distinction from siblings like edit_pvc_voice_sample or create_pvc_voice, and does not clarify what 'samples' means in this context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as edit_pvc_voice_sample, delete_pvc_voice_sample, or run_pvc_voice_training. No prerequisites or conditions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_rulesC

Add Rules To The Pronunciation Dictionary

ParametersJSON Schema
NameRequiredDescriptionDefault
rulesYesList of pronunciation rules. Rule can be either: an alias rule: {'string_to_replace': 'a', 'type': 'alias', 'alias': 'b', } or a phoneme rule: {'string_to_replace': 'a', 'type': 'phoneme', 'phoneme': 'b', 'alphabet': 'ipa' }
pronunciation_dictionary_idYesThe id of the pronunciation dictionary

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the mutation and non-idempotency profile is covered structurally. The description adds nothing beyond that — no mention of whether rules are appended, deduplicated, or validated, nor any auth/permission context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with zero padding, front-loaded with the action. It is efficient, though the brevity is partly under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent mutation on a nested rules array with no output schema, the description should say what happens on repeated calls, whether existing rules are preserved, and what a successful result looks like. None of that is present, leaving real gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including a detailed breakdown of alias vs phoneme rule shapes, so the schema already carries the parameter burden. The description adds no syntax, format, or constraint detail beyond that, which is the baseline 3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb (add) and resource (rules to the pronunciation dictionary), but it essentially restates the tool name/title with no differentiating detail. With close siblings like set_rules, remove_rules, and add_from_rules present, an agent cannot tell from this text alone which one to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus set_rules (which likely replaces rules), remove_rules, or add_from_rules. No prerequisites, no context about dictionary state, no indication of whether adding is additive or overwriting.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_sharing_voiceD

Add Shared Voice

ParametersJSON Schema
NameRequiredDescriptionDefault
new_nameYesThe name that identifies this voice. This will be displayed in the dropdown of the website.
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
bookmarkedNo
public_user_idYesPublic user ID used to publicly identify ElevenLabs users.

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true. The description adds no behavioral context beyond those annotations, such as permission requirements, effects on the voice library, or reverting behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but this reflects under-specification rather than effective conciseness. It is front-loaded only in the sense that it is a single phrase.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's mutation behavior, four parameters, lack of output schema, and dense sibling namespace, the description is completely inadequate. It does not explain what sharing a voice entails, what happens after adding, or how this differs from related voice-management tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, with three of four parameters documented in the schema. The description adds no parameter meaning or usage detail, and the undocumented optional bookmarked parameter is not addressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name and title without adding a specific verb+resource distinction. It does not differentiate from siblings like add_voice, create_voice, or share_resource_endpoint, leaving the exact meaning of 'Add Shared Voice' unclear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative guidance is provided. The description gives no context for selecting this tool over other voice or sharing tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_ticket_comment_routeC

Add Comment To Agent Conversation Ticket

ParametersJSON Schema
NameRequiredDescriptionDefault
commentYesA comment discussing how to resolve the ticket.
agentqa_ticket_idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false and destructiveHint=false, so the agent knows this is a non-idempotent, non-destructive write. The description adds nothing beyond that: no note that duplicate comments accumulate on repeat calls, no permissions or visibility notes, and no indication of whether the comment notifies anyone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no wasted words, which is good, but it reads as a heading rather than a front-loaded sentence with substance. Brevity here reflects under-specification rather than disciplined concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, non-idempotent write tool with no output schema and only half its parameters documented, the description should carry more weight. It omits the ticket identifier's semantics, the effect of the comment (visibility, notifications), and how it differs from sibling comment/update tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: the 'comment' parameter is documented in the schema, but 'agentqa_ticket_id' is undocumented. The description could have clarified the ticket identifier's format or provenance but instead provides no parameter information at all, leaving the required ID ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb and resource ('Add Comment To Agent Conversation Ticket'), so the basic action is inferable, but it is essentially the tool title restated in title case. It does nothing to separate this tool from close siblings such as add_turn_comment_route or update_agent_conversation_ticket_route, so an agent could easily pick the wrong one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when to prefer an alternative (e.g., add_turn_comment_route, update_agent_conversation_ticket_route), or any prerequisite such as the ticket needing to exist. Usage must be inferred entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_tool_routeD

Add Tool

ParametersJSON Schema
NameRequiredDescriptionDefault
tool_configYesConfiguration for the tool
response_mocksNo

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description adds no further behavioral context such as side effects, required IDs, approval policy impact, or what configuration gets created.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short but under-specified rather than efficiently concise. It is not front-loaded with useful information and does not earn its place for a complex tool with a large nested configuration schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that accepts a sizeable nested tool_config object with many optional fields and no output schema, the description is completely inadequate. An agent would need to infer almost everything from the schema and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, and the description provides no parameter information at all. It does not compensate for the undocumented top-level response_mocks parameter or clarify the nested tool_config fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is just "Add Tool", which essentially restates the tool name and annotation title "Add Tool Route" without adding scope, resource detail, or sibling differentiation. It does not distinguish this from update_tool_route, delete_tool_route, or get_tool_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as update_tool_route or add_mcp_tool_config_override_route. The description provides no context, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_turn_comment_routeC

Add Turn Comment To Agent Conversation Ticket

ParametersJSON Schema
NameRequiredDescriptionDefault
commentYesWhat went wrong at this turn.
turn_indexYesZero-based index of the transcript turn this comment refers to.
agentqa_ticket_idYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, openWorldHint=true, and destructiveHint=false. The description adds nothing beyond that profile—it does not explain whether comments are appended, whether duplicates are created on retry, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase, but it is under-specified rather than concise. No sentence earns extra value because there is essentially only a restated title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent mutation with three required parameters and no output schema, the description is inadequate. It omits when to use it, how it relates to the ticket lifecycle, and what the agentqa_ticket_id refers to.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%: comment and turn_index are described in the schema, but agentqa_ticket_id is undocumented. The description contributes no parameter meaning at all, so it fails to compensate for the gap and adds nothing over the structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Add Turn Comment To Agent Conversation Ticket' is essentially a title-cased restatement of the tool name. It identifies the verb (add) and resource (turn comment on a conversation ticket) but provides no detail about what a turn comment is or how this differs from the sibling add_ticket_comment_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives such as add_ticket_comment_route or update_agent_conversation_ticket_route. The agent is left to infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_voiceD

Add Voice

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe name that identifies this voice. This will be displayed in the dropdown of the website.
labelsNoLabels for the voice. Keys can be language, accent, gender, or age.
descriptionNoA description of the voice.
files_pathsNoA list of file paths to audio recordings intended for voice cloning. Local paths. Required for this call.
files_filenamesNoFilenames to send for "files". Some endpoints infer the audio format from them.
files_base64_listNoBase64 contents for "files", one entry per file. Use this when the server cannot read your local disk.
remove_background_noiseNoIf set will remove background noise for voice samples using our audio isolation model. If the samples do not include background noise, it can make the quality worse.

TDQS

D1.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false, but the description adds zero context: no mention that voice cloning consumes uploaded audio, that training may be asynchronous, or that repeated calls create duplicates (non-idempotent). It contributes nothing beyond the structured data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At two words it is technically short, but this is under-specification rather than conciseness; there is no front-loaded explanation of purpose, scope, or arguments. Nothing earns or fails to earn a place because nothing is there.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, non-idempotent mutation tool with no output schema, the definition is completely inadequate. It omits the cloning workflow, the required-vs-optional file input patterns, and any distinction from the many sibling voice tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all 7 parameters, so the schema itself documents name, labels, description, file paths/filenames/base64 lists, and background-noise removal. Baseline is 3 when the schema does the heavy lifting, and the description adds no additional parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The entire description is "Add Voice", which merely restates the tool name and the annotation title. It gives no verb-plus-resource specificity beyond the name itself and does nothing to distinguish this tool from siblings such as create_voice, edit_voice, delete_voice, or add_sharing_voice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance whatsoever about when to use this tool versus create_voice/add_sharing_voice/add_pvc_voice_samples, nor any stated prerequisites. The agent is left to infer everything from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_testing_bulk_move_routeC

Bulk Move Tests To Folder

ParametersJSON Schema
NameRequiredDescriptionDefault
move_toNo
entity_idsYesThe IDs of tests or folders to move.

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true. The description adds only 'bulk' and 'tests to folder', with no additional behavioral context such as error handling, permissions, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The phrase is very short, front-loaded, and has no wasted words. It is arguably under-specified for a bulk mutation tool, but it does not suffer from verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation operation with openWorldHint=true, no output schema, and one undocumented parameter, the description omits success/failure behavior, required permissions, and result format. It is too sparse for confident invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: entity_ids is described in the schema, but move_to has no description. The description vaguely implies a destination folder but does not clarify accepted values, null handling, or that entity_ids can include folders.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Bulk Move Tests To Folder'), enough to know it moves tests/folders. It does not differentiate from sibling move endpoints or explicitly mention the agent-testing context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, prerequisites, or alternative tools are mentioned. Usage is only implied by the operation name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assign_conversation_tags_routeD

Assign Conversation Tags

ParametersJSON Schema
NameRequiredDescriptionDefault
tag_idsYesTag IDs to add to the conversation. Re-assigning an existing tag is a no-op.
conversation_idYes

TDQS

D1.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, openWorldHint=true, and idempotentHint=false, so the safety profile is covered structurally. The description adds nothing — no note on whether existing tags are replaced or merged, no permission requirements, no indication of how many tags may be assigned at once.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but the brevity reflects under-specification rather than efficiency — three words that merely echo the tool name. There is nothing to be front-loaded because nothing is said.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with a partially documented schema, no output schema, and a large family of sibling tag/conversation tools, this description is inadequate. An agent cannot determine arguments, effect, or side effects from it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 2 parameters and only 50% schema description coverage, conversation_id is undocumented in both the schema and the description. The description supplies no parameter meaning at all, leaving the coverage gap entirely unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name with title-casing — 'Assign Conversation Tags' — adding no scope, target system, or distinguishing detail. It is a tautology: the agent learns nothing beyond what the identifier already conveys.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus clearly related siblings such as unassign_conversation_tag_route, update_conversation_tag_route, or create_conversation_tag_route. No prerequisites, no conditions, no alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audio_isolationB

Audio Isolation Spends ElevenLabs credits. Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
audio_pathNoThe audio file from which vocals/speech will be isolated from. Local path. Required for this call.
file_formatNoThe format of input audio. Options are 'pcm_s16le_16' or 'other' For `pcm_s16le_16`, the input audio must be 16-bit PCM at a 16kHz sample rate, single channel (mono), and little-endian byte order. Latency will be lower than with passing an encoded waveform.
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
preview_b64NoOptional preview image base64 for tracking this generation.
audio_base64NoBase64 contents for "audio". Use this when the server cannot read your local disk.
audio_filenameNoFilename to send for "audio". Some endpoints infer the audio format from it.

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations present (readOnlyHint=false, openWorldHint=true, idempotentHint=false), the description still adds non-obvious context: it consumes ElevenLabs credits and returns audio/mpeg bytes that are saved via output_path or returned inline. That cost/return behavior is not derivable from the annotations. It does not disclose rate limits or failure modes, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with essentially no filler, and the operative information (output_path saves the returned bytes) is placed clearly. The opening phrase is awkwardly capped and low-value, which keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter generation tool with a fully documented schema and no output schema, the description covers cost and return format, which is useful. However, it omits the core operation and any routing versus the parallel stream/stem-separation tools, leaving the agent to infer the tool's actual function.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters thoroughly. The description only restates the output_path save behavior and the inline-bytes default, adding no syntax or format detail beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The lead sentence is essentially the title restated ('Audio Isolation') plus a cost note, so the description never states the actual verb+resource (isolating vocals/speech from a source file) — that meaning lives only in the schema's audio_path description. The only genuinely informative content is about output format, not purpose. Siblings like audio_isolation_stream, separate_song_stems and start_speaker_separation are not distinguished.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus alternatives; the near-identical audio_isolation_stream sibling is never mentioned, nor is there any prerequisite or exclusion. The only quasi-usage note is that you should pass output_path for large files, which is really a parameter hint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audio_isolation_streamB

Audio Isolation Stream Spends ElevenLabs credits. Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
audio_pathNoThe audio file from which vocals/speech will be isolated from. Local path. Required for this call.
file_formatNoThe format of input audio. Options are 'pcm_s16le_16' or 'other' For `pcm_s16le_16`, the input audio must be 16-bit PCM at a 16kHz sample rate, single channel (mono), and little-endian byte order. Latency will be lower than with passing an encoded waveform.
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
audio_base64NoBase64 contents for "audio". Use this when the server cannot read your local disk.
audio_filenameNoFilename to send for "audio". Some endpoints infer the audio format from it.

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint false, openWorldHint true, idempotentHint false, destructiveHint false, giving the safety profile. The description adds useful behavior beyond annotations: it warns that the call spends ElevenLabs credits, states the return format (audio/mpeg bytes), and explains how to persist output via output_path. It does not cover auth or rate limits, but adds solid behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no waste. The cost and output handling are front-loaded and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with full schema coverage and annotations, the description covers cost and output handling. However, it omits the core purpose statement (isolation of vocals/speech from audio) and any distinction from sibling audio_isolation, leaving an agent to infer when to use the stream variant.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 even without parameter details in the description. The description mentions output_path only, which the schema already fully documents; no additional parameter meaning is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with "Audio Isolation Stream Spends ElevenLabs credits," which repeats the tool name and adds a cost note, but does not state a verb or what is isolated (vocals/speech from audio). The schema's audio_path description supplies the missing purpose, but the description alone lacks a clear action statement and does not distinguish from sibling audio_isolation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use/when-not, no alternatives, no prerequisite context. The only contextual cue is the cost warning, which is a side effect, not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audio_native_project_update_content_endpointC

Update Audio-Native Project Content

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathNoEither txt or HTML input file containing the article content. HTML should be formatted as follows '&lt;html&gt;&lt;body&gt;&lt;div&gt;&lt;p&gt;Your content&lt;/p&gt;&lt;h5&gt;More of your content&lt;/h5&gt;&lt;p&gt;Some more of your content&lt;/p&gt;&lt;/div&gt;&lt;/body&gt;&lt;/html&gt;' Local path
project_idYesThe ID of the Studio project.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
auto_convertNoWhether to auto convert the project to audio or not.
auto_publishNoWhether to auto publish the new project snapshot after it's converted.
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true, which already tells the agent this is a non-idempotent mutating write. The description adds nothing beyond that baseline: it doesn't disclose that it replaces existing project content, whether auto_convert/auto_publish trigger irreversible conversions or publication, or what side effects occur. With annotations carrying the safety profile, some latitude is allowed, but a bare name is still insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is maximally short and front-loaded, so it is not bloated, but the brevity is under-specification rather than genuine conciseness. Every sentence should earn its place, and here the single phrase conveys no information beyond the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter mutating endpoint with no output schema, no annotations explaining side effects, and no disclosure of the convert/publish behaviors, the description is materially incomplete. It leaves the agent dependent entirely on the schema and the tool name to understand the operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every one of the six parameters (file_path, project_id, file_base64, auto_convert, auto_publish, file_filename) is documented in the schema itself. The description adds no parameter context whatsoever, which is acceptable at high coverage but caps value at the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update Audio-Native Project Content' only restates the tool name. It says nothing specific about what kind of content is updated, how it's supplied (text/HTML file or base64), or how it differs from siblings like audio_native_update_content_from_url or edit_project_content. It is effectively a tautology of the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of the alternative audio_native_update_content_from_url (which updates content from a URL rather than a local file), and no prerequisites. An agent cannot tell from the description alone which of the two update-content endpoints to call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audio_native_update_content_from_urlC

Update Audio-Native Content From Url

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL of the page to extract content from.
titleNo
authorNo

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, and openWorldHint=true, so the mutation and network-fetch nature is covered structurally. The description adds nothing on top: it does not say whether the update replaces or merges existing content, whether it fetches/parses the URL immediately, or how errors are surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but brevity here reflects under-specification rather than efficient structure; the single phrase earns its place only as a name echo. Nothing is front-loaded because there is nothing substantive to front-load.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an open-world, non-idempotent mutation with no output schema and two undocumented parameters, the description is far too thin. An agent cannot determine the target resource, the effect on existing content, or the expected inputs for title/author.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% – url is documented but title and author are not – so the description should compensate and does not, listing no fields at all. It also gives no hint as to how the target audio-native project is identified, given there is no project identifier parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is essentially the title restated: 'Update Audio-Native Content From Url' repeats the tool name verbatim without stating what content is updated, from which page structure, or what 'Audio-Native Content' means. It does not distinguish this tool from the near-identical sibling audio_native_project_update_content_endpoint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no exclusions, and no reference to alternatives such as audio_native_project_update_content_endpoint, create_audio_native_project, or refresh_url_document_route. The agent is left to infer the scenario entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_batch_callC

Cancel A Batch Call.

ParametersJSON Schema
NameRequiredDescriptionDefault
batch_idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, but the description adds nothing about what cancelling does to in-progress calls, whether it is reversible, or what state results. destrictiveHint=false alongside a 'cancel' verb is potentially confusing and unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short and front-loaded, but its brevity is a result of under-specification rather than disciplined editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema and zero parameter documentation, the description leaves the agent without the state requirements, side effects, or parameter meaning needed to invoke it confidently among several batch-call siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter batch_id has no description anywhere. The description does not compensate by clarifying the format or source of the batch identifier.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb+resource (cancel a batch call), which is clear enough on its own but is essentially a restatement of the title. It does not distinguish itself from the closely-related siblings delete_batch_call, retry_batch_call, and get_batch_call, which the agent must choose between.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance about when to use this tool versus delete_batch_call or retry_batch_call, or what state the batch must be in to be cancellable. The agent is left to infer the difference from names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_crawl_job_routeC

Cancel Crawl Job

ParametersJSON Schema
NameRequiredDescriptionDefault
crawl_job_idYesThe id of the crawl job to retrieve

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true), so the bar is lower. The description adds no additional behavioral context, such as whether cancellation is immediate, whether it requires the job to be active, or what happens on an already-cancelled job.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At only two words, the description is concise but severely under-specified for a mutating tool with a required identifier. The dimension penalizes under-specification as much as verbosity, and this is too sparse to be useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent, non-read-only action with a required crawl_job_id and no output schema, the description should at least clarify usage context or basic effects. It supplies no information beyond the title, leaving the agent relying entirely on annotations and the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 even though the description says nothing about parameters. The schema itself contains the parameter description, though its wording ("to retrieve") is semantically off for a cancel operation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Cancel Crawl Job" is essentially a restatement of the tool name/title, making it a tautology rather than an elaboration. It does name the action and resource, but adds no differentiating detail beyond what the name already conveys.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is provided, nor are alternatives such as get_crawl_job_route or list_crawl_jobs_route mentioned. The implied usage is only that the tool cancels a crawl job, which an agent could already infer from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_file_upload_routeC
DestructiveIdempotent

Delete File Upload

ParametersJSON Schema
NameRequiredDescriptionDefault
file_idYes
conversation_idYes

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, openWorldHint=true, and readOnlyHint=false. The description only restates 'Delete File Upload' and adds no behavioral context such as what exactly gets destroyed, required permissions, or side effects. No annotation contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is extremely short and front-loaded, but the issue is under-specification rather than conciseness. A single phrase cannot carry a destructive two-parameter operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although annotations cover the safety profile and there is no output schema, the description still omits critical context: what file upload is deleted, how conversation_id scopes it, and when to choose this over sibling tools. Inadequate for a mutation tool with two required parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%; neither file_id nor conversation_id has a schema description. The description provides no format, scope, or relationship information for these required parameters, so it does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource, 'Delete File Upload', but it does not distinguish this operation from sibling tools like upload_file_route or the many delete_* routes, nor does it clarify scope (e.g., conversation-bound cancellation). It is essentially a short restatement of the operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, what prerequisites are required, or how it differs from alternatives such as upload_file_route or other delete endpoints. The agent receives no routing information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compile_procedures_routeC

Compile Procedures

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesAgent ID to get the procedure draft from
branch_idYesBranch ID to get the procedure draft from

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations state readOnlyHint=false, idempotentHint=false, openWorldHint=true, destructiveHint=false. The description adds nothing about side effects, such as whether compilation overwrites an existing procedure or is read-only. This is inadequate for a non-idempotent write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two words that merely repeat the title. It is concise but not informative, failing to front-load any useful details. A description should earn its place; this one does not.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with no output schema and moderate complexity (two required inputs), the description should clarify the operation's effect and any traversal or compilation semantics. It omits all such context, leaving an agent without enough information to safely invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both parameters documented as 'Agent ID to get the procedure draft from' and 'Branch ID to get the procedure draft from'. Per the rubric, this sets a baseline of 3. The description contributes nothing beyond the schema, so it cannot exceed the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is essentially a restatement of the tool name: 'Compile Procedures'. It does not specify what is being compiled, what the inputs represent (procedure draft for an agent and branch according to the schema), or what the outcome is (likely a compiled procedure from a draft). No resource beyond the tool name is clarified.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to use this tool or how it differs from siblings like get_procedure_draft_route, update_procedure_draft_route, create_procedure_route, or list_procedures_route. The description offers no context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compose_detailedC

Compose Music With A Detailed Response Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
promptNo
model_idNo
finetune_idNo
lyrics_textNo
music_promptNoComposition plan for the `music_v1` model. Using this field with any other model will result in an error.
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. Use "auto" (the default) to let the API pick the best format for the selected model: mp3_44100_128 for v1 models and mp3_48000_192 for v2 models.
sign_with_c2paNoWhether to sign the generated song with C2PA. Applicable only for mp3 files.
generation_modeNo
music_length_msNo
with_timestampsNoWhether to return the timestamps of the words in the generated song.
composition_planNo
finetune_strengthNoHow strongly the finetune influences the generation. Defaults to 1.0 (full strength). Lower values soften the influence of the finetune, leaving more room for prompt-level steering. Only meaningful when `finetune_id` is also provided.
force_instrumentalNoIf true, guarantees that the generated song will be instrumental. If false, the song may or may not be instrumental depending on the `prompt`. Can only be used with `prompt`.
model_style_prefixNo
use_phonetic_namesNoIf true, proper names in the prompt will be phonetically spelled in the lyrics for better pronunciation by the music model. The original names will be restored in word timestamps.
store_for_inpaintingNoWhether to store the generated song for inpainting.
with_waveform_visualNoWhether to return the visual waveform of the generated song.
respect_sections_durationsNoControls how strictly section durations in the `composition_plan` are enforced. Only used with `composition_plan` and only applies to `music_v1`; for `music_v2` and `music_v2_5` section durations are always enforced and this is ignored. When false for `music_v1`, the model may adjust individual sect

TDQS

C2.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so safety and idempotency are covered. The description's one extra fact – that it spends ElevenLabs credits – is genuinely useful beyond the annotations, though thin. It says nothing about latency, rate limits, or auth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single run-on sentence is short but poorly structured and mixes purpose with a cost note without punctuation. It is under-specification rather than effective conciseness, and not front-loaded with the key distinguishing behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 19-parameter music-generation tool with no output schema and complex nested inputs like music_prompt and composition_plan, the description is drastically incomplete. It omits model selection guidance, the music_v1-only constraint of music_prompt, and any notion of what 'detailed response' returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 53%, and the description itself mentions no parameters. For the roughly half of parameters that lack schema descriptions (seed, prompt, lyrics_text, generation_mode, music_length_ms, etc.), the description provides no compensating detail. Baseline 3 fits a mid-coverage schema with no param info in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Compose Music With A Detailed Response Spends ElevenLabs credits' restates the title and adds only a vague cost note. It doesn't specify what 'detailed response' means or how this differs from siblings like compose_plan, stream_compose, or compose_detailed_stream, which all involve music composition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no alternatives mentioned, and no preconditions stated. An agent cannot tell from this description when to choose compose_detailed over compose_plan or compose_detailed_stream.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compose_detailed_streamC

Stream Composed Music With A Detailed Response Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
promptNo
model_idNo
finetune_idNo
lyrics_textNo
music_promptNoComposition plan for the `music_v1` model. Using this field with any other model will result in an error.
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. Use "auto" (the default) to let the API pick the best format for the selected model: mp3_44100_128 for v1 models and mp3_48000_192 for v2 models.
generation_modeNo
music_length_msNo
with_timestampsNoWhether to return the timestamps of the words in the generated song.
composition_planNo
finetune_strengthNoHow strongly the finetune influences the generation. Defaults to 1.0 (full strength). Lower values soften the influence of the finetune, leaving more room for prompt-level steering. Only meaningful when `finetune_id` is also provided.
force_instrumentalNoIf true, guarantees that the generated song will be instrumental. If false, the song may or may not be instrumental depending on the `prompt`. Can only be used with `prompt`.
use_phonetic_namesNoIf true, proper names in the prompt will be phonetically spelled in the lyrics for better pronunciation by the music model. The original names will be restored in word timestamps.
store_for_inpaintingNoWhether to store the generated song for inpainting.
with_waveform_visualNoWhether to return the visual waveform of the generated song.

TDQS

C2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations provided (readOnlyHint=false, idempotentHint=false, destructiveHint=false, openWorldHint=true), the description notes it spends credits but adds little beyond that. It doesn't describe streaming behavior, response format, or the 'detailed response' mentioned. With annotations covering safety, this is a moderate gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with awkward capitalization, but it is brief. However, conciseness here comes at the cost of under-specification rather than efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 16 parameters, no output schema, partial schema coverage, and complex music composition functionality, the description is grossly inadequate. It doesn't explain streaming, detailed responses, parameter interactions, or credit costs beyond the vague mention.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

16 parameters with 50% schema coverage, and the description adds zero parameter information. Many parameters lack descriptions in the schema (seed, prompt, model_id, finetune_id, lyrics_text, generation_mode, music_length_ms, composition_plan), so the description should compensate but doesn't.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Stream Composed Music With A Detailed Response Spends ElevenLabs credits' restates the tool name rather than explaining its actual behavior. While it implies music composition with streaming, it doesn't distinguish this tool from close siblings like compose_detailed, compose_plan, compose_detailed_stream, or stream_compose. It's more tautology than differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no alternatives mentioned, no conditions for selection. Despite having siblings like compose_detailed, compose_plan, and stream_compose, the description offers no routing information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compose_planC

Generate Composition Plan Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesA simple text prompt to compose a plan from.
model_idNo
music_length_msNo
source_composition_planNoAn optional composition plan to use as a source for the new composition plan.

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is covered. The description's only added behavioral fact is that the call 'Spends ElevenLabs credits', which is a genuinely useful, non-obvious cost side effect that the schema does not convey. It says nothing further about latency, external service dependency, or auth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is only two fragments long, so there is no bloat, but the phrasing is ungrammatical ('Generate Composition Plan Spends ElevenLabs credits') and reads as two unpunctuated notes rather than a coherent statement. Brevity here reflects under-specification rather than disciplined editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a four-parameter generation tool with no output schema, so the description is the only place an agent could learn what a 'composition plan' actually is, which model to pick, how source_composition_plan is consumed, or what comes back. None of that is present, and the sole behavioral note is about credit consumption.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: 'prompt' and 'source_composition_plan' are documented, but 'model_id' (an enum with three versions) and 'music_length_ms' have no descriptions in either the schema or the tool description. The description contributes zero parameter meaning, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a verb ('Generate') and a resource ('Composition Plan'), so the basic action is inferable, but 'Composition Plan' is domain jargon that is never defined. More importantly, there is no differentiation from the closely related siblings compose_detailed, compose_detailed_stream, and stream_compose, leaving the agent unable to tell which entry point it wants.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description supplies no when-to-use guidance, no prerequisites, and no exclusions. Given that three sibling tools (compose_detailed, compose_detailed_stream, stream_compose) appear to cover the same composition-planterritory, the absence of any routing hint is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

convert_chapter_endpointB

Convert Chapter Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
chapter_idYesThe ID of the chapter.
project_idYesThe ID of the Studio project.

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true. The description adds one useful behavioral fact beyond annotations: the operation consumes ElevenLabs credits. However, it does not explain what conversion does, whether it is long-running, or whether it requires specific account state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence fragment with no filler, and the key information is front-loaded. It is slightly awkward grammatically, but it does not waste space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-readonly, open-world, credit-consuming mutation tool, the description is very thin. It does not explain prerequisites, output, expected result, or what happens to the chapter, and the absence of an output schema leaves the agent without enough context to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both chapter_id and project_id are already documented in the input schema. The description adds no parameter-level meaning beyond what the schema provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource (Convert Chapter) and adds that the operation spends ElevenLabs credits. It is clear enough to identify the action, but it does not distinguish this tool from the sibling convert_project_endpoint or explain what the conversion produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no when-not-to-use guidance, and no mention of the alternative convert_project_endpoint. The agent must infer usage context entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

convert_project_endpointC

Convert Studio Project Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesThe ID of the Studio project.

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish non-read-only, non-idempotent, open-world behavior, and the description usefully adds the credit-consumption cost trait that annotations do not cover. It does not disclose what conversion produces, whether it is reversible, or whether it can be re-run without incurring further spend.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler and the key cost fact included. It reads as slightly truncated/awkward ('Convert Studio Project Spends ElevenLabs credits'), which costs it the top mark, but it is front-loaded and economical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent, credit-consuming, open-world operation with no output schema, the description should tell the agent what the conversion yields and what to expect afterward. It omits the outcome, the cost magnitude, and any post-conditions, leaving the agent unable to predict the effect of the call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single project_id parameter is fully documented in the schema, so the baseline of 3 applies. The description adds no format, sourcing, or constraint detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a verb ('Convert') and a resource ('Studio Project'), so the general action is identifiable, but it never states what the project is converted into or what the result is. The sibling convert_chapter_endpoint suggests a family of conversion operations, yet nothing here explains how this one differs (project-level vs chapter-level).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to invoke this versus alternatives such as convert_chapter_endpoint, edit_project, or get_project_by_id. The only usage-adjacent signal is that credits are consumed, which the agent must infer as a caution rather than being told when or when not to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_agent_conversation_ticket_routeC

Create Agent Conversation Ticket

ParametersJSON Schema
NameRequiredDescriptionDefault
qa_commentNo
turn_commentsNoOptional turn-level comments on what went wrong.
conversation_idYesConversation this ticket is about.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose that this is a non-read-only, open-world, non-idempotent, non-destructive operation. The description adds no further behavioral context such as authorization needs, side effects, or what a created ticket entails, leaving the annotations as the only behavioral signal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but it is under-specified rather than concise. It does not front-load any useful detail beyond the tool name, so the brevity reflects missing information rather than efficient communication.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create tool with three parameters, no output schema, and only partial schema coverage, the description is far too thin. Annotations provide some safety context, but the description does not help an agent understand ticket semantics, required context, or how this differs from related ticket endpoints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 67% and the description provides no parameter meaning at all. The required conversation_id is documented in the schema, but qa_comment has no description in either the schema or the tool description, so the parameter semantics remain incomplete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create Agent Conversation Ticket' restates the tool name/title rather than adding specific scope or differentiation. It identifies a create operation on a ticket resource but does not distinguish this from related siblings such as create_manual_agent_ticket_route, list_agent_conversation_tickets_route, or update_agent_conversation_ticket_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or alternatives are given. The description offers no indication of when an agent should choose this tool over other ticket-related or create-related siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_agent_deployment_routeC

Create Or Update Deployments

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.
deployment_requestYes

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the operation is not read-only, not idempotent, open-world, and non-destructive. The description adds only 'Create Or Update Deployments' and does not explain side effects, required permissions, or what an update changes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely terse and essentially restates the title without adding context. It is under-specified rather than usefully concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with a nested request object and no output schema, the description is inadequate. It omits what a deployment request contains, how agent_id is used, and what behavior to expect when creating or updating deployments.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%, and the description provides no parameter meaning at all. The nested deployment_request structure and required agent_id are left to the schema, which itself does not fully document the deployment request format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb and resource ('Create Or Update Deployments'), but it omits the agent scope and route nature conveyed by the tool name. It does not distinguish this tool from the many other creation/update siblings, leaving the agent to infer what a deployment route is.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. It does not explain when to create versus update, nor does it name related tools or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_agent_draft_routeD

Create Agent Draft

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesName for the draft
tagsNo
agent_idYesThe id of an agent. This is returned on agent creation.
workflowYes
branch_idYesThe ID of the agent branch to use
platform_settingsYesPlatform settings for the draft
conversation_configYesConversation config for the draft

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true. The description adds no behavioral context such as required permissions, what a draft represents, side effects, or whether the operation is reversible. It does not contradict the annotations, but it also adds nothing beyond them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short and front-loaded, but the single phrase does not earn its place by adding information. This is under-specification rather than useful conciseness, similar to the low-calibration example that scored 2 on this dimension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex create operation with 6 required parameters, nested objects, and no output schema, the description is completely inadequate. It omits any explanation of what a draft route is, what the nested configs represent, or what the agent should expect after calling the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 7 parameters with 71% description coverage, including some nested object descriptions. The description adds no parameter meaning, syntax guidance, or clarifications for required fields like workflow, platform_settings, or conversation_config. With moderate schema coverage and no descriptive compensation, this is a clear gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Create Agent Draft" restates the tool name/title without adding specificity, scope, or differentiation from siblings such as create_agent_route, create_branch_route, or delete_agent_draft_route. It does communicate a verb and a resource, but only at the level of a tautology, which fits the rubric's score of 2.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and no alternatives. It does not explain how this differs from creating an actual agent via create_agent_route or why one would create a draft instead. There is no explicit or implied usage context beyond the literal name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_agent_response_test_routeD

Create Agent Response Test

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
typeNo
environmentNo
chat_historyNo
evaluation_modelNo
failure_examplesNoNon-empty list of example responses that should be considered failures
parent_folder_idNo
success_examplesNoNon-empty list of example responses that should be considered successful
tool_mock_configNoSimulation/preview-side config: tools are identified by IDs, resolved to names at runtime.
dynamic_variablesNoDynamic variables to replace in the agent config during testing
success_conditionNo
success_conditionsNoList of prompts that evaluate whether the simulation was successful. If provided, all criteria are evaluated and merged into a final result. Capped at the maximum number of evaluation criteria.
simulation_scenarioNoDescription of the simulation scenario and user persona for simulation tests.
tool_mock_overridesNoTest-specific response mocks, keyed by tool ID. Applied ahead of the tool's shared mocks and only within this test. Only take effect for tools that are mocked (see tool_mock_config).
simulated_user_modelNo
simulation_max_turnsNoMaximum number of conversation turns for simulation tests.
tool_call_parametersNo
check_any_tool_matchesNo
simulation_environmentNo
from_conversation_metadataNo
conversation_initiation_sourceNoEnum representing the possible sources for conversation initiation.

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false, but the description adds nothing beyond them. With 21 parameters covering simulation config, tool mocking, and evaluation criteria, the description should disclose side effects, whether tests are persisted, and what happens on repeated calls — none of which is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short, but this is under-specification rather than conciseness. A single restated phrase cannot front-load meaningful information for a tool this complex.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A 21-parameter, nested-object, mutation tool with no output schema and no annotations covering its semantics needs a substantial description. The one-line title restatement leaves the agent with essentially no information needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 43% across 21 parameters, so the description carries a heavy burden — and it provides zero parameter information. Fields like tool_mock_overrides, success_conditions, simulation_scenario, and from_conversation_metadata are complex and nested, yet the description explains none of them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Create Agent Response Test" is a verbatim restatement of the tool name/title, which is the definition of tautology. It does not specify what an 'agent response test' is, what it operates on, or how it differs from sibling tools like update_agent_response_test_route or create_agent_test_folder_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance whatsoever on when to use this tool versus the 100+ siblings, no prerequisites, no mention of the agent or workspace context required. The agent must infer everything from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_agent_routeD

Create Agent

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
tagsNo
workflowNo
enable_versioningNoDeprecated: all agents are versioned. This parameter is ignored.
platform_settingsNo
conversation_configYes

TDQS

D1.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that, though 'Create' is at least consistent with a non-read-only, non-idempotent mutation rather than contradicting it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words are not 'concise' here but under-specified; the phrase is front-loaded but carries no information an agent can act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex creation tool with nested conversation_config, workflow, and platform_settings objects, no output schema, and no annotation-to-description elaboration, the definition is entirely inadequate for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 6 parameters, a required conversation_config, nested objects, and only 17% schema description coverage, the description supplies zero parameter meaning and does not compensate for the widespread undocumented fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is the two-word phrase "Create Agent", which merely restates the tool name and title. It states a verb and resource but adds no scope, and does not distinguish this from siblings such as create_agent_draft_route, duplicate_agent_route, or create_agent_deployment_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many other agent-creation siblings, nor any prerequisites, required permissions, or context about draft vs. live agents.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_agent_test_folder_routeC

Create Agent Test Folder

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe name of the folder to create
parent_folder_idNo

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond the title: no note on permission requirements, whether creating an existing folder errors, or nesting behavior with parent_folder_id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase, which is concise but under-informative rather than front-loaded with useful context. It is essentially the title repeated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with one undocumented optional parameter, no output schema, and annotations that cover basic safety, the description is incomplete. It lacks any detail about the operation's effect, prerequisites, or return behavior, leaving the agent reliant on schema and annotations alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: the required 'name' parameter is documented, but parent_folder_id has no description in either schema or description. The tool description provides no additional parameter semantics to compensate for the undocumented optional parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name ('Create Agent Test Folder') with a clear verb+resource, so the purpose is unambiguous: it creates a folder. However, it adds no scope or distinguishing detail beyond the name, and its relationship to siblings like create_folder_route or create_agent_draft_route is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as create_folder_route or the general folder endpoints. The agent must infer usage from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_audio_native_projectD

Creates Audio Native Enabled Project.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesProject name.
imageNo(Deprecated) Image URL used in the player. If not provided, default image set in the Player settings is used.
smallNo(Deprecated) Whether to use small player or not. If not provided, default value set in the Player settings is used.
titleNoTitle used in the player and inserted at the top of the uploaded article. If not provided, the default title set in the Player settings is used.
authorNoAuthor used in the player and inserted at the start of the uploaded article. If not provided, the default author set in the Player settings is used.
model_idNoTTS Model ID used in the player. If not provided, default model ID set in the Player settings is used.
voice_idNoVoice ID used to voice the content. If not provided, default voice ID set in the Player settings is used.
file_pathNoEither txt or HTML input file containing the article content. HTML should be formatted as follows '&lt;html&gt;&lt;body&gt;&lt;div&gt;&lt;p&gt;Your content&lt;/p&gt;&lt;h3&gt;More of your content&lt;/h3&gt;&lt;p&gt;Some more of your content&lt;/p&gt;&lt;/div&gt;&lt;/body&gt;&lt;/html&gt;' Local path
text_colorNoText color used in the player. If not provided, default text color set in the Player settings is used.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
auto_convertNoWhether to auto convert the project to audio or not.
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.
sessionizationNo(Deprecated) Specifies for how many minutes to persist the session across page reloads. If not provided, default sessionization set in the Player settings is used.
background_colorNoBackground color used in the player. If not provided, default background color set in the Player settings is used.
apply_text_normalizationNoThis parameter controls text normalization with four modes: 'auto', 'on', 'apply_english' and 'off'. When set to 'auto', the system will automatically decide whether to apply text normalization (e.g., spelling out numbers). With 'on', text normalization will always be applied, while with 'off', it w
pronunciation_dictionary_locatorsNoA list of pronunciation dictionary locators (pronunciation_dictionary_id, version_id) encoded as a list of JSON strings for pronunciation dictionaries to be applied to the text. A list of json encoded strings is required as adding projects may occur through formData as opposed to jsonBody. To specif

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=false, destructiveHint=false, openWorldHint=true and idempotentHint=false, giving the safety profile for free. The description adds nothing beyond that: it does not mention that content is converted to audio, whether auto_convert triggers generation, that callers cannot retry idempotently, or that a file upload is expected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single sentence with no wasted words, but it is under-specified rather than concise — the brevity comes at the cost of all useful information for a 16-parameter project-creation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex creation endpoint with 16 parameters, a required file content field, deprecated options, and no output schema, one four-word sentence is completely inadequate. Nothing about expected inputs, the resulting project state, or follow-up steps (e.g., content update endpoints) is communicated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 16 parameters including the deprecated ones and the file/file_base64/file_filename trio. The description contributes no parameter-level meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Creates Audio Native Enabled Project" is essentially a restatement of the tool name create_audio_native_project, with only a verb added. It gives no distinguishing detail against siblings such as add_project, create_podcast, or dubbing_project_create, so an agent cannot tell what an "Audio Native Enabled" project is or how it differs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. With siblings like add_project and create_podcast that also create projects, the absence of any routing information is a serious gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_auth_connectionD

Create Workspace Auth Connection

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
tokenNo
issuerNoJWT issuer (iss claim)
key_idNo
scopesNoOAuth2 scopes to request when exchanging JWT for access token
subjectNoJWT subject (sub claim)
audienceNoJWT audience (aud claim)
passwordNo
providerNo
usernameNo
algorithmNoJWT signing algorithm
auth_typeNo
client_idNo
token_urlNoToken endpoint URL for exchanging JWT for access token
client_keyNo
secret_keyNo
header_nameNoThe name of the header to use for authentication (e.g., 'x-api-key')
extra_paramsNoAdditional custom claims to include in the JWT
client_secretNo
ca_certificateNo
custom_headersNo
key_passphraseNo
client_certificateNo
expiration_secondsNoToken expiration time in seconds
basic_auth_in_headerNoIf True, send client credentials in Authorization header as Basic Auth instead of request body
token_response_fieldNoToken field to extract from the token endpoint response.

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false and destructiveHint=false, which tells the agent this is a non-idempotent write reaching external systems. The description adds nothing on top of that: it does not mention that credentials or secret material are being persisted, that missing parameters may cause failures, or that repeats will create duplicates (non-idempotent). No contradiction, but no added value either.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single short phrase is technically concise and front-loaded, but this is under-specification rather than efficiency. Given the complexity of the operation, the brevity is a defect, not a virtue.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 26-parameter, zero-required mutation tool with no output schema and only partial annotation coverage, this description is completely inadequate. It cannot help an agent decide which of the mutually exclusive credential parameters (token, username/password, client_id/client_secret, secret_key) to use for a given auth_type.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 26 parameters and only 42% schema description coverage, the description carries a heavy compensation burden, yet it names zero parameters. Fields like auth_type, provider, token_url, and the many credential options are left entirely undocumented in prose, so an agent must guess which fields to supply for a given auth method.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create Workspace Auth Connection' is essentially a restatement of the tool name create_auth_connection with the word 'Workspace' added. It names a verb and a resource, but adds no distinguishing scope and does nothing to separate it from siblings like update_auth_connection, delete_auth_connection, or list_auth_connections. This is tautological rather than clarifying.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance whatsoever on when to use this tool, when to prefer update_auth_connection, or what prerequisites (credentials, provider selection, scopes) are needed. The description gives the agent no routing or context information at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_batch_callC

Submit A Batch Call Request.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes
timezoneNo
branch_idNo
call_nameYes
recipientsYes
environmentNo
whatsapp_paramsNo
scheduled_time_unixNo
agent_phone_number_idNo
telephony_call_configNo
target_concurrency_limitNo

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, covering the basic safety profile. The description adds no behavioral context beyond 'submit'—it does not mention side effects, required permissions, return behavior, or whether the batch is queued immediately.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise in form but severely under-specified for a tool with 11 parameters and nested configuration. The sentence does not earn its place with useful information, making this under-specification rather than effective conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex batch-call creation tool with 11 parameters, nested objects, no output schema, and only minimal annotations, the one-line description is completely inadequate. An agent cannot reliably determine required inputs, scheduling behavior, or provider-specific options from the description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for 11 parameters, including required ones like agent_id and recipients plus nested objects such as telephony_call_config and whatsapp_params. The description provides no parameter meaning whatsoever, failing to compensate for the complete lack of schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb 'Submit' and resource 'Batch Call Request', so the general purpose is identifiable. However, it offers no detail about what a batch call is, how it differs from sibling operations like cancel_batch_call or retry_batch_call, and reads close to a restatement of the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no guidance on when to use this tool versus alternatives such as cancel_batch_call, get_batch_call, or get_workspace_batch_calls. No prerequisites, context, or exclusions are mentioned, leaving the agent to infer usage entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_branch_routeC

Create A New Branch

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesName of the branch. It is unique within the agent.
agent_idYesThe id of an agent. This is returned on agent creation.
workflowNo
descriptionYesDescription for the branch
include_draftNoWhen true, the new branch uses the caller's draft procedure set instead of the branch tip. Requires parent_version_id to be the branch tip.
parent_version_idYesID of the version to branch from
platform_settingsNo
conversation_configNo

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is known. The description adds no behavioral context whatsoever — nothing about branch creation semantics, whether the parent version must be the branch tip, or what include_draft implied constraints exist, even though the schema hints these constraints matter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single four-word sentence, which is brief but not concise in the useful sense — it omits everything an agent would need and contains no information beyond the tool name. It is under-specification rather than economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an eight-parameter mutation tool with no output schema and several non-trivial fields (parent_version_id tip constraint, include_draft, workflow definition), a bare title-like sentence is wholly inadequate. Nothing about the creation flow, constraints, or return value is conveyed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 63%, and the description supplies zero parameter information. Eight parameters include non-obvious ones like include_draft and platform_settings whose semantics the agent must glean purely from the schema; the description does nothing to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Create A New Branch" states a specific verb and resource, so the basic operation is identifiable. However, it offers no differentiation from siblings such as get_branch_route, update_branch_route, merge_branch_into_target, or rebase_branch_onto_main, leaving the agent to infer from the name alone which branch operation is intended.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, what prerequisites exist (agent_id, parent_version_id), or how it relates to the merge/rebase branch siblings. The agent gets no context for selecting it over alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_clipB

Create A Segment For The Speaker Spends ElevenLabs credits. Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
textNo
end_timeYes
dubbing_idYesID of the dubbing project.
speaker_idYesID of the speaker.
start_timeYes
translationsNo

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnly=false, destructive=false, idempotent=false, and openWorld=true, so safety basics are covered. The description adds meaningful behavioral context beyond the annotations by disclosing that the operation spends ElevenLabs credits and is deprecated upstream. It does not, however, explain permissions, side effects on existing segments, or response behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loads the core action. The deprecation and credit notes are valuable and do not feel padded. The awkward capitalization slightly harms readability but does not impede selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter mutation tool with no output schema and low schema description coverage, the description is too thin. It does not explain how the required time and text parameters relate, what the translations parameter does, or what happens after creation. The deprecation and cost notes help, but key invocation details remain missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, and the description does not compensate. It implies a speaker and a segment but gives no meaning for text, translations, start_time, or end_time. An agent must rely mostly on the partially documented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: creating a segment for a speaker. It also names the cost model (ElevenLabs credits) and lifecycle status (deprecated upstream). However, it does not distinguish the tool from siblings such as dubbing_transcript_segment_add or create_speaker, and the clip/segment terminology is not reconciled.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes that the tool is deprecated upstream, which implicitly warns against use, but it never states when to use this tool instead of alternatives. It also provides no explicit exclusions or replacement tool guidance. This falls short of useful usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_conversation_tag_routeC

Create Conversation Tag

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYesDisplay title of the tag.
descriptionNo

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true), so the description only needed to add context such as the scope of the created tag or error behavior on duplicate titles. It provides none, leaving the description with zero added value beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single phrase has no padding, but it is under-specified rather than concise. Brevity here comes from omitting required information, not from efficient expression.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation with no output schema and half its parameters undocumented, the description should explain what a conversation tag is and what happens on creation. It supplies none of this, leaving the definition materially incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50% – 'title' is documented in the schema, but 'description' has no description anywhere. The tool description adds nothing to compensate, so an agent gets no meaning for half the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create Conversation Tag' merely restates the tool name and title verbatim, adding no scope, target, or distinguishing detail. It does not differentiate the tool from siblings like update_conversation_tag_route, list_conversation_tags_route, or delete_conversation_tag_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives. Nothing states that this creates a brand-new tag rather than updating an existing one, nor any prerequisite such as uniqueness of the title.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_crawl_job_routeD

Create Crawl Job

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to a page of documentation that the agent will have access to in order to interact with users.
patternNo
max_depthNoMaximum depth for crawling (1-5), defaults to 3.
max_pagesNoMaximum number of pages to crawl (1-10,000), defaults to 1000.
auto_removeNoWhether to automatically remove the document if the URL becomes unavailable. Only applicable when auto-sync is enabled.
sitemap_urlsNo
enable_auto_syncNoWhether to enable auto-sync for this URL document.
parent_folder_idNo
minimum_frequency_daysNo

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose the safety profile (readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false). The description adds nothing beyond that: it does not say the crawl runs asynchronously, that it returns a job identifier, or how auto_remove/auto-sync behavior affects the created resource.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The three-word phrase is technically short, but it is under-specified rather than concise — it omits information the agent needs rather than front-loading useful content. Brevity here reflects a missing description, not efficient writing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter mutation tool with no output schema and partial field documentation, the description is completely inadequate; an agent cannot form a correct call from it alone. Nothing about the crawl lifecycle, required URL semantics, or parameter defaults is conveyed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 9 parameters and only 56% schema description coverage, several parameters (pattern, sitemap_urls, parent_folder_id, minimum_frequency_days) have no documented meaning in either location. The description provides no compensating detail, so it fails to cover the schema's gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is just "Create Crawl Job", which restates the tool name with no additional scope, target, or distinguishing detail. It states a verb and resource, but an agent learns nothing beyond the identifier itself, and it is not differentiated from siblings like create_url_document_route or list_crawl_jobs_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to create a crawl job versus the sibling create_url_document_route, nor any mention of prerequisites (URL required, auto-sync interactions) or follow-up tools like get_crawl_job_route or cancel_crawl_job_route. Usage is left entirely to inference from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_dubbingC

Dub A Video Or An Audio File Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoThe mode in which to run this Dubbing job. Defaults to automatic, use manual if specifically providing a CSV transcript to use. Note that manual mode is experimental and production use is strongly discouraged.
nameNoName of the dubbing project.
csv_fpsNoFrames per second to use when parsing a CSV file for dubbing. If not provided, FPS will be inferred from timecodes.
end_timeNoEnd time of the source video/audio file.
file_pathNoA list of file paths to audio recordings intended for voice cloning Local path.
watermarkNoWhether to apply watermark to the output video.
source_urlNoURL of the source video/audio file.
start_timeNoStart time of the source video/audio file.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
source_langNoSource language. Expects a valid iso639-1 or iso639-3 language code.
target_langNoThe Target language to dub the content into. Expects a valid iso639-1 or iso639-3 language code.
num_speakersNoNumber of speakers to use for the dubbing. Set to 0 to automatically detect the number of speakers
csv_file_pathNoCSV file containing transcription/translation metadata Local path.
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.
target_accentNo[Experimental] An accent to apply when selecting voices from the library and to use to inform translation of the dialect to prefer.
dubbing_studioNoWhether to prepare dub for edits in dubbing studio or edits as a dubbing resource.
csv_file_base64NoBase64 contents for "csv_file". Use this when the server cannot read your local disk.
csv_file_filenameNoFilename to send for "csv_file". Some endpoints infer the audio format from it.
highest_resolutionNoWhether to use the highest resolution available.
use_profanity_filterNo[BETA] Whether transcripts should have profanities censored with the words '[censored]'
disable_voice_cloningNoInstead of using a voice clone in dubbing, use a similar voice from the ElevenLabs Voice Library. Voices used from the library will contribute towards a workspace's custom voices limit, and if there aren't enough available slots the dub will fail. Using this feature requires the caller to have the '
drop_background_audioNoAn advanced setting. Whether to drop background audio from the final dub. This can improve dub quality where it's known that audio shouldn't have a background track such as for speeches or monologues.
background_audio_file_pathNoFor use only with csv input Local path.
foreground_audio_file_pathNoFor use only with csv input Local path.
background_audio_file_base64NoBase64 contents for "background_audio_file". Use this when the server cannot read your local disk.
foreground_audio_file_base64NoBase64 contents for "foreground_audio_file". Use this when the server cannot read your local disk.
background_audio_file_filenameNoFilename to send for "background_audio_file". Some endpoints infer the audio format from it.
foreground_audio_file_filenameNoFilename to send for "foreground_audio_file". Some endpoints infer the audio format from it.

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true, so the safety profile is covered. The description adds one genuinely useful behavior trait beyond annotations — that the call consumes ElevenLabs credits — but says nothing about the async nature of the job, required languages, or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is very short and the credit-cost clause carries real information, but the phrasing is a run-on fragment with inconsistent capitalization. It is tersely front-loaded yet reads more like a title than a structured statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 28-parameter dubbing job with zero required fields and no output schema, the description is far too thin: it omits how to supply input (URL/base64/local path), the language-code requirements, the automatic vs manual mode trade-off, and the credit-cost magnitude. Only the credit-consumption warning is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 28 parameters are already documented in the schema. The description adds no parameter meaning (no explanation of source_url vs file_path vs file_base64 selection, no language-code requirements), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ("Dub a video or an audio file"), so an agent can tell what the tool performs. However, it does not differentiate from close siblings like `dub` or `dubbing_project_create`, and the garbled Title Case phrasing with the run-on "Spends ElevenLabs credits." makes the scope slightly ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no routing between this and the sibling `dub`/`dubbing_project_create` tools. The credit note hints at cost but does not tell the agent when invoking this is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_environment_variableD

Create Environment Variable

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNo
labelNoUnique label for the environment variable.
valuesNoEnvironment-specific auth connection references. Must include 'production' key.

TDQS

D1.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose that this is a non-read-only, non-idempotent, open-world, non-destructive write. The description adds nothing beyond that — no mention of required permissions, what happens on duplicate labels, or how auth connection references are created.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single phrase is technically brief but it is under-specification rather than conciseness; it omits any front-loaded useful information about the operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with a nested 'values' object, an undocumented 'type' parameter, no required-parameter guidance, and no output schema, the description is completely inadequate — an agent cannot safely call this correctly from the given text.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, with 'label' and 'values' documented but the 'type' parameter carrying no description and no enum. The description provides zero parameter information, so it fails to compensate for the undocumented 'type' field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create Environment Variable' merely restates the tool name and title verbatim, adding no scope, target, or distinguishing detail. It does not differentiate this tool from siblings like update_environment_variable, get_environment_variable, or list_environment_variables.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, what prerequisites exist, or how it relates to alternatives such as update_environment_variable. The agent is left to infer everything from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_file_document_routeC

Create File Document

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoA custom, human-readable name for the document.
file_pathNoDocumentation that the agent will have access to in order to interact with users. Local path. Required for this call.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.
parent_folder_idNoIf set, the created document or folder will be placed inside the given folder.

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations (readOnlyHint=false, idempotentHint=false, openWorldHint=true, destructiveHint=false) already convey that this is a non-idempotent write in an open world. The description adds no behavioral context beyond that – no mention of whether the base64/path inputs are mutually exclusive, what errors occur on failure, or what is created.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words with zero waste, but this is under-specification rather than conciseness. The description is front-loaded only in the trivial sense that there is nothing behind it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation route that accepts either a local path or base64 content, with no required parameters and no output schema, the description supplies none of the disambiguation an agent needs. It is indistinguishable in intent from several sibling create_*_document_route tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters (name, file_path, file_base64, file_filename, parent_folder_id) are already documented in the schema. The description contributes nothing extra about them, so the baseline 3 for full schema coverage applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Create File Document" only restates the tool name/title with no additional specification. It gives a verb and a resource but nothing that separates it from siblings like create_text_document_route, create_url_document_route, or create_folder_route, which all create documents in the same workspace.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisite, and no mention of the sibling routes an agent should choose between. The only signal about selection comes from the name itself, leaving the agent to guess how a "file document" differs from a text or URL document.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_finetuneB

Create Music Finetune Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesName for the finetune (5-200 characters).
tagsNoTags to associate with the finetune.
model_idNo
visibilityNoFinetune visibility. Only 'private' and 'workspace' can be set.
files_pathsNoAudio files to train on. Local paths.
primary_genreYesPrimary musical genre of the finetune.
files_filenamesNoFilenames to send for "files". Some endpoints infer the audio format from them.
files_base64_listNoBase64 contents for "files", one entry per file. Use this when the server cannot read your local disk.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, and openWorldHint=true, so the mutation profile is covered. The description adds one genuinely useful behavioral fact — that the call consumes ElevenLabs credits — but says nothing about training time, data requirements, or irreversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short and front-loaded with the action, but it is a single run-on fragment ('Create Music Finetune Spends ElevenLabs credits.') with awkward capitalization and a missing clause boundary, which slightly hurts readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter training operation with no output schema and no annotations beyond the standard hints, the description is thin — it omits the mutual-exclusivity of the three file-input parameters and any notion of training/cost implications. The rich schema partly compensates.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 88%, so the schema already documents nearly all eight parameters, including file-input variants. The description adds no parameter-level meaning (e.g. that exactly one of files_paths/files_base64_list/files_filenames must be supplied). Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ('Create') and resource ('Music Finetune'), which separates it from get_finetune, update_finetune, and delete_finetune in the sibling list. It does not explicitly name those alternatives, but the action+resource pairing is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to create a finetune versus using an existing model, no prerequisites (training data required), and no mention of the sibling read/update/delete operations. The only usage-adjacent fact is the credit cost.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_folder_routeD

Create Folder

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesA custom, human-readable name for the document.
auto_removeNoWhether to automatically remove the document if the URL becomes unavailable. Only applicable when auto-sync is enabled.
enable_auto_syncNoWhether to enable auto-sync for this URL document.
parent_folder_idNo
minimum_frequency_daysNo

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description adds nothing beyond them: it does not explain permissions, side effects, idempotency, or what folder creation entails. With annotations present, the bar is lower, but zero added context still earns a low score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two words, which is concise but severely under-specified for a five-parameter mutation tool. It is front-loaded but does not earn its place by conveying useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter folder-creation tool with no output schema and only partial schema descriptions, the description is completely inadequate. It omits any behavioral, parameter, or usage context an agent would need to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 60%, leaving some parameter semantics undocumented, and the description adds no parameter meaning at all. It does not clarify the required name, parent_folder_id, auto_remove, enable_auto_sync, or minimum_frequency_days, so it fails to compensate for the partial schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create Folder' is essentially a restatement of the tool name and title. While it does indicate a create-folder operation, it offers no additional specificity, no distinction from other create_* siblings, and no scope beyond what is already in the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites, and no exclusions. The description provides no context for selection among the many sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_image_generationD

Create Image Generation

ParametersJSON Schema
NameRequiredDescriptionDefault
maskNo
seedNo
imagesNoUp to 10 reference images to edit or draw from.
promptNoA text description of the image to generate.
qualityNoThe quality of the output image.
webhookNo
model_idNoThe model to use for the generation.
backgroundNoThe background of the output image. With `auto`, the model picks the background that suits the image.
resolutionNoThe resolution of the output image.
aspect_ratioNoThe aspect ratio of the output image. With `auto`, the model picks an aspect ratio based on the inputs.

TDQS

D1.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false, so the safety profile is covered structurally. The description adds nothing beyond that — no mention of async behavior, cost implications, model selection consequences, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short but not in a good way; this is under-specification rather than conciseness. There is nothing front-loaded because there is nothing substantive to load.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter generative tool with no output schema and an async webhook parameter, the definition is completely inadequate. An agent cannot know required inputs, output format, or completion semantics from this description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 70%, below the 80% baseline, so the description should compensate for the undocumented parameters (seed, model_id, images items) — but it provides zero parameter meaning. An agent must rely entirely on the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create Image Generation' merely restates the tool name and title verbatim, providing no additional specificity beyond the tautology. It does not clarify scope, distinguish from siblings like create_video_generation or text_to_voice_design, or state what kind of image generation is performed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance whatsoever about when to use this tool versus alternatives. With many sibling creation tools (create_video_generation, sound_generation, generate), the agent receives no routing signal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_manual_agent_ticket_routeD

Create Manual Agent Ticket

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes
qa_commentYesWhat the ticket is about, e.g. a follow-up task for the agent. This is shown as the ticket title.

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=false, openWorldHint=true and destructiveHint=false, so the agent knows this is a non-idempotent write. The description adds nothing beyond that — no mention of side effects, required permissions, or what a 'manual' ticket means versus an auto-created one. With annotations carrying the safety profile, a low score reflects the complete absence of added behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single fragment is short but not concise in any useful sense — it is under-specification, not economy. There is no front-loaded scope or actionable content for the agent to act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, non-idempotent ticket-creation tool with no output schema and only 50% schema coverage, the description provides nothing: no return behavior, no relation to the other ticket routes, no note on what 'manual' implies. An agent cannot safely place this call from the definition alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: qa_comment is documented but agent_id has no description in either schema or description text. The description supplies no parameter meaning at all, so it fails to compensate for the undocumented agent_id. Baseline 3 is not reachable given the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create Manual Agent Ticket' is a verbatim restatement of the tool name, offering no scope, object, or distinguishing detail. It does not differentiate this tool from near-identical siblings such as create_agent_conversation_ticket_route or create_agent_response_test_route. This is a tautology rather than an explanation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when this tool should be used versus alternatives, nor any prerequisite or context. The crowded 'ticket' sibling family makes this omission costly. No misleading claim is made, but no guidance whatsoever is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_mcp_server_routeC

Create Mcp Server

ParametersJSON Schema
NameRequiredDescriptionDefault
configYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, covering the basic safety profile. The description adds nothing beyond that — no word on whether creation is persistent, what happens on duplicate names, or auth requirements for the config payload.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but only because it is under-specified rather than efficiently phrased. A three-word phrase cannot be considered well-structured for a tool with a complex nested config object.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool whose sole parameter is a large nested config object, the description supplies none of the context an agent needs — not the effect of creation, nor the required sub-fields, nor the response. Annotations cover the safety hints but not the operational context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single top-level parameter 'config' has no description, and the tool description provides no explanation of it either. The nested properties do carry their own schema descriptions, but the description does nothing to compensate for the undocumented top-level wrapper.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Create Mcp Server" essentially restates the tool name and title without adding scope, resource location, or any distinguishing detail from siblings like create_auth_connection or update_mcp_server_config_route. It is a tautology rather than a clarifying statement of purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as update_mcp_server_config_route (for changing an existing server) or list_mcp_servers_route. No prerequisites, no context about where the server gets registered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_phone_number_routeD

Import Phone Number

ParametersJSON Schema
NameRequiredDescriptionDefault
sidNoTwilio Account SID (starts with `AC`) or API Key SID (starts with `SK`)
labelNoLabel for the phone number
tokenNoSecret paired with `sid`: the Account Auth Token for an Account SID, or the API Key Secret for an API Key SID
app_idNoExotel applet identifier used in Calls/connect
api_keyNoExotel API Key
agent_idNo
providerNo
api_tokenNoExotel API Token
applet_urlNo
enable_smsNoRoute inbound SMS to ElevenLabs. On by default; set to false to skip SMS configuration for numbers that don't support it.
account_sidNoExotel Account SID
phone_numberNoPhone number
api_subdomainNo
region_configNo
supports_inboundNoThis field is deprecated and will be removed in the future. Whether this phone number supports inbound calls
supports_outboundNoThis field is deprecated and will be removed in the future. Whether this phone number supports outbound calls
account_auth_tokenNo
inbound_trunk_configNo
outbound_trunk_configNo

TDQS

D1.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false and openWorldHint=true, so the safety profile is carried by structured data and the description does not contradict it. However, the description adds nothing on top: it does not mention that credentials (Account SID/token, Exotel keys, region config) are required, that repeated calls are non-idempotent, or what side effects occur on the workspace.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words are technically waste-free, but this is under-specification rather than conciseness: a 19-parameter, multi-provider onboarding tool cannot be adequately framed in a bare noun phrase. There is no front-loaded scope or context for the agent to anchor on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 19 parameters, zero required fields, a provider-variant schema (Twilio vs Exotel vs SIP trunk, plus region and media-encryption options), and no output schema, the description supplies none of the orientation an agent needs. It does not explain that this provisions a phone number route, which credential set applies, or what happens on success or failure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description mentions no parameters at all despite 19 of them and only 58% schema description coverage. Key choices such as provider, enable_sms, region_config, inbound/outbound trunk config, and the deprecated supports_inbound/supports_outbound flags are left entirely to the schema, with several fields (agent_id, provider, applet_url, api_subdomain, account_auth_token) undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Import Phone Number" essentially restates the tool name (create_phone_number_route) with a different verb; it is close to a tautology. It gives no indication of scope, such as which provider (Twilio, Exotel, SIP trunk) is being onboarded or what 'import' means versus a plain create, so an agent cannot distinguish it from siblings like update_phone_number_route or list_phone_numbers_route on the basis of the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance whatsoever: no mention of prerequisites, of the alternative tools (update_phone_number_route, delete_phone_number_route, list_phone_numbers_route), or of the conditions that select this route. Nothing in the description helps an agent decide between this and any sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_podcastC

Create Podcast Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesThe type of podcast to generate. Can be 'conversation', an interaction between two voices, or 'bulletin', a monologue.
introNo
outroNo
sourceYesThe source content for the Podcast.
languageNo
model_idYesThe ID of the model to be used for this Studio project, you can query GET /v1/models to list all available models.
highlightsNo
callback_urlNo
duration_scaleNoDuration of the generated podcast. Must be one of: short - produces podcasts shorter than 3 minutes. default - produces podcasts roughly between 3-7 minutes. long - produces podcasts longer than 7 minutes.
quality_presetNo
safety-identifierNoUsed for moderation. Your workspace must be allowlisted to use this feature.
instructions_promptNo
apply_text_normalizationNo

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-readonly, non-idempotent, open-world write, so the safety profile is covered. The description adds one genuinely useful behavioral fact — that the call consumes ElevenLabs credits — but omits key traits such as async/callback behavior (callback_url exists), moderation allowlisting, or long-running generation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short line, so nothing is wasted, but it reads as a run-on fragment ('Create Podcast Spends ElevenLabs credits') with no punctuation separating the action from the cost note. The brevity is a symptom of under-specification rather than disciplined economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter, nested-object, credit-consuming generation tool with no output schema and only 38% schema coverage, this description is far too thin. It should at minimum explain the mode/source requirement and the callback flow, none of which appear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 13 parameters and only 38% schema description coverage, the description needs to compensate for undocumented fields like intro, outro, language, instructions_prompt, and callback_url. It provides zero parameter semantics, leaving several required and optional inputs unexplained in both the description and schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description conveys a verb+resource ('Create Podcast'), matching the tool name, but adds nothing to differentiate it from siblings like create_audio_native_project or create_video_generation. The only non-obvious content is the credit-spend note, leaving the core purpose essentially a restatement of the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites mentioned, and no indication of required inputs. The credit-cost sentence hints at a gating factor but does not tell the agent when to prefer this tool over other content-creation siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_procedure_routeD

Create Procedure

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoProcedure name
typeNo
contentNoInitial procedure content
triggerNo
agent_idYesAgent ID to get the procedure draft from
branch_idYesBranch ID to get the procedure draft from
folder_parent_idNo

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already state readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, giving the safety profile. The description adds no behavioral context beyond those annotations, such as permission needs, what is created from the supplied agent_id/branch_id draft, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but it is under-specified rather than concise. It lacks the structure needed to convey purpose, requirements, or behavior for a seven-parameter creation route.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent, open-world creation tool with seven parameters and no output schema, the description is completely inadequate. It omits the draft-source role of the required agent_id and branch_id, the meaning of creation, and any usage context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers only 57% of the seven parameters with descriptions, and the description contributes no additional parameter meaning. Undocumented fields such as type, trigger, and folder_parent_id receive no explanation from either source.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Create Procedure" simply restates the tool name create_procedure_route and adds no scope, object, or distinguishing detail. It does not differentiate this tool from sibling route-creation tools such as create_agent_draft_route or create_branch_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites, and no mention of the draft/branch context implied by the required parameters. The agent is left to infer everything from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_pvc_voiceD

Create Pvc Voice

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe name that identifies this voice. This will be displayed in the dropdown of the website.
labelsNo
languageYesLanguage used in the samples.
descriptionNo

TDQS

D1.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, which tells the agent this is a non-destructive write operation. The description adds no behavioral context beyond the word 'Create,' such as authentication needs, whether the operation is asynchronous, or what happens to existing voices. It does not contradict the annotations, but it adds essentially no transparency value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but its brevity reflects under-specification rather than efficient communication. It does not front-load any useful information beyond the tool name, so there is no structure or content worth preserving.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that creates a PVC voice with four parameters (two required) and only 50% schema coverage, no output schema, and no usage guidance, the description is completely inadequate. It omits prerequisites, expected behavior, return information, and any distinction from closely related sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description says nothing about the four parameters. Schema description coverage is only 50%: name and language have descriptions in the schema, but labels and description do not. Since half of the parameters are undocumented in both the schema and the description, the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description merely restates the tool name as 'Create Pvc Voice' and provides no additional specificity about what a PVC voice is or how it differs from sibling tools like create_voice, add_voice, or edit_pvc_voice. It is a tautology that gives an agent little beyond the tool name itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. The description does not mention prerequisites, when to choose create_pvc_voice over create_voice or add_pvc_voice_samples, or any contextual conditions for invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_secret_routeC

Create Convai Workspace Secret

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
typeYes
valueYes

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, openWorldHint=true, and destructiveHint=false, covering the basic safety profile. The description adds nothing beyond that — no mention of required auth, what a 'type' entails, whether value is stored encrypted, or what happens on name collisions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is short but under-specified rather than concise; it omits all operational detail and wastes no words only because it says almost nothing. It is not misleading, but it does not earn its place as the sole guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool creating a workspace secret with three undocumented required parameters and no output schema, the description is far too thin. Annotations cover the safety profile, but the agent still lacks information on parameter meaning and any side effects or prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for three required parameters (name, type, value). The description does not compensate at all — it never mentions that 'type' selects a secret kind or that 'value' is the secret payload, so an agent gets no guidance on accepted type values or value format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb ('Create') and resource ('Secret') and scopes it to 'Convai Workspace', which distinguishes it from generic secret tools in other systems. However, it adds no detail about what a secret route is or how it differs from siblings like update_secret_route or create_environment_variable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as create_environment_variable or update_secret_route, nor any prerequisites, permissions, or conditions described. The agent is left to infer everything from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_service_accountD

Create Service Account

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
default_sharing_groupsNo

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the mutation and open-world nature is captured structurally. The description adds nothing further — no note on required permissions, what the created account can do, or whether default_sharing_groups grants access.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words that mirror the title, which is under-specification rather than conciseness. There is no front-loaded content because there is no content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutation tool with no output schema, undocumented parameters (including a nested free-form array), and no sibling differentiation is left entirely described by its name. An agent cannot know what this account is, what it can access, or what identifies it on return.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters, yet the description says nothing about 'name' or 'default_sharing_groups'. The meaning of default_sharing_groups — an array of free-form objects with additionalProperties unconstrained — is completely opaque and unmitigated by any text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create Service Account' is a verbatim restatement of the tool name and title, adding no information beyond what the identifier already conveys. It does not distinguish this tool from close siblings such as create_service_account_api_key or get_workspace_service_accounts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives is provided. The crowded sibling namespace (create_service_account_api_key, edit_service_account_api_key, delete_service_account_api_key) makes the absence of routing guidance costly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_service_account_api_keyD

Create Service Account Api Key

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
allowed_ipsNo
permissionsYesThe permissions of the XI API.
character_limitNo
service_account_user_idYes
third_party_disable_allowedNo

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, and the description does not contradict them. However, the description adds zero behavioral context beyond the annotations: it does not say whether the generated API key secret is returned only once, whether existing keys are affected, or how permissions are enforced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The four-word description is short, but this is under-specification rather than useful conciseness. There is no front-loaded detail, no structure, and no information beyond the tool name, so the brevity is a liability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a mutation tool with six parameters, a nested permissions object, three required fields, and no output schema. Given that complexity and the very low schema description coverage, the description is completely inadequate: an agent has no way to know required formats, the meaning of permissions, or what the call returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% (only the 'permissions' object has a description), and the description contributes no parameter meaning at all. Six parameters including a nested permissions object, allowed_ips, character_limit, and third_party_disable_allowed are left completely unexplained, so the description fails to compensate for the sparse schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is the tool name restated verbatim ('Create Service Account Api Key' with title 'Create Service Account Api Key'), making it a tautology. While it names a verb and resource, it provides no scope, no mention of the owning service account, and no differentiation from close siblings like create_service_account or edit_service_account_api_key.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as create_service_account (which likely creates the parent account) or edit_service_account_api_key. No prerequisites, no mention that a service_account_user_id must refer to an existing account, and no when-not-to-use conditions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_speakerC

Create A New Speaker Spends ElevenLabs credits. Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idNo
dubbing_idYesID of the dubbing project.
voice_styleNo
speaker_nameNo
voice_stabilityNo
voice_similarityNo

TDQS

C2.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, idempotentHint=false, openWorldHint=true), the description discloses two facts the agent could not infer: the call consumes ElevenLabs credits, and the endpoint is deprecated upstream. Both are decision-relevant. It still omits error/failure behavior, but with annotations covering the safety profile this is above the bar.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short, but the first sentence is a run-on fragment ('Create A New Speaker Spends ElevenLabs credits') whose missing punctuation blurs the action and the cost caveat together. The deprecation note is tacked on rather than front-loaded, which is the opposite of what an agent needs to see first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A six-parameter mutating tool with 17% schema coverage, no output schema and no annotation coverage of cost or lifecycle needs far more than two fragments. The description never explains the dubbing context, what a 'speaker' maps to, or how the five undocumented voice parameters behave.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 17% — only dubbing_id is documented, while voice_id, speaker_name, voice_style, voice_stability and voice_similarity are bare. The description compensates with nothing: no parameter names, no formats, no valid ranges for the numeric voice-tuning fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Create A New Speaker' gives a verb and a resource, so the basic action is identifiable, but 'speaker' is ambiguous in a catalog full of dubbing, voice and speaker-separation tools. The description never says this creates a dubbing-project speaker, nor does it distinguish itself from siblings like update_speaker or create_voice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Deprecated upstream' is the only usage signal, and it is not actionable: it doesn't say whether to avoid the tool, what replaces it, or under what conditions it still works. There is no when-to-use guidance and no alternative named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_speech_engineD

Create Speech Engine

ParametersJSON Schema
NameRequiredDescriptionDefault
asrNo
ttsNo
vadNo
nameNoName of the speech engine
tagsNoTags for categorization
turnNo
privacyNo
languageNoLanguage for the speech engine
overridesNo
call_limitsNo
conversationNo
speech_engineYes

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, which provides a basic safety profile. However, the description adds no behavioral context such as what gets created, whether the operation is reversible, authentication requirements, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At three words, it is not wasteful in length, but it is severely under-specified rather than concise. A helpful description would need more content to earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 12 parameters, deep nested objects, no output schema, and sparse schema coverage, the description is completely inadequate. It gives an agent no meaningful context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% across 12 parameters, including many complex nested objects. The description provides zero parameter information and therefore fails to compensate for the significant documentation gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a tautology: it merely restates the tool name/title as 'Create Speech Engine' and adds no specific scope, resource detail, or sibling differentiation. It does not distinguish itself from update_speech_engine, delete_speech_engine, or get_speech_engine.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or alternatives are provided. The agent must infer from the name alone that this creates a speech engine rather than updating or retrieving one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_text_document_routeD

Create Text Document

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
textYesText content to be added to the knowledge base.
parent_folder_idNo

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description adds no behavioral context beyond this, such as side effects, knowledge base insertion behavior, or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but it is under-specified rather than concise. It omits necessary context and does not earn its place as a useful tool definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter create operation with no output schema and only sparse annotations, the description is critically incomplete. It does not explain that the document is added to a knowledge base, the optionality or role of parent_folder_id, or any return behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, with only the 'text' parameter documented in the schema. The description provides no meaning for 'name' or 'parent_folder_id', so it fails to compensate for the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create Text Document' restates the tool name/title and provides no scope or distinction from sibling tools such as create_file_document_route or create_url_document_route. It states a verb and resource but is effectively a tautology rather than a differentiating purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. The description does not help an agent decide between this tool and the many other create/update document routes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_text_to_speech_generationD

Create Speech Generation

ParametersJSON Schema
NameRequiredDescriptionDefault
textNoThe text to synthesize into speech.
voiceNoThe ID of the voice to speak with.
webhookNo
model_idNoThe model to use for the generation.
language_codeNo
output_formatNoThe audio encoding of the output, as `codec_sampleRateHz_bitrateKbps`. `mp3_44100_192` requires the Creator tier or above.
voice_settingsNoOverrides for the voice's saved settings, applied to one generation.
pronunciation_dictionary_locatorsNoPronunciation dictionaries to apply to the text, in order of precedence. Up to 3.

TDQS

D1.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false, so the safety/mutation profile is covered. The description adds nothing: it does not mention that this is an async generation job, whether a webhook result is required, credit/cost implications, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short, but this is under-specification rather than conciseness; there is no structure or front-loaded detail to earn its brevity. The single phrase conveys no actionable information beyond the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter generation/creation tool with no output schema and no required-parameter guidance, the description is wholly inadequate. An agent cannot determine inputs, side effects, or result retrieval from it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Eight parameters with 75% schema description coverage; the schema documents text, voice, output_format, and webhook reasonably well. The description contributes zero parameter meaning, leaving the lower-coverage fields (language_code, model_id, voice_settings, pronunciation_dictionary_locators) without any supplementary context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Create Speech Generation" essentially restates the tool name (create_text_to_speech_generation) with no added specificity. It states a verb and resource but nothing about scope, output, or how it differs from the many sibling TTS tools (text_to_speech_full, text_to_speech_stream, text_to_dialogue, etc.).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no exclusions, and no reference to any alternative despite a crowded sibling landscape of text-to-speech and speech-to-speech tools. An agent has no basis to prefer this tool over the others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_url_document_routeD

Create Url Document

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to a page of documentation that the agent will have access to in order to interact with users.
nameNo
auto_removeNoWhether to automatically remove the document if the URL becomes unavailable. Only applicable when auto-sync is enabled.
enable_auto_syncNoWhether to enable auto-sync for this URL document.
parent_folder_idNo
minimum_frequency_daysNo

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, covering the safety profile. The description adds nothing beyond those annotations — no mention of auto-sync behavior, the external URL dependency, or what happens on failure. With the annotation bar low, some added context was expected and none is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a three-word fragment with no wasted words, but this reflects under-specification rather than effective conciseness. It fails to front-load any useful information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter creation tool with no output schema and only 50% schema description coverage, the description is completely inadequate. It conveys nothing about the auto-sync/auto-remove lifecycle or the external URL-fetching behavior implied by openWorldHint.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: url, auto_remove, and enable_auto_sync are documented in the schema, but name, parent_folder_id, and minimum_frequency_days are not. The description provides no compensating information about these undocumented parameters or their relationships (e.g., auto_remove only applies when auto-sync is enabled).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Create Url Document" is essentially a restatement of the tool name create_url_document_route. It names a verb and resource but offers no detail that would distinguish it from siblings such as create_text_document_route or create_file_document_route. This is a tautology rather than a real statement of purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the text/file document creation siblings, nor any mention of prerequisites such as auto-sync setup. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_video_generationD

Create Video Generation

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
audioNo
imageNo
audiosNoUp to 10 reference audios, e.g. for lipsync. Cannot be combined with `start_frame`/`end_frame`.
imagesNoUp to 30 reference images to draw subjects from. Cannot be combined with `start_frame`/`end_frame`.
promptNoA text description of the video to generate.
videosNoUp to 10 reference videos to draw subjects or motion from. Cannot be combined with `start_frame`/`end_frame`.
webhookNo
model_idNoThe model to use for the generation.
end_frameNo
resolutionNoThe resolution of the output video.
start_frameNo
aspect_ratioNoThe aspect ratio of the output video. With `auto`, the model picks an aspect ratio based on the inputs. First-frame / first-and-last-frame tasks always use `auto`.
duration_secsNoThe duration of the output video in seconds.
enhance_promptNoWhether the model may rewrite the prompt to improve results.
generate_audioNoWhether to generate audio with the video.
guidance_scaleNo
negative_promptNo
audio_guidance_scaleNo

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true. The description adds nothing about asynchronous generation behavior, asset prerequisites, webhook delivery, long-running job handling, or prompt/model requirements, so it provides no behavioral context beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a two-word title with no front-loaded information or structure. It is under-specified rather than concise, since it omits every useful detail an agent would need.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 19 parameters, nested media objects, no output schema, and no required parameters, the description provides no context about what generation entails or how to invoke it correctly. It is completely inadequate for the complexity of the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 19 parameters and only 53% schema description coverage, leaving fields such as seed, type discriminators, guidance_scale, negative_prompt, and audio_guidance_scale without documented meaning. The description mentions no parameters and adds no semantics beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description exactly restates the tool name and title, 'Create Video Generation,' giving no scope or detail beyond the name. It does not distinguish this tool from siblings such as create_image_generation, get_video_generation, or list_video_generations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no indication of when to choose this tool over alternatives like create_image_generation or list_video_generations. Usage is left entirely to inference from the name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_voiceB

Create A New Voice From Voice Preview Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
labelsNo
voice_nameYesName to use for the created voice.
voice_descriptionYesDescription to use for the created voice.
generated_voice_idYesThe generated_voice_id to create; obtain it from POST /v1/text-to-voice/design, POST /v1/text-to-voice/:voice_id/remix, or the response headers when generating previews.
played_not_selected_voice_idsNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-readonly, non-idempotent, non-destructive, open-world write. The description adds the important cost context that it spends ElevenLabs credits, which annotations do not cover. However, it says nothing about what happens to existing voices or the preview-to-voice relationship.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loads the verb, but the run-on phrasing 'Voice From Voice Preview Spends ElevenLabs credits' reads as two clipped clauses fused together, hurting clarity. No wasted sentences, but structure is clumsy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With five parameters, only 60% schema coverage, and no output schema, the description should carry more weight but does not. It omits the preview-to-creation workflow and any return/confirmation behavior, leaving gaps an agent must infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60% and the schema itself already explains generated_voice_id's provenance and the other required fields. The description adds no parameter-level meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Create A New Voice From Voice Preview.' This distinguishes it from sibling tools like create_pvc_voice and add_voice, which use different sources. No explicit sibling naming, but the 'from voice preview' scope is reasonably differentiating.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage context by saying the source is a voice preview and warns it spends credits, but never states when to use this vs. design/remix/text_to_voice siblings or what prerequisites exist. Usage is only weakly implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_workspace_webhook_routeC

Create Workspace Webhook

ParametersJSON Schema
NameRequiredDescriptionDefault
settingsYesSettings for creating an HMAC-authenticated webhook

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, covering the safety and side-effect profile. The description adds no behavioral context such as authentication requirements, response behavior, or what happens to existing routes, so it contributes nothing beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is maximally short but also under-specified: it merely repeats the tool title and earns no place as functional description. Conciseness here reflects a lack of content rather than efficient communication.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with a nested settings object and no output schema, the description is inadequate. While annotations cover safety and the schema covers parameters, the description provides no context about what a webhook route is, what it is for, or what the agent should expect after calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is high (100% per context signals), so the schema already documents all parameters including the nested settings object and its fields. The description adds no parameter meaning beyond what the schema supplies, which meets the baseline of 3 for high-coverage schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create Workspace Webhook' is a tautological restatement of the tool name and title, providing no additional distinguishing detail from siblings like edit_workspace_webhook_route or delete_workspace_webhook_route. It states a verb and resource but adds no scope or specificity beyond the name itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites, and no exclusions. An agent must infer usage entirely from the tool name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_agent_conversation_ticket_routeC
DestructiveIdempotent

Delete Agent Conversation Ticket

ParametersJSON Schema
NameRequiredDescriptionDefault
agentqa_ticket_idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered. The description adds no behavioral context beyond that—it doesn't mention that deletion is permanent, whether a ticket must exist, or any side effects. Without annotations this would be a 1, but annotations carry the burden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, short phrase with no waste and is front-loaded with the action. It is appropriately sized, though it borders on being too terse to be functional.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation tool with one undocumented parameter, no output schema, and no annotations beyond safety hints, the description omits prerequisites, side effects, and parameter meaning. It is not sufficient for an agent to call it confidently without opening the schema and inferring context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the sole required parameter (agentqa_ticket_id) is undocumented in both the schema and the description. The description does not say what a 'ticket id' is, its format, or how to obtain it, leaving a critical gap for the only parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Delete Agent Conversation Ticket' states a clear verb (Delete) and resource (Agent Conversation Ticket), so the purpose is identifiable. However, it largely restates the tool name and provides no differentiation from siblings like delete_conversation_route or delete_conversation_tag_route, leaving the agent to infer distinctions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when or when not to use this tool, nor are alternatives such as update_agent_conversation_ticket_route or list_agent_conversation_tickets_route mentioned. The agent must infer context from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_agent_draft_routeC
DestructiveIdempotent

Delete Agent Draft

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.
branch_idYesThe ID of the agent branch to use

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false, and openWorldHint=true, so the safety profile is fully covered by structured fields. The description adds no behavioral context beyond them – nothing about what a draft is, whether deletion is permanent, or any auth/permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The four-word description is efficiently front-loaded and wastes no space, but it is under-specified rather than genuinely concise – there is simply not enough content to be useful. Brevity here reflects omission, not discipline.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world delete operation with two required identifiers and no output schema, the description provides no consequential context – no explanation of scope, permanent effects, or relation to other agent-delete tools. Annotations cover safety but do not compensate for the missing task context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both required parameters (agent_id, branch_id) are fully documented in the schema at 100% coverage, so the schema carries the semantics. The description adds nothing about parameters, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource (delete + agent draft), but it is essentially the tool name restated and does not explain what an 'agent draft' is or how it differs from siblings like delete_agent_route, delete_agent_hold_audio_route, or delete_agent_test_folder_route. An agent can guess the action but cannot confidently distinguish it from the many other agent-delete variants.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. Given the dense cluster of delete_*agent* siblings, the absence of any routing guidance is a notable gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_agent_hold_audio_routeC
DestructiveIdempotent

Delete Agent Hold Audio

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that — no indication of whether the hold audio must exist, whether the deletion affects live calls, or what auth/permissions are required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short with no wasted words, but this is under-specification rather than conciseness — a four-word phrase that duplicates the title provides no front-loaded value for the agent to act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world mutation endpoint with no output schema, the description should at minimum clarify what resource is removed and any side effects. None of that is present, leaving key behavioral details to the annotations alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter with 100% schema description coverage ("The id of an agent. This is returned on agent creation."), so the schema fully documents it. The description contributes no additional parameter meaning, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Delete Agent Hold Audio" is essentially the tool name restated with underscores removed. It names a verb and resource, but adds no information an agent couldn't already derive from the identifier itself, and does not distinguish it from sibling post_agent_hold_audio_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, no prerequisites, and no reference to alternatives such as post_agent_hold_audio_route (setting the audio) or delete_agent_route. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_agent_routeC
DestructiveIdempotent

Delete Agent

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true. The description adds nothing beyond that, such as what exactly is destroyed, whether related data is removed, or whether the action is reversible. It does not contradict the annotations, but it fails to enrich behavioral understanding.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short and front-loaded, with no wasted words. However, for a destructive operation with many sibling tools, this terseness borders on under-specification rather than being appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the destructive nature, the many related delete siblings, and the absence of an output schema, the description is too minimal. The annotations cover safety hints and the schema covers the parameter, but the description does not clarify scope, irreversibility, or side effects, leaving the agent with insufficient context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single agent_id parameter is fully documented in the schema. The description adds no parameter meaning beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: delete an agent. It is specific enough to distinguish from read or create operations, but it does not differentiate from other delete siblings such as delete_agent_draft_route or delete_agent_hold_audio_route, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance, no prerequisites, and no alternatives. With many sibling delete tools in the same namespace, an agent receives no help in choosing this tool over others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_agent_test_folder_routeC
DestructiveIdempotent

Delete Agent Test Folder

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoForce delete. Required for deleting non-empty folders.
folder_idYesThe folder ID.

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered structurally. The description adds no behavioral context beyond the name, such as what gets deleted, whether deletion cascades to test cases, or whether permissions are required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short and front-loaded, but it is under-specified rather than appropriately concise. For a destructive operation, a single title-like phrase does not structure enough information to help an agent invoke it confidently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the destructive nature and the presence of annotations, the description should at least clarify the scope of deletion or any key constraints. It omits that force is required for non-empty folders and does not explain what an 'Agent Test Folder' is, leaving important context missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the two parameters (folder_id and force) are fully documented in the schema. The description adds no extra parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Delete') and resource ('Agent Test Folder'), so an agent can identify the operation. It does not explicitly contrast with sibling tools like create_agent_test_folder_route or update_agent_test_folder_route, but the name itself is descriptive enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, no prerequisites, and no warnings. It only restates the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_asset_endpointC
DestructiveIdempotent

Delete Asset

ParametersJSON Schema
NameRequiredDescriptionDefault
asset_idYesID of the asset.

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, openWorldHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds zero behavioral context beyond that — it does not explain permanence, side effects, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words with no waste, but this is under-specification rather than conciseness. There is no front-loaded explanation of scope or effect to anchor the agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive deletion tool, the description is far too thin: it omits what is removed, whether the action is recoverable, and what happens to dependent resources. The rich annotations partially compensate, but the description itself is inadequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter asset_id is fully documented in the schema (100% coverage). The description adds nothing beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Delete Asset" gives a specific verb and resource, which technically distinguishes it from siblings like get_asset, list_assets, and upload_asset. However, it says nothing about what an asset is or what the deletion scope is, so it is only minimally informative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, no prerequisites (e.g., permissions), and no reference to alternatives. The agent must infer everything from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_audio_isolation_history_itemC
DestructiveIdempotent

Delete Audio Isolation History Item

ParametersJSON Schema
NameRequiredDescriptionDefault
history_item_idYesIdentifier of the audio isolation history item.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false, and openWorldHint=true, so the safety and repeat-behavior profile is covered structurally. The description adds nothing on top of that — no mention of permanence, authorization requirements, or side effects — so it contributes no behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence with no redundancy or filler, so there is nothing wasteful to trim. However, minimal length here reflects under-specification rather than efficient front-loading, so it earns only an adequate score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world delete operation, the description offers no guidance on irreversibility, error conditions, or what a caller should have in hand before invoking it. With no output schema, the annotations carry most of the useful signal, leaving the description effectively empty of context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with a single required parameter (history_item_id) that is fully documented in the schema. Since the structured schema does the heavy lifting and the description adds no additional meaning, the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name and annotation title ('Delete Audio Isolation History Item'), adding no verb-resource scope beyond the identifier itself. It does not distinguish this delete from the many other delete_* siblings such as delete_speech_history_item or delete_dubbing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative routing guidance. An agent must infer from the name alone that this deletes a specific audio isolation history record, with no hint about prerequisites such as first listing history items via get_audio_isolation_history.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_auth_connectionC
DestructiveIdempotent

Delete Workspace Auth Connection

ParametersJSON Schema
NameRequiredDescriptionDefault
auth_connection_idYes

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, covering the key behavioral traits. The description adds no additional context about what gets destroyed, whether the action is reversible, or any rate limits or auth requirements, so it does not enhance the behavioral picture beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient phrase that is front-loaded and contains no unnecessary words. It is concise but under-specified for a deletion operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive operation with no output schema and an undocumented parameter, the description is insufficient. It does not explain the consequences of deletion, required permissions, or how to obtain the auth_connection_id, leaving significant gaps for safe tool invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'auth_connection_id' has 0% schema description coverage, meaning neither the schema nor the description explains what this ID represents or its format. The description provides no information about the parameter, leaving it completely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb (Delete) and resource (Workspace Auth Connection), which is specific enough. However, it provides no differentiation from sibling tools like delete_secret_route or delete_workspace_webhook_route, which are also workspace-level deletion operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidance is provided. The description does not indicate when to use this tool versus alternatives like deleting a secret or webhook, nor does it mention any prerequisites such as required permissions or the impact of deletion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_batch_callC
DestructiveIdempotent

Delete A Batch Call.

ParametersJSON Schema
NameRequiredDescriptionDefault
batch_idYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered. The description adds no context about what deletion removes, whether it is reversible, or what permissions are required; it merely restates the destructive action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. However, its extreme brevity leaves little useful structure beyond the core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with one undocumented parameter and many related batch-call siblings, the description is too thin. It does not explain the parameter, consequences, or when this deletion is appropriate instead of cancellation or retry.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the sole parameter batch_id has no documentation. The description does not identify the required batch identifier or its expected format, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: delete a batch call. It is clear but does not differentiate from sibling batch-call tools such as cancel_batch_call or get_batch_call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like cancel_batch_call, retry_batch_call, or get_batch_call. Usage is only implied by the word 'Delete'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_chapter_endpointC
DestructiveIdempotent

Delete Chapter

ParametersJSON Schema
NameRequiredDescriptionDefault
chapter_idYesThe ID of the chapter.
project_idYesThe ID of the Studio project.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose destructiveHint=true, idempotentHint=true, and openWorldHint=true. The description adds no further behavioral context, such as what gets deleted, whether associated data is affected, or any irreversibility warning.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

'Delete Chapter' is extremely short and front-loads the operation, but it is under-specified for a destructive tool. The terseness avoids waste, yet it lacks the structure and safety context that would make it appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with only two required parameters, and the schema plus annotations cover the structured details and safety profile. However, the description provides no usage context or consequences of deletion, leaving an agent with only the bare operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both project_id and chapter_id are fully documented in the schema. The description does not add any parameter meaning beyond what the schema already provides, which matches the baseline for high-coverage schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource combination: 'Delete Chapter.' It distinguishes the operation from siblings like add_chapter and edit_chapter, but it does not clarify scope or differentiate from other delete endpoints beyond the chapter resource.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. It does not mention prerequisites, required permissions, or related operations such as edit_chapter or get_chapter_by_id_endpoint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_chat_response_test_routeC
DestructiveIdempotent

Delete Agent Response Test

ParametersJSON Schema
NameRequiredDescriptionDefault
test_idYesThe id of a chat response test. This is returned on test creation.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered structurally. The description adds nothing beyond that — no statement that the deletion is permanent, whether the test_id must exist, or what an error looks like. It is consistent with the annotations, not contradictory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A three-word phrase is maximally concise and front-loaded, but it is under-specified rather than efficiently concise. There is no wasted sentence, yet it also carries no useful information beyond the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, single-parameter delete whose safety semantics are fully carried by annotations, the description is just barely adequate. It omits irreversibility, not-found behavior, and any mention of the paired create/get/update tools an agent might need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single test_id parameter is documented as 'returned on test creation'. The description adds no additional meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Delete) and resource (Agent Response Test), which is enough to place it among the delete_* siblings. However, it uses 'Agent Response Test' while the tool name and schema say 'chat response test', creating mild terminology ambiguity, and it does not differentiate itself from update_agent_response_test_route or get_agent_response_test_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus the sibling test-management tools (create/get/update/list response tests), nor any prerequisite or caution. The agent must infer usage entirely from the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_conversation_routeC
DestructiveIdempotent

Delete Conversation

ParametersJSON Schema
NameRequiredDescriptionDefault
conversation_idYesThe id of the conversation you're taking the action on.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false, and openWorldHint=true, so the safety profile is partly covered. The description adds nothing beyond that: no mention of irreversibility, cascade effects on messages/transcripts, required permissions, or confirmation requirements for a permanently destructive action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded phrase with no wasted words, but its brevity stems from under-specification rather than disciplined editing. It is concise but earns little.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, irreversible-seeming delete with no output schema, the description should clarify what exactly is removed, whether deletion is permanent or soft, and any side effects. None of that is present, leaving a significant gap despite annotation coverage of the safety hints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter and schema description coverage is 100%, with the schema itself documenting conversation_id. The description contributes no additional meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Delete Conversation' names a verb and a resource, so the basic operation is intelligible. However, it is effectively a human-readable restatement of the tool name with no scope or differentiator, leaving ambiguous boundaries against siblings such as delete_conversation_tag_route or delete_chat_response_test_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when this operation should be used, what preconditions exist (e.g. conversation must be resolved/closed), or which sibling to choose for related deletions. The description offers zero routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_conversation_tag_routeC
DestructiveIdempotent

Delete Conversation Tag

ParametersJSON Schema
NameRequiredDescriptionDefault
tag_idYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, openWorldHint=true, and readOnlyHint=false, so the safety profile is covered structurally. The description adds nothing beyond that—no mention of permanence, cascade effects, or auth requirements—so it contributes no behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase, but it is under-specified rather than concise—it omits essential context and does not front-load any useful information. Brevity here reflects emptiness, not efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive operation with only one required parameter and no output schema, an agent still needs to know what is deleted, whether it is reversible, and how it differs from unassigning. The description supplies none of this, relying entirely on annotations for the bare minimum safety signal.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the schema does not explain the single required parameter 'tag_id'. The description also says nothing about it, leaving the agent to infer that it is the ID of the tag to delete, with no format or source guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Delete Conversation Tag' is a verbatim restatement of the tool name with the '_route' suffix removed, making it a tautology rather than an independent statement of purpose. It does not distinguish this tool from siblings like unassign_conversation_tag_route or delete_conversation_route, and adds no scope or constraint information.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as unassign_conversation_tag_route (which removes a tag from a conversation) or update_conversation_tag_route. The description implies deletion but provides no context, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_dubbingC
DestructiveIdempotent

Delete Dubbing

ParametersJSON Schema
NameRequiredDescriptionDefault
dubbing_idYesID of the dubbing project.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive=true, idempotent=true, and readOnly=false, providing the safety profile structurally. The description adds no context about what is deleted, whether associated data is cascaded, or any permission requirements, so it offers nothing beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and front-loaded, with no wasted words. However, it is under-specified for a destructive tool, making its brevity insufficient rather than structurally strong.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one parameter, rich annotations, and no output schema, the structured data covers much of the invocation surface. Still, the description fails to clarify which dubbing resource is deleted or how it differs from similar delete siblings, which is context an agent needs for safe selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single required parameter dubbing_id has 100% schema description coverage ('ID of the dubbing project.'). The description adds no syntax, format, or lookup guidance, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the name and title ('Delete Dubbing'), so it provides no information beyond what the tool name already conveys. It does not distinguish this from siblings such as dubbing_project_delete, dubbing_language_delete, or dubbing_transcript_segment_delete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool or when to prefer alternatives. For a destructive delete operation, routing guidance is especially important, yet the description says nothing beyond the operation name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_finetuneC
DestructiveIdempotent

Delete Music Finetune

ParametersJSON Schema
NameRequiredDescriptionDefault
finetune_idYes

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is fully covered by structured data. The description adds nothing beyond that - no note on irreversibility, whether the underlying model/data is reclaimed, or permission requirements. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with zero filler, so nothing is wasted, but at this length the brevity comes at the cost of specification rather than being genuinely tight writing. Structure is acceptable; substance is thin.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no output schema, no annotation-independent detail, and an undocumented required parameter, the description leaves too much unspecified. An agent cannot tell what is destroyed, whether the action is recoverable, or what identifier format to supply.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter finetune_id is an undocumented string. The description provides no clarification of what identifier is expected or where it comes from, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Delete Music Finetune' essentially restates the tool name delete_finetune, adding only the qualifier 'Music'. It names a verb and a resource, but there is no differentiation from siblings such as get_finetune, get_finetunes, or update_finetune, and it is vague about whether a finetune model, job, or dataset is being removed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. With four finetune-related siblings (create_finetune, update_finetune, get_finetune, get_finetunes) in the toolset, an agent gets no help on routing; it can only infer usage from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_inviteC
DestructiveIdempotent

Delete Existing Invitation

ParametersJSON Schema
NameRequiredDescriptionDefault
emailYesThe email of the customer

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds nothing on top of that — it doesn't say whether deletion revokes a pending invitation, whether it fails for unknown emails, or whether it is scoped to a workspace, so there is no value beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words, front-loaded and waste-free, with the verb leading. It is efficient, though the terseness edges toward under-specification rather than genuine concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter delete whose annotations already carry the destructive/idempotent profile and with no output schema required, a minimal description could suffice. However, it omits invocation context and the effect of deletion on the invitation lifecycle, leaving gaps an agent would need filled before calling it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (the single 'email' parameter is documented in the schema as 'The email of the customer'). The description adds no syntax, format, or scoping meaning beyond that, so the baseline 3 for full schema coverage applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource ('Delete ... Invitation'), which sets it apart from reads, but it is essentially an expansion of the tool name/title 'delete_invite'/'Delete Invite'. It gives no scope detail (whose invitation, pending vs. accepted) and does not distinguish itself from nearby tools like remove_member or invite_user, so it is only minimally clearer than the identifier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance, no prerequisites (e.g. ownership or admin requirement), and no mention of alternatives such as remove_member or invite_users_bulk. The agent must infer everything about invocation context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_knowledge_base_documentC
DestructiveIdempotent

Delete Knowledge Base Document Or Folder

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoIf set to true, the document or folder will be deleted regardless of whether it is used by any agents and it will be removed from the dependent agents. For non-empty folders, this will also delete all child documents and folders.
documentation_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, openWorldHint=true, and readOnlyHint=false, so the safety profile is covered structurally. The description adds nothing beyond that — it does not mention the cascade behavior or the force override that determines whether dependent agents are affected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded fragment with no wasted words, which is structurally efficient. But it is under-specified rather than truly concise — the brevity comes at the cost of any actionable context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a destructive, cascading delete with no output schema and no usage context in the description. Annotations and the schema cover safety and the force flag, but an agent gets no guidance on verifying the target or on when deletion is the right operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both documentation_id and force (including the cascade and dependent-agent semantics) are fully documented in the schema. The description only hints at folder scope via "Or Folder" and adds no syntax or format detail, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase names the verb (Delete) and the resource (Knowledge Base Document), and the addition of "Or Folder" signals a scope broader than the tool name implies. However, it is essentially a restatement of the tool name with no differentiation from siblings such as post_knowledge_base_bulk_delete_route or delete_rag_index.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites, and no routing to alternatives. The folder-deletion capability appears only as an unexplained suffix, with no indication of when a folder vs. a document is the intended target.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_mcp_server_routeC
DestructiveIdempotent

Delete Mcp Server

ParametersJSON Schema
NameRequiredDescriptionDefault
mcp_server_idYesID of the MCP Server.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the essential safety profile is covered by structured data. The description adds nothing beyond the name—no statement about permanence, cascading effects on related routes/tools, or auth requirements. With annotations doing the heavy lifting, the low score reflects the absence of any extra context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short and front-loaded, but it is under-specified rather than truly concise. It earns no penalty for length but also conveys minimal structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, idempotent tool in a large toolset with no output schema, the description is too thin. It omits any behavioral or operational context beyond what annotations and the schema already state.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and includes a description for mcp_server_id, so the schema fully documents the single parameter. The description adds no meaning beyond the schema, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb+resource ('Delete Mcp Server'), which is clear enough about what it does, but it uses the entity name rather than the tool's route concept and adds no scope or sibling differentiation. Given the huge sibling list containing many 'delete_*' tools, it does not help an agent distinguish this from the others.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, no mention of alternatives. The agent gets no help deciding between this and sibling delete tools or list_mcp_servers_route/get_mcp_route.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_phone_number_routeC
DestructiveIdempotent

Delete Phone Number

ParametersJSON Schema
NameRequiredDescriptionDefault
phone_number_idYesThe phone number ID. This is returned when a phone number is imported.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered. The description adds no further behavioral context such as permanence, side effects on associated routes, or required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two words. While brief, it is under-specified rather than optimally concise, omitting necessary context for a destructive operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter destructive delete tool, the annotation and schema coverage make the definition minimally viable. Still, the description itself leaves consequences, prerequisites, and usage scope unstated, which are clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents phone_number_id, including how the ID is obtained. The description adds no parameter meaning beyond what the structured field provides, so the baseline for high schema coverage applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Delete Phone Number' states a specific verb and resource, so an agent can tell what the tool does. However, it provides no differentiation from the many other delete_* siblings or scope details beyond the resource name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus siblings like create_phone_number_route, update_phone_number_route, or get_phone_number_route. It also omits any prerequisites or confirmation context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_procedure_draft_routeC
DestructiveIdempotent

Delete Procedure Draft

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesAgent ID to get the procedure draft from
branch_idYesBranch ID to get the procedure draft from
procedure_idYesThe procedure ID

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that—no note on irreversibility, permissions required, or what happens to the underlying procedure/branch state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with zero waste, but the brevity here reflects under-specification rather than efficient conciseness—there is no front-loaded scope or constraint for an agent to act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, three-parameter deletion route with no output schema, the description omits identifying which draft is deleted, what the branch/agent scoping implies, and any confirmation or irreversibility context. Annotations cover safety hints but the description leaves the call semantics thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three required parameters (agent_id, branch_id, procedure_id) are documented in the schema. The description contributes no additional parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Delete Procedure Draft' simply restates the tool name (delete_procedure_draft_route) without adding scope, target, or differentiating detail. It conveys verb+resource only through the name itself, which fits the tautology criterion rather than a clear, differentiated purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many sibling delete tools (delete_agent_draft_route, remove_procedure_route, delete_agent_route, etc.) or versus update_procedure_draft_route. No prerequisites or context are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_projectC
DestructiveIdempotent

Delete Studio Project

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesThe ID of the Studio project.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that: it does not say the deletion is irreversible, nor whether dependent assets (chapters, snapshots, tracks) are cascaded or orphaned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short fragment with no wasted words, but the brevity here reflects under-specification rather than disciplined concision. There is no structure or front-loaded constraint because there is no content to structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive operation on a container entity that has many dependent resources, the description omits the consequences an agent should know before calling it. The annotations carry the safety signal and the schema is complete, but the risk context of an irreversible project deletion is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single parameter at 100% schema description coverage, the schema fully documents project_id. The description contributes no additional meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Delete Studio Project' states a verb and a resource, but it is essentially the tool name and annotation title restated with the word 'Studio' added. That qualifier is the only real signal, and it is never explained as distinguishing this from the many dubbing_project_* and audio-native project tools in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites (ownership/permissions), and no routing to alternatives such as deleting a dubbing project or archiving. An agent must infer everything from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_pvc_voice_sampleC
DestructiveIdempotent

Delete Pvc Voice Sample

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
sample_idYesSample ID to be used

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false, and openWorldHint=true, so the agent knows this is a destructive, idempotent write. The description adds nothing beyond the name: it doesn't state what gets destroyed, whether it's recoverable, or any rate/permission constraints. With annotations carrying the safety profile, the description still fails to add any behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short and front-loaded, which is efficient, but it is under-specified rather than truly concise. The single phrase carries no useful information beyond the name, so brevity here reflects a lack of content rather than tight writing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive deletion tool with no output schema, the description should clarify what a PVC voice sample is, the relationship to voice_id, and any irreversibility. It omits all of this, leaving annotations and schema to carry the full burden. Inadequate for a destructive operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both voice_id and sample_id are documented in the schema (voice_id even includes the endpoint to list voices). The description repeats the term 'sample' but adds no syntax or meaning beyond the schema. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name in title case, giving the verb+resource ('delete' + 'Pvc Voice Sample'). It conveys the purpose but is essentially a tautology of the name with no additional differentiation from siblings like delete_sample, edit_pvc_voice_sample, or delete_voice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no conditions, no alternative tools named. An agent must infer from the name alone that this deletes a sample belonging to a PVC voice, with no exclusions or prerequisites stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_rag_indexC
DestructiveIdempotent

Delete Rag Index.

ParametersJSON Schema
NameRequiredDescriptionDefault
rag_index_idYesThe id of RAG index of document from the knowledge base.
documentation_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true, idempotentHint=true, openWorldHint=true, and readOnlyHint=false, so safety behavior is covered. The description adds no extra context about what is permanently removed, authorization needs, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The one-phrase description is concise but under-specified for a destructive operation with two required identifiers. Brevity here is not effective conciseness; it fails to front-load useful selection or caution information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive deletion tool in a large sibling set, the description is incomplete. Annotations and schema cover safety and parameters, but the description does not help an agent choose this tool over other RAG-index operations or understand the deletion target.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both required parameters are fully documented in the schema. The description adds no parameter meaning, so the baseline of 3 applies for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb and resource, but it is an exact restatement of the tool name and title. It adds no scope, target details, or differentiation from the many sibling delete and RAG-index tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as get_rag_indexes, get_or_create_rag_indexes, rag_index_status, or query_agent_knowledge_base_rag_route. Prerequisites and consequences are absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_sampleC
DestructiveIdempotent

Delete Sample

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
sample_idYesSample ID to be used, you can use GET https://api.elevenlabs.io/v1/voices/{voice_id} to list all the available samples for a voice.

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false, and openWorldHint=true, so the safety profile is covered elsewhere. The description adds nothing beyond that: no statement about irreversibility, required voice ownership, or whether the sample's audio is purged. No contradiction, but no added value either.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words with zero waste, but that brevity comes from under-specification rather than tight writing. There is no front-loaded useful information to preserve.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a destructive, non-read-only mutation with no output schema and no description of consequences, confirmation requirements, or failure modes. For an irreversible delete, the definition is far too thin for an agent to invoke it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both voice_id and sample_id are documented in the schema (including how to list valid values), making 3 the baseline. The description contributes nothing about the required pairing of the two IDs or how sample_id is scoped to voice_id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Delete Sample" restates the tool name and title without elaboration. It gives no indication of what a sample belongs to (a voice), which matters here because siblings such as delete_pvc_voice_sample and delete_audio_isolation_history_item are also 'delete a sample-like thing' tools. An agent gets verb+resource but nothing to distinguish this from them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, prerequisites, or alternative routing is provided. The description never explains how this differs from delete_pvc_voice_sample or delete_voice, so the agent must guess based on the name and required parameters alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_secret_routeC
DestructiveIdempotent

Delete Convai Workspace Secret

ParametersJSON Schema
NameRequiredDescriptionDefault
secret_idYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so safety aspects are covered. The description adds nothing beyond the operation name - it does not disclose what happens to dependent agents or resources (notably a get_secret_dependencies_route sibling exists), whether deletion is reversible, or any auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded phrase with zero wasted words. It is efficient, though the brevity borders on under-specification rather than concise coverage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no output schema and an undocumented identifier, the description is too thin. It omits dependency/cascade behavior (relevant given get_secret_dependencies_route), confirmation semantics, and any parameter guidance, leaving the agent to infer too much.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there is one required parameter (secret_id). The description never references the parameter, its format, or where the ID comes from, so it fails to compensate for the schema's lack of documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Delete) and resource (Convai Workspace Secret), scoping it to the correct workspace. It does not explicitly distinguish itself from siblings like update_secret_route or get_secret_route, but the delete action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives (update_secret_route, create_secret_route, get_secret_dependencies_route), no prerequisites, and no warning about what a delete implies. The agent gets a purpose but no situational direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_segmentC
DestructiveIdempotent

Deletes A Single Segment Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
dubbing_idYesID of the dubbing project.
segment_idYesID of the segment

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true and openWorldHint=true, so the safety profile is covered. The one addition beyond structured data is "Deprecated upstream," which is genuinely useful but unexplained (no replacement, no timeline). It also fails to state what exactly gets destroyed or any permission requirements, so it only modestly exceeds the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, which is good, but the two clauses are jammed together ("Deletes A Single Segment Deprecated upstream") with odd title-casing and no punctuation or clear separation. The deprecation note is not front-loaded or clearly flagged, reducing scannability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter destructive mutation with no output schema, the annotations cover the safety profile and the schema covers the parameters, so the essentials are present. What is missing is the deprecation story (alternative tool, migration path) and any indication of blast radius, so the description is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both required parameters (dubbing_id, segment_id) documented in the schema itself, so the description carries no additional burden. The description adds no parameter-level meaning, which matches the baseline 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and resource ("Deletes A Single Segment"), so the basic action is clear. However, it never says which kind of segment (dubbing transcript segment, chapter segment, etc.) despite the sibling list containing dubbing_transcript_segment_delete, dubbing_target_transcript_segment_update and migrate_segments. Without that scoping it is hard to separate from its siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no named alternative or exclusion are given. The fragment "Deprecated upstream" hints the tool may be unsafe to adopt, but it is cryptic and doesn't point the agent at a replacement or state conditions under which this tool should still be called.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_service_account_api_keyC
DestructiveIdempotent

Delete Service Account Api Key

ParametersJSON Schema
NameRequiredDescriptionDefault
api_key_idYes
service_account_user_idYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false and openWorldHint=true, so the safety profile is covered by structured data. The description adds nothing on top of that — no statement about irreversibility, permission requirements, or what happens to the service account after key deletion.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but this is under-specification rather than conciseness — the single phrase is simply the tool name repeated. There is no front-loaded purpose statement, precondition, or consequence to justify the brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, idempotent delete operation with no output schema and two undocumented identifiers, the description supplies none of the missing context (auth scope, irreversibility, effect on dependent keys). Only the annotations carry any behavioral signal, leaving the description substantially incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both required parameters (service_account_user_id, api_key_id). The description does not compensate at all, offering no format, source, or relationship guidance for either identifier, which is exactly the case where free-text parameter help is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name, capitalizing each word but adding no information. While 'delete service account api key' does identify a verb and resource, it is a tautology against the name/title rather than a description, and it does nothing to distinguish this tool from the many other delete_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as delete_secret_route, delete_auth_connection, or delete_service_account. No prerequisites, no mention of the create/edit counterparts (create_service_account_api_key, edit_service_account_api_key) that an agent might confuse it with.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_speech_engineC
DestructiveIdempotent

Delete Speech Engine

ParametersJSON Schema
NameRequiredDescriptionDefault
speech_engine_idYesThe speech engine ID (accepts seng_ or agent_ prefix)

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is known without the description. The description contributes nothing extra - no note on permanence, what happens to agents still referencing the engine, or required permissions - so it fails to add value over structured data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short but for the wrong reason: the two-word string is a restatement of the title rather than a load-bearing sentence. Brevity here reflects under-specification, not efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world, irreversible delete operation with no output schema, the description should at minimum warn about permanence and side effects. It supplies none of that, leaving the agent with only the annotated hints and the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter is fully documented in the schema, including the 'seng_ or agent_' prefix acceptance. With no param info in the description, the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is the tool name restated verbatim: 'Delete Speech Engine'. While that is technically a verb+resource, it adds no information beyond the title and does nothing to distinguish this tool from the many other delete_* siblings in the list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool rather than alternatives (e.g., update_speech_engine or get_speech_engine), no prerequisites, and no conditions. The agent is left to infer usage entirely from the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_speech_history_itemC
DestructiveIdempotent

Delete History Item

ParametersJSON Schema
NameRequiredDescriptionDefault
history_item_idYesHistory item ID to be used, you can use GET https://api.elevenlabs.io/v1/history to receive a list of history items and their IDs.

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that: no note on permanence, whether the underlying audio is removed, or any confirmation requirement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The three-word phrase is front-loaded and wastes no words, but it is terse to the point of under-specification rather than efficient communication; nothing is said that the title does not already say.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter destructive tool, the schema documents the ID and the annotations cover safety and idempotency, so the essentials are present. However, given that this permanently deletes user content, the description should at minimum reinforce irreversibility or point at the retrieval step.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema itself explains history_item_id and even points to the GET /v1/history endpoint for retrieving IDs. The description adds no parameter meaning of its own, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Delete History Item" states a verb and a resource, so the basic action is clear, but it is essentially a restatement of the tool name and does not distinguish it from the many other delete_* siblings (e.g. delete_audio_isolation_history_item) or clarify what a "history item" is in this API.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get_speech_history_item_by_id, download_speech_history_items, or delete_audio_isolation_history_item. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_tool_routeC
DestructiveIdempotent

Delete Tool

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoIf set to true, the tool will be deleted regardless of whether it is used by any agents and it will be removed from the dependent agents and branches.
tool_idYesID of the requested tool.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false and openWorldHint=true, so the safety profile is covered by structured data. The description adds nothing on top — no mention of permanence, permissions, or what happens to dependent agents (that detail lives only in the schema's 'force' parameter).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At two words the text is technically short but is under-specified rather than concise — it omits information an agent needs rather than trimming redundancy. There is no structure or front-loading to evaluate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive delete operation, the description supplies no context about consequences, dependency removal, or irreversibility. Annotations and the schema's 'force' description carry most of the load, but the description itself leaves an agent with nothing beyond the tool name.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the 'force' parameter is well documented in the schema itself, including its cascade behavior. The description contributes no additional parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Delete Tool' is a bare restatement of the tool name and title 'Delete Tool Route'. It conveys a verb and resource but adds no scope, target, or differentiation from the many other delete_* siblings such as delete_agent_route or delete_project.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no mention of prerequisites, and no warning about the cascading effects of deletion. The agent must infer everything from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_transcript_by_idC
DestructiveIdempotent

Delete Transcript By Id

ParametersJSON Schema
NameRequiredDescriptionDefault
transcription_idYesThe unique ID of the transcript to delete

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=true, and openWorldHint=true, so the safety profile is covered. However, the description adds zero context beyond that — it does not state whether deletion is permanent, whether it cascades, or what happens on a non-existent ID, despite the destructive nature of the operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four words with no padding, but this is under-specification rather than conciseness — there is no front-loaded scope, no constraint, and nothing that earns the space a longer description would occupy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, irreversible-style operation, the agent gets no information about permanence, side effects, or confirmation requirements from the description. With annotations covering the safety profile and no output schema, the definition is technically callable but leaves meaningful behavioral gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is only one parameter and schema description coverage is 100%, so the schema fully documents 'transcription_id'. The phrase 'By Id' in the description implies an ID parameter but adds no syntax, format, or sourcing detail. Baseline 3 is appropriate when the schema carries the burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name/title, adding no information beyond the identifier itself. While the name does convey verb+resource, the description text itself is a tautology and offers no differentiation from the sibling get_transcript_by_id.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance whatsoever on when to use this tool, when not to, or what alternatives exist (e.g., get_transcript_by_id for reading). Prerequisites, permissions, and the distinction from other delete_* siblings are entirely absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_voiceD
DestructiveIdempotent

Delete Voice

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations carry the safety profile (destructiveHint=true, idempotentHint=true, openWorldHint=true), so the agent knows this is a destructive, idempotent, open-world write. The description adds nothing beyond that — it does not mention irreversibility, cascading effects on agents/workspaces using the voice, or required permissions. Repeating the annotation-covered operation is the only signal present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words that merely echo the name and annotation title. Brevity here is under-specification, not efficient communication — there is no content to front-load.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world deletion tool with no output schema, the description must explain scope and consequences. It provides none, leaving the agent unable to determine what is destroyed or what state changes result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema itself documents voice_id including a pointer to the list endpoint. Baseline 3 applies; the description adds no extra meaning about the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Delete Voice' restates the tool name and title with zero added specificity. It does not say what a 'voice' is in this system, what deletion entails, or distinguish this from sibling delete tools like delete_pvc_voice_sample, delete_sample, or delete_audio_isolation_history_item.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no alternatives named, no prerequisites or warnings about irreversibility. The description gives the agent nothing to route on.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_whatsapp_accountC
DestructiveIdempotent

Delete Whatsapp Account

ParametersJSON Schema
NameRequiredDescriptionDefault
phone_number_idYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, openWorldHint=true, and readOnlyHint=false, so the safety profile is largely covered structurally. The description adds nothing on top: it does not say what gets destroyed (account record vs. associated phone numbers/conversations), whether the action is reversible, or what permissions are required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words contain no waste but the brevity stems from under-specification rather than tight editing. There is nothing front-loaded to sequence because there is essentially no content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, irreversible-ish delete operation with no output schema and a completely undocumented parameter, the description is far too thin. Even with annotations covering the safety hints, an agent needs to know what the deletion removes and what confirmation guarantees apply.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions phone_number_id. The single parameter is self-descriptively named and the tool name implies which account is targeted, so an agent can partially infer intent, but no semantics (format, whether it can be a WABA id, whether it must exist) are supplied by either source.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name ('delete_whatsapp_account' → 'Delete Whatsapp Account') with capitalization only. It conveys a verb and resource, but adds no scope, target, or differentiation from siblings like get_whatsapp_account, list_whatsapp_accounts, or update_whatsapp_account. This is the tautology case the rubric scores 2.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this rather than the sibling get_whatsapp_account, list_whatsapp_accounts, update_whatsapp_account, or the generic delete_phone_number_route. No prerequisites, no confirmation/warning conditions for an irreversible operation are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_workspace_webhook_routeC
DestructiveIdempotent

Delete Workspace Webhook

ParametersJSON Schema
NameRequiredDescriptionDefault
webhook_idYesThe unique ID for the webhook

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond that — no disclosure of irreversibility, cascade effects, auth requirements, or what happens to associated data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words is concise but this is under-specification rather than effective brevity; nothing is front-loaded because there is nothing substantive to front-load.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive single-param tool with no output schema, the annotations carry safety but the description leaves the agent without any operational context — no confirmation of irreversibility, no note about dependent resources.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single required parameter (webhook_id) with 100% schema description coverage, so the schema fully documents it. Per the baseline rule for high coverage with no additional description-level param info, a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Delete Workspace Webhook' merely restates the tool name and adds no distinguishing information. It does convey verb+resource, so an agent knows it deletes a workspace webhook, but nothing separates it from siblings like create_workspace_webhook_route or edit_workspace_webhook_route beyond the obvious verb.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, no mention of alternatives or the permanent nature of deletion. The agent gets zero routing help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

disableC

Disable Api Key

ParametersJSON Schema
NameRequiredDescriptionDefault
api_key_nameYesMust be set to `self` to disable the API key used to authenticate this request. Required as an explicit confirmation to avoid accidentally disabling the wrong key.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is partly covered externally. The description adds nothing beyond that: it does not say whether disabling is reversible, whether it requires elevated permissions, or what effect it has on existing requests. For a mutation tool, this is a meaningful gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a three-word fragment with no waste, but it is under-specified rather than genuinely concise. It is front-loaded by default but carries too little information to be considered well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one required parameter, no nested objects, and no output schema, the tool is simple, but the description still omits the behavioral context an agent needs: reversibility of disabling, and how it relates to the delete/edit API-key siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema itself is unusually rich, explaining that api_key_name must be `self` as an explicit confirmation safeguard. The description contributes no additional parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Disable Api Key' names a verb and a resource, so the basic operation is inferable. However it is essentially a restatement of the tool name/title with no differentiation from the many sibling auth/API-key tools such as delete_service_account_api_key or edit_service_account_api_key. An agent cannot tell from this text how 'disable' differs from deleting or editing a key.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, no prerequisites, and no mention of the delete/edit siblings that operate on the same resource. Usage is left entirely to inference from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_speech_history_itemsB

Download History Items Returns application/zip bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
output_formatNo
history_item_idsYesA list of history items to download, you can get IDs of history items and other metadata using the GET https://api.elevenlabs.io/v1/history endpoint.

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false and openWorldHint=true, so the safety profile is covered. The description usefully adds that the payload is application/zip bytes (a multi-item archive), which the annotations do not convey, but it omits auth requirements, size/batch limits, and whether re-running is safe.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence, front-loaded with the action and return type. The phrasing is slightly run-on ("Download History Items Returns application/zip bytes"), but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a filesystem-writing operation, the description should at minimum clarify what the zip contains and what output_format accepts. It does disclose the zip return type, but leaves output_format and multi-item/batch behavior undocumented, so the definition is only partially complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 67%, and output_format carries no description anywhere. The description merely restates the already-documented output_path behavior ("pass output_path to save them") and adds nothing about output_format or valid ID sources, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Download History Items") and adds the return format (application/zip bytes). However, it does not distinguish itself from near-siblings such as get_audio_full_from_speech_history_item, get_speech_history_item_by_id, or get_speech_history, so the agent cannot tell which history-fetching tool applies without reading schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance is "pass output_path to save them," which is a parameter note rather than a when-to-use rule. There is no statement of when this bulk-download tool should be chosen over the single-item audio siblings, nor any prerequisite or exclusion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubB

Dubs All Or Some Segments And Languages Spends ElevenLabs credits. Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
segmentsYesDub only this list of segments.
languagesYes
dubbing_idYesID of the dubbing project.

TDQS

B3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds real behavioral context beyond the annotations: it consumes ElevenLabs credits (a cost side effect that idempotentHint=false and destructiveHint=false do not convey) and it is deprecated upstream, which is decision-relevant. It stops short of describing auth requirements or partial-failure behavior, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loads the action, which is good, but the sentence is a garbled run-on with stray title casing that forces the reader to re-parse it. Being terse is not the same as being economical when the single sentence is ambiguous.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-required-parameter mutating tool with no output schema, the description covers the two highest-value facts (credit spend, deprecation) but omits what the call actually returns and how segments/languages should be populated. Adequate, but several gaps remain for a paid mutation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%: segments and dubbing_id are described in the schema, while languages has no schema description. The phrase 'All Or Some Segments And Languages' vaguely implies optional filtering, but the description adds essentially no format or semantics beyond what the schema already provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

It names a verb (dubs) and the resource being modified (segments and languages of a dubbing project), so the core action is inferable. However, the run-on title-case phrasing is garbled, and it gives no differentiation from the many sibling dubbing tools (create_dubbing, migrate_segments, dubbing_project_*), so an agent cannot easily place it in the family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is a weak implicit signal — 'Deprecated upstream' hints the tool may be avoided — but it never says when to use this versus create_dubbing or the dubbing_language_* tools, nor what 'All Or Some' means as a selection condition. No prerequisites or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_language_createC

Create Dubbing Language Target Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesIdentifier of the parent dubbing project.
translationsNo
voice_settingsNo
target_languageYesBCP-47 language tag to dub the project into (for example, `fr` or `es-MX`). Must be one of the [languages the project's dubbing model supports](https://elevenlabs.io/docs/help-center/product/dubbing/which-languages-are-supported-in-dubbing), and a region-qualified tag must be one of the supported di

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true. The description adds a useful behavioral detail: it spends ElevenLabs credits. It does not explain what a Dubbing Language Target is, whether repeated calls duplicate resources, or what permissions are required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loads the action, but it is ungrammatical and fragmented: 'Create Dubbing Language Target Spends ElevenLabs credits.' It is concise but under-specified for a tool that creates a resource and incurs cost.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool that requires project_id and target_language and spends credits, the description is far too sparse. It does not explain the effect on the dubbing project, return value, or any constraints beyond the implicit cost warning.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, with 'translations' having no description at all. The description provides no parameter meaning, so it does not compensate for the schema gap. The agent must rely entirely on the schema for project_id, target_language, translations, and voice_settings.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Create Dubbing Language Target.' This distinguishes it from sibling tools like dubbing_language_get, dubbing_language_list, and dubbing_language_delete. It does not mention the parent project or scope, but the core purpose is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. The only hint is that it 'Spends ElevenLabs credits,' which is a cost warning rather than usage direction. No prerequisites or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_language_deleteC
DestructiveIdempotent

Delete Dubbing Language Target

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesIdentifier of the parent dubbing project.
language_idYesIdentifier of the language target to delete.

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that: it does not say what is destroyed alongside the language target (transcripts, dubbed audio), whether deletion is reversible, or what authorization is required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four words, front-loaded with the verb and resource, with zero padding. It is concise, though the brevity reflects under-specification rather than disciplined editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no output schema, the description should at minimum state the blast radius of the deletion and any precondition. As written, an agent cannot tell whether deleting a language target also removes its transcript or dubbed output, which is material before invoking a destructive tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both project_id and language_id fully documented in the schema, so the baseline is 3. The description contributes no additional meaning about the identifiers or their expected format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (Delete) and resource (Dubbing Language Target), so the operation is identifiable. However, it is essentially a restatement of the tool name and title, and offers no scope or differentiation from siblings like delete_dubbing, dubbing_project_delete, or dubbing_transcript_segment_delete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the other dubbing deletion tools, nor any stated prerequisites. The agent is left to infer that this removes a language target within a project and nothing more.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_language_getC
Read-onlyIdempotent

Get Dubbing Language Target

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesIdentifier of the parent dubbing project.
language_idYesIdentifier of the language target to fetch.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false. The description adds no behavioral context beyond those structured hints, such as error behavior, access constraints, or what a 'target' represents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded phrase with no wasted words. It is appropriately concise for a simple GET operation, though its extreme terseness leaves little context beyond the operation name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, two-parameter read tool with full schema coverage and rich read-only annotations, the definition is minimally complete. However, it does not clarify what a dubbing language target contains or how it relates to the parent project, which leaves a small context gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both project_id and language_id are documented directly in the schema. The description adds no parameter meaning beyond the schema, so the baseline of 3 is appropriate when structured fields already do the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (Dubbing Language Target), making the core operation understandable. It does not differentiate from sibling tools such as dubbing_language_list, dubbing_language_create, or dubbing_language_delete, but it is clear enough for an agent to identify it as a single-resource fetch.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no guidance on when to use this tool versus alternatives. It does not mention dubbing_language_list for enumeration or dubbing_project_get for parent project data, nor does it state prerequisites beyond the schema's required IDs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_language_listC
Read-onlyIdempotent

List Dubbing Language Targets

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoPass the `next_cursor` from a previous response to fetch the page after it. Omit for the first page.
statusNoFilter to targets in this status: `queued`, `processing`, `completed`, `stale`, or `failed`. Omit to return every status.
page_sizeNoNumber of language targets per page. Clamped to between 1 and 100 rather than rejected, so a larger value returns a full page.
project_idYesIdentifier of the parent dubbing project.

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive, and openWorld, so the safety profile is covered structurally. The description adds nothing about pagination behavior, result ordering, or scoping beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no wasted words, which satisfies conciseness, but the terseness borders on under-specification rather than efficient front-loading of useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter, paginated list tool with no output schema, the description should indicate what a language target is and how pagination/scoping works. None of that context is present, leaving the definition inadequate despite the rich schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema documents cursor, status, page_size, and project_id thoroughly (including clamping and enum-like status values). Per the rubric, the baseline is 3 when the schema does the heavy lifting; the description contributes nothing additional.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Dubbing Language Targets' merely restates the tool name (dubbing_language_list) and title, adding no distinguishing detail. It does convey a read verb and resource, but it is effectively a tautology rather than a clarifying statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as dubbing_language_get, dubbing_project_list, or dubbing_language_create. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_project_createB

Create Dubbing Project Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
keytermsNoKey terms to bias transcription and translation toward (for example, product or brand names). At most 1,000 terms; each term at most 50 characters and 5 words; the characters `<>{}[]\` are not allowed. Terms are trimmed and deduplicated.
model_idNoDubbing model (`dubbing_v1` or `dubbing_v2`) every language target of this project is dubbed with. Defaults to `dubbing_v2`. Fixed at create time — the source is prepared for this model, so neither the project nor an individual target can change it later.
file_pathNoThe source media file to dub: an audio or video file of at most 3 GiB. Provide this or `source_url`, not both. Local path.
referenceNoOptional free-form string (at most 500 characters) to identify the project on your end. Stored and echoed back verbatim; it does not affect the dub.
source_urlNoPublic HTTP(S) URL the source media is fetched from server-side, subject to the same size and format limits as an upload. Provide this or `file`, not both.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
webhook_idsNoIDs of workspace webhooks to notify as this project progresses — the alternative to polling, and what we recommend. Each receives a `dubbing_project_ready` or `dubbing_project_failed` event for the project, and a `dubbing_language_completed` or `dubbing_language_failed` event for every language unde
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.
source_languageNoBCP-47 language tag of the source media; must be a language the transcription model supports. Any region or script subtag is ignored, since transcription is per-language. Omit to auto-detect.
target_languageNoOptional shortcut: also create a language target in this BCP-47 language, queued to start once the project is ready — equivalent to creating the project and then creating one language target. Must be one of the [languages the dubbing model supports](https://elevenlabs.io/docs/help-center/product/dub
transcript_pathNoEnterprise only. Optional JSON transcript to use instead of transcribing the source: a `{"segments": [...]}` document, at most 20,000 segments and 4 MiB. See [Bring your own transcript](https://elevenlabs.io/docs/eleven-api/guides/how-to/dubbing/bring-your-own-transcript) for the segment fields and
transcript_base64NoBase64 contents for "transcript". Use this when the server cannot read your local disk.
transcript_filenameNoFilename to send for "transcript". Some endpoints infer the audio format from it.

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, openWorldHint=true, and idempotentHint=false. The description adds the important side effect that it 'Spends ElevenLabs credits,' which is beyond annotations. However, it does not cover auth requirements, rate limits, or failure behavior, so it is only partially transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loads the core action 'Create Dubbing Project.' The second fragment 'Spends ElevenLabs credits.' is grammatically incomplete but still concise and informative. There is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex creation tool with 13 parameters and no output schema, the description is far too sparse. It does not explain what the tool returns (e.g., a project identifier), how it relates to language targets or webhooks, or when to choose it over similar creation tools. The rich schema covers parameters, but the description leaves significant contextual gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 13 parameters are fully documented in the input schema. The description adds no parameter-specific meaning beyond what the schema already provides, making the baseline score of 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Create Dubbing Project' states a clear verb and resource. However, it does not distinguish this tool from siblings like create_dubbing or dub, and it omits any scope details such as what a 'dubbing project' encompasses.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives such as create_dubbing or dub, nor does it mention prerequisites or context. It only notes the credit cost, which is not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_project_deleteC
DestructiveIdempotent

Delete Dubbing Project

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesIdentifier of the dubbing project to delete.

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry the important traits: destructiveHint=true, idempotentHint=true, readOnlyHint=false. The description adds nothing beyond them — no note on irreversibility, cascade effects on transcripts/languages, or required permissions. With annotations covering safety, the bar is lower, but zero added context still warrants a low score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase with no filler, so nothing is wasted, but its brevity reflects under-specification rather than disciplined conciseness. There is no front-loaded rationale or scope statement to structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world delete operation with no output schema, the description should at least convey consequences or prerequisites; it conveys none. The annotations supply the safety signal, but the surrounding context for a destructive mutation is thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single project_id parameter is fully documented in the schema, so the description is not required to compensate. It adds no syntax, format, or sourcing hint beyond the schema, which lands at the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Delete Dubbing Project' is a verbatim restatement of the tool name, which is the textbook tautology case. It does not distinguish itself from adjacent siblings like delete_dubbing, delete_project, or dubbing_language_delete, so an agent gets no help choosing among them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus alternatives, no prerequisites, and no mention of the related getter/creator siblings (dubbing_project_get, dubbing_project_create). The description provides literally no guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_project_getC
Read-onlyIdempotent

Get Dubbing Project

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesIdentifier of the dubbing project to fetch.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered by structured data. The description adds no behavioral context beyond that, not even confirming a single-project fetch returns one record.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is a four-word fragment with no real structure or front-loaded information. It is short, but only because it is under-specified rather than because every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter read tool with full schema coverage and rich annotations, the sparse description is minimally adequate, but with no output schema the agent gets no hint about what the fetched project contains. Slight gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is a single required parameter (project_id) with an in-schema description ('Identifier of the dubbing project to fetch'). Per the baseline rule, a 3 is appropriate; the description adds nothing further.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Dubbing Project' merely restates the tool name and title, effectively a tautology. It does hint at fetching a dubbing project, but gives no scope, no differentiation from siblings like dubbing_project_list, dubbing_project_create, or dubbing_project_delete, so an agent gains nothing beyond the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as dubbing_project_list (to enumerate projects) or dubbing_project_create/delete. The agent must infer usage purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_project_listC
Read-onlyIdempotent

List Dubbing Projects

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoPass the `next_cursor` from a previous response to fetch the page after it. Omit for the first page.
statusNoFilter to projects in this status: `queued`, `preparing`, `ready`, or `failed`. Omit to return every status.
page_sizeNoNumber of projects per page. Clamped to between 1 and 100 rather than rejected, so a larger value returns a full page.
sort_directionNoSort by creation time; newest first by default.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds no further behavioral context, such as pagination behavior or return shape, leaving it entirely dependent on structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded phrase with zero waste. While extremely terse, it is appropriately concise for a list tool and avoids repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should at least hint at what is returned (a paginated list of dubbing projects) and how pagination works. The current description omits return value semantics and filtering context, leaving gaps for an agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (cursor, status, page_size, sort_direction) are fully documented in the schema. The description contributes no additional meaning, which is acceptable given the high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Dubbing Projects' uses a clear verb and resource, distinguishing it from create/delete/get siblings. However, it does not explicitly differentiate itself from other list-like tools such as list_dubs or dubbing_language_list, leaving sibling selection to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives (e.g., dubbing_project_get for a single project, or list_dubs). It states only what the tool does, not the context or conditions that select it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_target_transcript_getC
Read-onlyIdempotent

Get Dubbing Target Transcript

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesIdentifier of the dubbing project.
language_idYesIdentifier of the language target.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that – no mention of what the transcript contains, whether it may be empty while generation is pending, or whether it requires the project/language to already exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no filler, so nothing needs trimming. However, its brevity reflects under-specification rather than discipline – there is no front-loaded scope, return info, or routing detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with full annotation coverage and no output schema, the structured data carries most of the load, but the description still fails to distinguish this tool from its many transcript siblings or to note any precondition (e.g. transcript must exist or be generated). It is too thin to route an agent reliably in a 200+ tool namespace.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both project_id and language_id documented in the schema itself, so the description is not required to compensate. It adds no meaning beyond the schema, which is the baseline-3 case for fully documented parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Dubbing Target Transcript' only restates the tool name/title in slightly reworded form. It identifies a resource and a retrieval verb, but provides zero differentiation from the many sibling retrieval tools in the dubbing/transcript family (dubbing_transcript_get, get_dubbing_transcripts, get_dubbed_transcript_file). An agent cannot tell which transcript variant to pick from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus siblings such as dubbing_transcript_get or dubbing_target_transcript_regenerate. Retrieval intent is implied by 'Get', but no prerequisites, context, or alternatives are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_target_transcript_regenerateB

Regenerate Dubbing Target Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesIdentifier of the dubbing project.
language_idYesIdentifier of the language target.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the mutation and non-idempotence profile is partly covered. The description adds genuinely useful non-annotation context - that the call spends ElevenLabs credits - but omits whether an existing transcript is overwritten or merged, which matters given the non-destructive hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the verb front-loaded and zero filler; the awkward phrasing ('Regenerate Dubbing Target Spends ElevenLabs credits') is grammatically clumsy but costs little.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent mutation with no output schema, the description covers the key cost side effect but leaves the most consequential unknown - whether regeneration replaces existing transcript segments or only fills gaps - unstated, which the sibling edit tools make a relevant question.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both project_id and language_id documented in the schema itself. The description adds no parameter-level information, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Regenerate') and resource ('Dubbing Target' transcript), which distinguishes it from the sibling read tools like dubbing_target_transcript_get and the manual edit tools dubbing_target_transcript_segment_update. However, it never explicitly contrasts itself with those siblings, so an agent must infer the boundary from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this regeneration versus manually editing segments via dubbing_target_transcript_segment_update/segments_update, and no mention of prerequisites or when-not-to-use. The only guidance is the implied cost warning.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_target_transcript_segments_updateC

Update Dubbing Target Transcript Segments Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
segmentsYesMap of segment ID to the translation edit to apply to that segment. At least one entry and at most 500.
project_idYesIdentifier of the dubbing project.
language_idYesIdentifier of the language target.

TDQS

C2.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false and destructiveHint=false, so the safety profile is covered. The description's one additive fact -- that the call spends ElevenLabs credits -- is genuinely useful context not present in the annotations. However, it says nothing about whether the operation is synchronous, triggers re-dubbing, or why non-idempotency matters for credit consumption.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but the two clauses are jammed into one unpunctuated string ('Segments Spends ElevenLabs credits'), which reads as a formatting defect rather than deliberate brevity. The cost warning should be its own sentence, front-loaded or clearly separated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent mutation that consumes credits and has no output schema, the description is too thin: it does not say what the update affects, whether it kicks off regeneration or a new dub, or what the caller should expect on success. The credit warning is the only substantive content.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: project_id, language_id, and the segments map (with the 1-500 entry bound) are all documented in the schema itself. The description contributes nothing about parameter meaning. Per the high-coverage baseline, 3 is the correct score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description largely restates the tool name ('Update Dubbing Target Transcript Segments' vs. dubbing_target_transcript_segments_update) without adding a verb+resource explanation an agent couldn't already infer. The only new information is the trailing credit-cost note. It also fails to distinguish this from the very similar sibling dubbing_target_transcript_segment_update (singular) or dubbing_transcript_segments_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives despite several near-identical siblings (segment vs. segments updates, regenerate, transcript_get). The agent must guess whether to use this bulk update or the single-segment variant.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_target_transcript_segment_updateC

Update Dubbing Target Transcript Segment Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesIdentifier of the dubbing project.
segment_idYesIdentifier of the segment to edit.
language_idYesIdentifier of the language target.
translationNo

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly=false, destructive=false, idempotent=false, and openWorld=true. The description adds a useful credit-cost warning, but does not explain replacement semantics, null translation behavior, or what else changes in the segment after update.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is very short and front-loads the action, but the sentence is malformed and ambiguous ('Update ... Segment Spends ElevenLabs credits'), which weakens structure despite the lack of fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent mutation with no output schema and incomplete parameter descriptions, the definition omits what fields are updated, how null translation behaves, and what the update affects. Only the credit-cost warning is provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema documents 3 of 4 parameters, but the translation parameter is undocumented and nullable while the description provides no field-level meaning. The description does not compensate for the missing parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb ('Update') and resource ('Dubbing Target Transcript Segment'), and adds a credit-consumption note. It does not distinguish this singular segment update from sibling dubbing_target_transcript_segments_update or dubbing_transcript_segment_update, so sibling differentiation is missing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this instead of sibling update, regenerate, delete, or plural-segment tools. The only additional statement is a cost warning, not a usage condition or alternative-selection rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_transcript_getC
Read-onlyIdempotent

Get Dubbing Transcript

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesIdentifier of the dubbing project.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered structurally. The description adds nothing beyond that — no indication of what transcript is selected (source vs target language), whether it can be empty, or remote-fetch behavior implied by openWorldHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a four-word sentence with zero waste, but it is not appropriately sized for the information it must carry — it is under-specified rather than concise, conveying no more than the tool name already does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no explanation of what the returned transcript contains or how it relates to the sibling target/source transcript tools, the definition leaves key questions unanswered for a dubbing-domain retrieval tool. An agent would have to guess which of the several transcript endpoints to call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single project_id parameter, which the schema documents as the identifier of the dubbing project. With full schema coverage and only one parameter, the baseline of 3 applies; the description adds no format or sourcing detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Dubbing Transcript" merely restates the tool name (title is even identical) with no added specificity. It does not distinguish this tool from close siblings such as get_dubbing_transcripts, get_dubbed_transcript_file, dubbing_target_transcript_get, or get_transcript_by_id, so an agent cannot tell which transcript resource this returns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives, no preconditions, and no mention of the dubbing project state required. In a tool family with at least four near-identical transcript getters, the absence of routing guidance is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_transcript_segment_addB

Add Dubbing Transcript Segment Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe text of the new segment.
end_sYesEnd time of the segment, in seconds.
start_sYesStart time of the segment, in seconds.
project_idYesIdentifier of the dubbing project.
speaker_idYesIdentifier of the segment's speaker.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-idempotent, non-destructive, open-world write operation. The description adds a meaningful behavioral trait not in annotations: it spends ElevenLabs credits. However, it does not describe validation, duplicate handling, or what the tool returns, so it is useful but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It is slightly awkward as a run-on ('Add Dubbing Transcript Segment Spends ElevenLabs credits'), but both clauses earn their place by covering the action and the cost.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With full schema coverage and annotations covering the write safety profile, the description only needs to add contextual value. It does add the credit cost, but it omits prerequisites such as needing an existing project and speaker, and does not clarify that this adds rather than replaces segments. It is minimally viable but thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each of the five required parameters clearly described. The description adds no further parameter meaning, so the baseline of 3 is appropriate because the schema already carries the full parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Add Dubbing Transcript Segment'), so an agent knows it creates a new transcript segment. It does not explicitly differentiate from sibling tools like dubbing_transcript_segment_update or dubbing_transcript_segment_delete, so it is clear but lacks sibling routing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance is the cost side effect ('Spends ElevenLabs credits'), which implies use should be deliberate, but there is no when-to-use condition, when-not-to-use condition, or mention of alternatives such as update or delete. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_transcript_segment_deleteC
DestructiveIdempotent

Delete Dubbing Transcript Segment

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesIdentifier of the dubbing project.
segment_idYesIdentifier of the segment to remove.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, idempotentHint=true, and openWorldHint=true, covering the safety profile. The description adds nothing about permanence, whether surrounding segment timings shift, or required permissions on top of that structured data. No contradiction, but no added value either.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no waste, but it is under-specified rather than genuinely concise. Being short is not itself a virtue when no information is conveyed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world mutation with no output schema, the description should at minimum state that the segment is permanently removed and what context it affects. The annotations carry the safety signal, but the agent still lacks confirmation of scope and irreversibility from the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters (project_id, segment_id), so the schema already carries the semantics. The description contributes no additional meaning, which is the baseline 3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a restatement of the tool title, adding no detail beyond the name itself. It identifies no scope, no distinguishing behavior from the many sibling delete tools (delete_segment, delete_transcript_by_id, dubbing_project_delete), so an agent must open the schema to know what it operates on.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative guidance. Nothing tells the agent how this relates to dubbing_transcript_segment_update, dubbing_transcript_segments_update, or delete_segment.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_transcript_segments_updateC

Update Dubbing Transcript Segments Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
segmentsYesMap of segment ID to the partial update to apply to that segment. At least one entry and at most 500.
project_idYesIdentifier of the dubbing project.

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description does add one genuinely new behavioral fact beyond annotations: the call spends ElevenLabs credits. However, it omits batching semantics, partial-update behavior, failure modes, and credit magnitude, so it is only a partial supplement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is very short, which is good, but the single sentence is a broken run-on that fuses the action with the cost clause ('...Segments Spends ElevenLabs credits'). The credit warning would land better as a separate, clearly punctuated sentence rather than being spliced onto the operation name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, non-idempotent tool taking a nested map of segment IDs to partial updates, the description should at least explain the partial-update semantics, the 500-entry cap consequence, or the response behavior. The schema documents the map shape but the description leaves behavioral gaps beyond credit cost, and there is no output schema to fill them.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (project_id and the segments map) are already documented in the schema, including the 1-to-500 entry bound. The description adds no parameter-level meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the verb (Update) and resource (Dubbing Transcript Segments) but is essentially the tool name with underscores removed, so it reads close to a tautology. Critically, it does not distinguish itself from the near-identical siblings dubbing_transcript_segment_update, dubbing_target_transcript_segments_update, and dubbing_target_transcript_segment_update, so an agent cannot tell which one to pick from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the singular segment update or the target-transcript variants, and no prerequisites or context are stated. The only usage-adjacent signal is the credit cost, which warns of expense but does not route the agent to the correct alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dubbing_transcript_segment_updateC

Update Dubbing Transcript Segment Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
textNo
end_sNo
start_sNo
project_idYesIdentifier of the dubbing project.
segment_idYesIdentifier of the segment to edit.
speaker_idNo

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses one non-obvious behavioral fact not present in the annotations: the call spends ElevenLabs credits. That is genuine added value beyond the structured readOnly/openWorld/idempotent hints. However, it is vague (no amount, no condition, no note on whether the dubbed audio/target transcript is regenerated) and the rest of the mutation behavior is left unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and wastes few words, but the phrasing is a malformed run-on ('...Segment Spends ElevenLabs credits.') that blurs the action with the side effect. It is terse rather than well front-loaded, and the grammar hurts readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent mutation tool with 6 parameters, low schema coverage, and no output schema, the description is insufficient. It never states which fields are editable, what an update implies for the project/segment, whether credits are consumed unconditionally, or how it relates to the sibling segment tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%: only project_id and segment_id are documented in the schema, while text, start_s, end_s, and speaker_id are undocumented. The description adds no information about any parameter, so it fails to compensate for the coverage gap on a 6-parameter editing tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Update Dubbing Transcript Segment'), which is clear on its own. However, it does not distinguish itself from very close siblings such as dubbing_transcript_segments_update (plural), dubbing_target_transcript_segment_update, dubbing_transcript_segment_add, or dubbing_transcript_segment_delete, so an agent must infer which segment tool applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. Given the dense cluster of near-identical segment tools (add/delete/update, singular/plural, target vs. non-target), the absence of routing guidance is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

duplicate_agent_routeC

Duplicate Agent

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
agent_idYesThe id of an agent. This is returned on agent creation.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, openWorldHint=true, so safety is partly covered structurally. The description adds nothing, though: it does not say whether the duplicate copies settings, knowledge base links, or triggers side effects, which matters for a non-idempotent mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words with no structure. Brevity here reflects under-specification rather than economy; nothing is front-loaded because nothing is stated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a non-idempotent write with an undocumented optional parameter, no output schema, and no explanation of what the duplicated result contains. Given that complexity, the description is far too thin to let an agent invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: agent_id is documented as 'the id of an agent', but the optional 'name' parameter has no description anywhere. The tool description does not compensate by explaining what name does (e.g. naming the clone or defaulting to the original name).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Duplicate Agent' merely restates the tool name. It conveys the verb (duplicate) and object (agent route), but adds no scope, no detail on what gets copied, and no differentiation from siblings like create_agent_route or create_agent_draft_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites (e.g. the referenced agent must exist), and no routing to alternatives such as create_agent_route for building a fresh agent. The agent is left to infer everything from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_chapterC

Update Chapter Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
contentNo
chapter_idYesThe ID of the chapter.
project_idYesThe ID of the Studio project.

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, covering the mutation profile. The description adds one genuinely useful non-annotation fact: the call spends ElevenLabs credits, a cost implication. It still omits what happens to existing content, whether name/content are full replacements, and any auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short, front-loaded sentence with no wasted words, which is structurally sound. But for a four-parameter mutation tool the terseness crosses into under-specification rather than genuine conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent mutation with no output schema, 50% schema coverage, and undocumented nested content, the description should do considerably more. It leaves the agent without knowledge of which fields are editable, how content is applied, or the response shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: project_id and chapter_id are documented, but name and content (a nested blocks object) are not. The description adds no meaning for any parameter, so it fails to compensate for the undocumented half, including the important question of whether content replaces or merges.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource ('Update Chapter'), so an agent knows it mutates a chapter object. However it never says which fields can be changed (name/content) and does not distinguish itself from siblings like edit_project_content, update_document_route, or add_chapter. Purpose is identifiable but not differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as edit_project_content, add_chapter, or convert_chapter_endpoint. No prerequisites, no conditions, no exclusions are given. The agent must infer applicability entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_projectC

Update Studio Project Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe name of the Studio project, used for identification only.
titleNo
authorNo
project_idYesThe ID of the Studio project.
isbn_numberNo
volume_normalizationNoWhen the Studio project is downloaded, should the returned audio have postprocessing in order to make it compliant with audiobook normalized volume requirements
default_title_voice_idYesThe voice_id that corresponds to the default voice used for new titles.
default_paragraph_voice_idYesThe voice_id that corresponds to the default voice used for new paragraphs.

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true. The description adds a useful behavioral note that this operation spends ElevenLabs credits, which is not captured in the annotations, but it omits permissions, reversibility, and side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short and front-loads the action. It is efficient, though the second clause reads as an abrupt fragment and could be integrated more clearly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter mutation tool with no output schema, the description is incomplete. It does not explain prerequisites, how unspecified fields are handled, return behavior, or why an agent should choose this over edit_project_content.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds no parameter meaning beyond the schema. With 8 parameters and 63% schema description coverage, several fields such as title, author, and isbn_number lack schema descriptions, and the tool description does not compensate for that gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: update a Studio Project. However, it does not distinguish this tool from closely named siblings such as edit_project_content, add_project, or delete_project, so sibling differentiation is missing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool, when not to use it, or which alternative sibling to choose. There is no usage context beyond the bare action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_project_contentC

Update Studio Project Content Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
from_urlNoAn optional URL from which we will extract content to initialize the Studio project. If this is set, 'from_url' and 'from_content' must be null. If neither 'from_url', 'from_document', 'from_content' are provided we will initialize the Studio project as blank.
project_idYesThe ID of the Studio project.
auto_convertNoWhether to auto convert the Studio project to audio or not.
from_content_jsonNoAn optional content to initialize the Studio project with. If this is set, 'from_url' and 'from_document' must be null. If neither 'from_url', 'from_document', 'from_content' are provided we will initialize the Studio project as blank. Example: [{"name": "Chapter A", "blocks": [{"sub_type": "p", "no
from_document_pathNoAn optional .epub, .pdf, .txt or similar file can be provided. If provided, we will initialize the Studio project with its content. If this is set, 'from_url' and 'from_content' must be null. If neither 'from_url', 'from_document', 'from_content' are provided we will initialize the Studio project as
from_document_base64NoBase64 contents for "from_document". Use this when the server cannot read your local disk.
from_document_filenameNoFilename to send for "from_document". Some endpoints infer the audio format from it.

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, covering the safety profile. The description adds one genuinely useful behavioral fact not in annotations — that the operation spends ElevenLabs credits — but says nothing about whether existing content is replaced or preserved, or about prerequisites.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and both ideas (purpose and credit cost) are relevant, but they are concatenated into a single ungrammatical fragment ('Update Studio Project Content Spends ElevenLabs credits.'), so the front-loading is unclear and the sentence reads awkwardly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 100% schema coverage and annotations carrying the safety profile, the description is minimally sufficient to invoke the tool. But for a non-idempotent mutation that consumes credits, it omits what the update actually does to existing project content and any prerequisite conditions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all seven parameters (including the content-source options from_url/from_content_json/from_document_*) are documented in the schema. The description adds no parameter meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Update Studio Project Content' gives a verb (update) and a resource (Studio Project Content), so the general intent is recoverable. However, 'content' is left undefined against siblings like edit_project, edit_chapter, and audio_native_project_update_content_endpoint, and the run-on wording muddies the statement, so no sibling differentiation is provided.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative routing guidance. The credit-cost note is a warning rather than usage guidance, and nothing tells the agent how this differs from edit_project or the audio-native content endpoints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_pvc_voiceD

Edit Pvc Voice

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoThe name that identifies this voice. This will be displayed in the dropdown of the website.
labelsNo
languageNoLanguage used in the samples.
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
descriptionNo

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the mutation and safety profile is partially covered by structured data. The description contributes nothing on top of this — no mention of which fields are mutated, whether the change is reversible, or any requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words that are technically concise but convey no information; this is under-specification rather than effective brevity. No meaningful content is front-loaded because there is no content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter mutation tool with no output schema and gaps in parameter documentation, the description omits everything an agent needs: what is editable, side effects, and prerequisites. It is effectively absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 60% and the description supplies no parameter meaning at all. name, language, and voice_id have inline schema descriptions, but labels and description fields are undocumented in both the schema and the description, leaving the agent to guess.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Edit Pvc Voice" merely restates the tool name in title case; it is a tautology with no added verb, scope, or constraint. It does nothing to distinguish this tool from alphabetically adjacent siblings such as edit_pvc_voice_sample, add_pvc_voice_samples, or delete_pvc_voice_sample.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool, when not to, or which sibling handles which voice-editing scenario. An agent must infer everything from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_pvc_voice_sampleD

Update Pvc Voice Sample

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
file_nameNo
sample_idYesSample ID to be used
trim_end_timeNo
trim_start_timeNo
selected_speaker_idsNo
remove_background_noiseNoIf set will remove background noise for voice samples using our audio isolation model. If the samples do not include background noise, it can make the quality worse.

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond that: it doesn't say what is altered (audio, trim range, speaker selection, file name) or whether changes are reversible, which matters for a non-idempotent mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but this is under-specification rather than conciseness — a three-word restatement of the title that front-loads no useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutation tool with seven parameters, no output schema, and undocumented fields needs far more than a restated title to be called correctly. Nothing about prerequisites, effect scope, or the trim/speaker options is conveyed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Seven parameters with only 43% schema description coverage: voice_id, sample_id, and remove_background_noise are documented, while file_name, trim_start_time, trim_end_time, and selected_speaker_ids are undocumented in both schema and description. The description contributes no meaning at all, leaving half the surface unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is the tool name restated as a phrase ('Update Pvc Voice Sample'); it names no resource meaningfully beyond the title and gives no scope or distinguishing detail. Non-English/duplicate casing aside, it is a tautology rather than an explanation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus siblings such as delete_pvc_voice_sample, add_pvc_voice_samples, get_pvc_sample_audio, or edit_pvc_voice. An agent must infer entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_service_account_api_keyD

Edit Service Account Api Key

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
api_key_idYes
is_enabledNoWhether to enable or disable the API key.
allowed_ipsNoList of IP addresses or CIDR ranges allowed to use this API key. Each entry may be a CIDR range (e.g. '10.0.0.0/24') or a bare IP address (normalized to /32 or /128). On create, omit or pass null to allow all IPs. On update, omit to leave the allowlist unchanged, or pass "clear" to remove it.
permissionsNoThe permissions of the XI API.
character_limitNoThe character limit of the XI API key. If provided this will limit the usage of this api key to n characters per month where n is the chosen value. Requests that incur charges will fail after reaching this monthly limit.
service_account_user_idYes
third_party_disable_allowedNoWhether the holder of this key may disable it via the self-disable endpoint. On create, omit or pass null to use the workspace's default (enabled for non-Enterprise plans, disabled for Enterprise plans). On update, omit to leave it unchanged, or pass "clear" to reset it to the workspace default. Onl

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-readonly, open-world, non-idempotent, non-destructive mutation, but the description adds nothing beyond that. It does not disclose which fields can be modified, whether changes are reversible, what permissions are required, or any side effects of the operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but it is under-specified rather than concise. Every sentence should earn its place, and this single phrase repeats the name without adding information, making it a wasted opportunity for front-loaded guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with 8 parameters, nested objects, and no output schema, the description is completely inadequate. It communicates nothing about what an edit does, what values are valid, or how to interpret success, leaving the agent with no context beyond the name and partial schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 63%, with several parameters (name, api_key_id, service_account_user_id) lacking any description in the schema. The tool description does not compensate by explaining the meaning, format, or behavior of any parameter, leaving gaps for an agent trying to construct a valid edit request.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Edit Service Account Api Key' is a verbatim restatement of the tool name, offering no additional detail about what fields can be edited, what the operation affects, or how it differs from creation/deletion siblings. It is a tautology rather than a clarifying specification.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like create_service_account_api_key, delete_service_account_api_key, or get_service_account_api_keys_route. No prerequisites, no conditions, no exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_voiceD

Edit Voice

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe name that identifies this voice. This will be displayed in the dropdown of the website.
labelsNoLabels for the voice. Keys can be language, accent, gender, or age.
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
descriptionNoA description of the voice.
files_pathsNoAudio files to add to the voice Local paths.
files_filenamesNoFilenames to send for "files". Some endpoints infer the audio format from them.
files_base64_listNoBase64 contents for "files", one entry per file. Use this when the server cannot read your local disk.
moderate_metadataNoRun synchronous LLM moderation over the voice name and description when they change. Has no effect unless the voice_library_metadata_moderation feature flag is enabled for the user.
remove_background_noiseNoIf set will remove background noise for voice samples using our audio isolation model. If the samples do not include background noise, it can make the quality worse.

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond that: no note on which fields are mutated, whether omitted fields are preserved, or that files_paths/files_base64_list upload audio into an existing voice.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words with zero waste but also zero substance. This is under-specification rather than conciseness, matching the calibration case for a name-only description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A nine-parameter mutation tool with file-upload and moderation side effects, no output schema, and many overlapping voice siblings needs far more than a two-word description. Nothing an agent needs in order to call it safely or correctly is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema itself documents all nine parameters (name, labels, voice_id, description, files_paths, files_filenames, files_base64_list, moderate_metadata, remove_background_noise). Per the high-coverage baseline, 3 is correct even though the description contributes no parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Edit Voice" is a bare restatement of the tool name/title and gives no more information than the identifier already does. It does not state what aspect of the voice is edited, what the required voice_id+name pair implies, or how it differs from siblings like edit_voice_settings, edit_pvc_voice, or update_speaker.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. With ~40 voice-related siblings (add_voice, create_voice, delete_voice, edit_pvc_voice, edit_voice_settings, get_voice_by_id), an agent has no basis for choosing this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_voice_settingsD

Edit Voice Settings

ParametersJSON Schema
NameRequiredDescriptionDefault
speedNo
styleNo
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
stabilityNo
similarity_boostNo
use_speaker_boostNo

TDQS

D1.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, covering the safety profile. The description adds no behavioral context beyond the word Edit, such as what changes, whether settings can be reset, or what permissions are required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but that brevity reflects under-specification rather than efficiency. It lacks the information an agent needs to select and invoke the tool correctly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with six parameters, low schema coverage, no output schema, and only partial annotations, the description is completely inadequate. It omits what settings can be changed, valid ranges, and expected behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% for six parameters, so the description must compensate heavily. It mentions no parameters at all, leaving fields like speed, stability, similarity_boost, style, and use_speaker_boost undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a near-verbatim restatement of the tool name and title. It states a generic verb (Edit) and resource (Voice Settings) but does not specify what settings, how they are edited, or how this differs from siblings such as edit_voice or update_settings_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives like edit_voice, get_voice_settings, or update_settings_route. There are no prerequisites, exclusions, or contextual cues.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_workspace_webhook_routeD

Update Workspace Webhook

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe display name of the webhook (used for display purposes only).
eventsNo
webhook_idYesThe unique ID for the webhook
is_disabledYesWhether to disable or enable the webhook
retry_enabledNo
request_headersNo

TDQS

D1.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered externally. The description adds nothing beyond that: no note on what is replaced or overwritten, no auth/permission requirements, no mention that updates are non-idempotent despite the annotation. It neither contradicts nor enriches the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a four-word fragment with no wasted words, but this is under-specification rather than genuine conciseness. There is no front-loaded statement of behavior because there is no content at all.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a non-idempotent mutation tool with six parameters (three required), half of them undocumented, no output schema, and no annotations-level detail repeated or extended in the description. Nothing an agent needs to invoke it correctly is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50% (events, retry_enabled, and request_headers are undocumented), so the description is expected to compensate — and it does not. It provides zero information about any of the six parameters, including which three are required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update Workspace Webhook' is essentially a restatement of the tool name and title (edit_workspace_webhook_route / 'Edit Workspace Webhook Route'). It conveys only the bare verb+resource and offers nothing to distinguish it from siblings such as create_workspace_webhook_route, get_workspace_webhooks_route, or delete_workspace_webhook_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance whatsoever on when to use this tool, no prerequisites, and no mention of the sibling routes (create/get/delete_workspace_webhook_route) an agent must choose between. The agent is left to infer everything from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_batch_callC
Read-onlyIdempotent

Export Batch Call Results

ParametersJSON Schema
NameRequiredDescriptionDefault
batch_idYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered by structured data. The description adds nothing beyond that: it says nothing about export format, size/pagination, whether the export is synchronous, or where results are delivered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single four-word phrase with no wasted words, but this is under-specification rather than conciseness. There is no structure to front-load because there is essentially no content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter export tool with no output schema and no parameter documentation, the description should at least state what is exported, in what form, and how it relates to get_batch_call / get_workspace_batch_calls. None of that is present, leaving the agent unable to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter batch_id has no description anywhere. The phrase 'Batch Call Results' implies batch_id identifies the batch, but adds no format, constraints, or example values, so the description fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Export Batch Call Results' is essentially a restatement of the tool name export_batch_call, adding only the word 'Results'. It does not distinguish this tool from siblings like get_batch_call, get_workspace_batch_calls, retry_batch_call, or cancel_batch_call, nor does it explain what form the export takes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites, and no routing to alternatives such as get_batch_call for metadata versus this tool for results. The agent must infer everything from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

forced_alignmentC

Create Forced Alignment Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe text to align with the audio. The input text can be in any format, however diarization is not supported at this time.
file_pathNoThe file to align. All major audio formats are supported. The file size must be less than 1GB. Local path. Required for this call.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the mutation and safety profile is covered. The description adds one genuinely useful behavioral fact not present in the structured data: the call consumes ElevenLabs credits. It says nothing about the 1GB file limit, required inputs, or processing/Latency characteristics, but that content is largely in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two fragments and front-loads the action, with zero filler. However, the phrasing "Create Forced Alignment Spends ElevenLabs credits" is grammatically garbled, reading as a run-on rather than two clean statements, which slightly undercuts clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering safety/write mode and a fully documented schema, the description needs mainly to state purpose and notable costs. It does capture the credit cost but omits any routing guidance against the numerous sibling audio tools, leaving the agent to infer when alignment is the right choice.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each parameter (text, file_path, file_base64, file_filename) is documented in the schema itself. The description adds no parameter meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a verb ("Create") and resource ("Forced Alignment"), so the agent can infer this is an audio/text alignment operation. However, it offers no scope detail and no differentiation from siblings (e.g., transcribe, speech_to_text), which would help an agent distinguish alignment from plain transcription.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as transcribe or speech_to_text. The only added sentence is about billing, not usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generateC

Compose Music Spends ElevenLabs credits. Returns audio/* bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
promptNo
model_idNo
finetune_idNo
lyrics_textNo
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
music_promptNoComposition plan for the `music_v1` model. Using this field with any other model will result in an error.
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. Use "auto" (the default) to let the API pick the best format for the selected model: mp3_44100_128 for v1 models and mp3_48000_192 for v2 models.
sign_with_c2paNoWhether to sign the generated song with C2PA. Applicable only for mp3 files.
generation_modeNo
music_length_msNo
composition_planNo
finetune_strengthNoHow strongly the finetune influences the generation. Defaults to 1.0 (full strength). Lower values soften the influence of the finetune, leaving more room for prompt-level steering. Only meaningful when `finetune_id` is also provided.
force_instrumentalNoIf true, guarantees that the generated song will be instrumental. If false, the song may or may not be instrumental depending on the `prompt`. Can only be used with `prompt`.
use_phonetic_namesNoIf true, proper names in the prompt will be phonetically spelled in the lyrics for better pronunciation by the music model. The original names will be restored in word timestamps.
store_for_inpaintingNoWhether to store the generated song for inpainting.
respect_sections_durationsNoControls how strictly section durations in the `composition_plan` are enforced. Only used with `composition_plan` and only applies to `music_v1`; for `music_v2` and `music_v2_5` section durations are always enforced and this is ignored. When false for `music_v1`, the model may adjust individual sect

TDQS

C2.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare it is a non-read-only, non-idempotent, open-world operation, so the description earns real credit by disclosing that it spends ElevenLabs credits and returns audio/* bytes. The credit-consumption cost is exactly the kind of trait annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loads purpose, but the first sentence is an unpunctuated run-on ("Compose Music Spends ElevenLabs credits") that reads as two fragments fused together, hurting clarity more than length does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter generative tool with overlapping model_ids, music_prompt vs composition_plan, and multiple generation modes, two terse sentences leave major decisions unexplained. With no output schema and only 53% parameter coverage, the description should have carried far more of the complexity burden.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 17 parameters and 53% schema coverage, the schema already documents the heavier fields, and the description only echoes output_path behavior that the schema itself explains in more detail. It adds no model-selection or mode guidance beyond the structured data, so baseline 3 fits.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Compose Music" gives a verb and resource, so the core purpose is legible despite the generic name "generate". However, it offers no differentiation from the many composition siblings (compose_detailed, compose_plan, stream_compose, sound_generation), so an agent cannot tell which composition tool to reach for from this text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives like compose_detailed or stream_compose. The only usage-shaped hint is the mechanical note about output_path, which is about output handling rather than tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_conversation_ticket_routeC
Read-onlyIdempotent

Get Agent Conversation Ticket

ParametersJSON Schema
NameRequiredDescriptionDefault
agentqa_ticket_idYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, covering the safety and idempotency profile. The description adds nothing beyond that, not even what a 'ticket' contains or whether the lookup can fail for a missing/foreign ID.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase, but the brevity comes from under-specification rather than efficiency. Nothing is front-loaded because nothing of substance is stated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description would need to say what is returned, and the lone parameter is undocumented. For a retrieval tool in a large sibling set, this leaves the agent without enough to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter agentqa_ticket_id has 0% schema description coverage, and the description does not compensate at all — it never mentions the ID, its format, or where it comes from. With one undocumented param, the description should carry the burden and does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Agent Conversation Ticket' essentially restates the tool name and title without adding specificity. It names a verb and resource, but gives no clue how it differs from closely related siblings such as list_agent_conversation_tickets_route, get_conversation_history_route, or get_conversation_tag_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. An agent must infer from the name alone that this retrieves a single ticket by ID rather than listing them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_knowledge_base_sizeC
Read-onlyIdempotent

Returns The Size Of The Agent'S Knowledge Base

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint. The description adds no additional behavioral context such as what 'size' measures, auth requirements, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single concise sentence with no wasted words. It front-loads the action and is appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read endpoint with rich annotations, the description is mostly adequate. It leaves the meaning of 'size' and the required agent_id undocumented, and with no output schema some ambiguity remains acceptable but noticeable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The one required agent_id parameter has 0% schema description coverage. The description only implies an agent via the possessive 'Agent's' and gives no format, source, or required nature, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear specific verb and resource: returns the size of an agent's knowledge base. It does not explicitly differentiate from siblings such as get_agent_knowledge_base_summaries_route or get_knowledge_base_content, so scope must be inferred.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no when-to-use guidance or alternatives. It does not say when to prefer this over summaries or content endpoints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_knowledge_base_summaries_routeC
Read-onlyIdempotent

Get Knowledge Base Summaries By Ids

ParametersJSON Schema
NameRequiredDescriptionDefault
document_idsYesThe ids of knowledge base documents.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is fully covered. The description adds no behavioral context beyond that — not what a 'summary' contains, nor any batching/error behavior for missing ids.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short and front-loaded, but it is a bare fragment lacking structure or detail. There is no wasted text, yet nothing extra earns its place either.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description should clarify what the returned 'summaries' are and whether they align with the requested ids. With a 1-param required schema and no return-value detail, it is under-specified for what an agent needs to interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single document_ids parameter, which is documented as 'The ids of knowledge base documents.' Baseline 3 applies since the schema does the heavy lifting and the description merely echoes 'By Ids' without adding format or limit details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb (Get) and resource (Knowledge Base Summaries) scoped to ids, so the basic purpose is discernible. However, it largely restates the tool name/title and gives no differentiation from closely named siblings like get_agent_summaries_route, get_agent_response_tests_summaries_route, or get_knowledge_base_content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives such as get_knowledge_base_content or get_agent_knowledge_base_size. No prerequisites, context, or exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_llm_expected_cost_calculationC

Calculate Expected Llm Usage For An Agent

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes
rag_enabledNo
prompt_lengthNo
number_of_pagesNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is partly covered. The description adds nothing on top of that: it does not say whether the result is an estimate, whether it requires prior usage data, or whether it consumes billable resources. For a calculation tool this is a meaningful gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, and the purpose is front-loaded. It is efficient, though efficiency here borders on under-specification rather than true conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With four undocumented parameters, no output schema, and a non-obvious calculation whose semantics matter (what the estimate covers, what units, what happens when optional params are null), the description is far too thin. An agent has almost nothing to act on beyond the name.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Four parameters exist with 0% schema description coverage, so the description carries the full burden and fails it. rag_enabled, prompt_length, and number_of_pages are critical inputs for a cost estimate and none are explained or even mentioned. The phrase 'for an agent' only loosely gestures at agent_id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb (Calculate) and a resource (Expected LLM usage for an agent), so the basic purpose is derivable. However, it is essentially a title-cased restatement of the tool name and offers no differentiation from the very close sibling get_public_llm_expected_cost_calculation. An agent cannot tell the two apart from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the closely related public cost-calculation sibling. The agent is left to infer entirely from the name which of the two cost tools to call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_response_test_routeC
Read-onlyIdempotent

Get Agent Response Test By Id

ParametersJSON Schema
NameRequiredDescriptionDefault
test_idYesThe id of a chat response test. This is returned on test creation.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety and idempotency profile is fully covered. The description adds nothing beyond that, saying nothing about failure behavior (e.g., unknown test_id), authorization requirements, or what the returned test object contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase, so there is no bloat and the key action is front-loaded. But it is so terse that it merely echoes the tool name, earning little beyond brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the burden of describing what is returned (the test's configured prompt/config, status, metadata), and it says nothing about the response shape. For a single-resource getter with no output schema, this leaves a meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single test_id parameter is documented in the schema as the id returned on test creation. The description adds no format, provenance, or lookup semantics beyond that, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a retrieval verb and the resource ('Agent Response Test') scoped by id, so the basic action is identifiable. However, it is essentially a restatement of the tool name with no differentiation from siblings like get_agent_response_tests_summaries_route, list_chat_response_tests_route, or update_agent_response_test_route. It is minimally viable but adds no clarifying detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the list, summaries, create, update, or delete variants that clearly exist among the siblings. The agent must infer from the name alone that this retrieves a single test by id rather than a collection or a related resource.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_response_tests_summaries_routeC

Get Agent Response Test Summaries By Ids

ParametersJSON Schema
NameRequiredDescriptionDefault
test_idsYesList of test IDs to fetch. No duplicates allowed. Prefer at most 1000 IDs per request.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=false and openWorldHint=true, while the description's 'Get' strongly implies a plain read; the description does nothing to clarify that mismatch, nor does it mention pagination, batching limits, or auth. Since annotations carry some signal already, this is not scored 1, but the description adds no behavioral value beyond the title.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short, front-loaded phrase with no padding, which is good. But it is a title fragment rather than an informative sentence, so the brevity reflects under-specification rather than efficient communication.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple by-ID fetch with no output schema, the description should at least say what a 'summary' contains or when to prefer it over the singular/list siblings. Given the dense sibling namespace around agent response tests, the current text leaves the agent guessing about selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter is fully documented in the schema (no duplicates, prefer at most 1000 IDs per request). The description's 'By Ids' merely echoes this, adding no format or batching guidance, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a verb and resource ('Get Agent Response Test Summaries') and adds the scoping qualifier 'By Ids', so the basic intent is clear. However, it reads as a near-restatement of the tool title and does nothing to distinguish it from close siblings such as get_agent_response_test_route (singular test) or list_chat_response_tests_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus the singular get_agent_response_test_route or the list route, and no prerequisites or exclusions. The agent must infer the batch-by-ID use case entirely from the parameter schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_routeC
Read-onlyIdempotent

Get Agent

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.
branch_idNoThe ID of the branch to use
version_idNoThe ID of the agent version to use

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is covered structurally. The description adds no behavioral context beyond those annotations, such as scope, permissions, or what the route represents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two words, but it is under-specified rather than efficiently concise. It lacks the front-loaded useful information an agent needs to distinguish and invoke the tool confidently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool taking agent_id plus optional branch_id and version_id, the description does not explain what is being retrieved or how branch and version affect the result. Annotations cover the safety profile, but the description remains inadequate for an agent to understand the operation's purpose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents agent_id, branch_id, and version_id. The description adds no parameter meaning, but with high schema coverage the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Agent" essentially restates the tool name and title rather than stating what resource is retrieved or what "route" means. It provides no differentiation from the many sibling tools such as get_agents_route, get_agent_summaries_route, or get_branch_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool, when not to use it, or which alternatives exist. It is not misleading, but it supplies no usage context at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agents_routeC
Read-onlyIdempotent

List Agents

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNoFilter agents by tag. Repeat the parameter to match any of several tags.
cursorNoUsed for fetching next page. Cursor is returned in the response.
searchNoSearch by agents name.
sort_byNoThe field to sort the results by
archivedNoFilter agents by archived status
page_sizeNoHow many Agents to return at maximum. Can not exceed 100, defaults to 30.
sort_directionNoThe direction to sort the results
created_by_user_idNoFilter agents by creator user ID. When set, only agents created by this user are returned. Takes precedence over show_only_owned_agents. Use '@me' to refer to the authenticated user.
show_only_owned_agentsNoIf set to true, the endpoint will omit any agents that were shared with you by someone else and include only the ones you own. Deprecated: use created_by_user_id instead.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered by structured data. The description adds nothing on top of that — no mention of pagination behavior, default page size, or that agents shared with you are included by default, which matters for a list endpoint with 9 optional filters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words are not concise, they are under-specified. Brevity here removes information rather than eliminating waste; a caller cannot act on the description alone.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, zero-required list endpoint with no output schema, the description should at least touch on default result size, pagination via cursor, and default inclusion of shared agents. Annotations cover safety but not retrieval semantics, leaving the agent with the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema itself documents each filter, including the '@me' shortcut and the deprecation of show_only_owned_agents. The description contributes no additional parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"List Agents" states a verb and a resource, so the core operation is identifiable. However it gives no scope information (all vs. filtered, owned vs. shared, paginated) and does nothing to distinguish it from siblings like get_agent_summaries_route or get_tool_dependent_agents_route, despite the tool's rich filter set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance whatsoever on when to use this tool versus the many adjacent agent-retrieval tools, no prerequisites, and no mention of the archived/ownership filters that would shape a caller's choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_summaries_routeC
Read-onlyIdempotent

Get Agent Summaries

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idsYesList of agent IDs to fetch summaries for

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that — no note on batching behavior, per-ID error handling, or what happens for unknown/inaccessible agent IDs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but conciseness here reflects under-specification rather than efficiency — a two-word phrase that restates the title earns no informational place. There is no front-loaded detail to guide invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, with full schema coverage and annotations on safety, so the description needn't explain return values. Still, for a batch-fetch route it should clarify that summaries are returned per requested agent ID, which is left entirely implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is only one parameter and schema description coverage is 100%, so the schema already documents agent_ids fully. Baseline 3 applies since the description contributes no additional meaning about ID format, batch limits, or ordering.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Agent Summaries' is essentially a restatement of the tool name with no added specificity. It does not distinguish this route from close siblings such as get_agent_route, get_agents_route, or get_agent_response_tests_summaries_route, nor does it explain what a 'summary' contains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to use this tool versus alternatives, no prerequisites, and no mention of the batch/multi-ID nature implied by the required agent_ids array. The agent must infer all routing decisions from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_test_folder_routeC
Read-onlyIdempotent

Get Agent Test Folder By Id

ParametersJSON Schema
NameRequiredDescriptionDefault
folder_idYesThe folder ID.

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe read operation. The description adds nothing beyond that, but with annotations covering the safety profile, a baseline 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise phrase with no wasted words. It is front-loaded, though it lacks any detail to earn a higher score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that this is a read operation with one required parameter and no output schema, the description should ideally explain what the tool returns or how it differs from similar tools. It provides insufficient context for an agent to confidently select it over siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single parameter folder_id is fully documented in the schema. The description adds no additional meaning about the parameter, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Agent Test Folder By Id' states a verb and resource, but it is essentially a restatement of the tool name and title. It does not distinguish the tool from its sibling update_agent_test_folder_route, create_agent_test_folder_route, or delete_agent_test_folder_route beyond the obvious 'get' verb.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as listing folders or updating them. The description offers no context about prerequisites or typical use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_topics_routeC
Read-onlyIdempotent

Get Agent Conversation Topics

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
sort_byNoColumn to rank topics by. Use conversations for volume, sentiment with sort_direction=asc for the most negative topics, and frustration with sort_direction=desc for the most frustrated ones. Topics with no score are always ranked last.
agent_idYesID of the agent
page_sizeNoNumber of top-level topic groups to return.
to_unix_secsNoEnd of the window to view topics for.
from_unix_secsNoStart of the window to view topics for. When set with to_unix_secs, the completed daily topic-discovery runs in the range are aggregated together, so the window scopes the metrics as well as the topic set. Floored to the start of its UTC day because runs cover whole UTC days; aggregated_run_count re
sort_directionNoDirection to sort topics.
include_evaluation_criteriaNoInclude the per-criteria evaluation breakdown on each topic's metrics. Pass false to drop it: it dominates the payload and the weighted success_rate is returned either way.

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond them, such as that results depend on completed daily topic-discovery runs, that windows snap to UTC day boundaries, or that the criteria breakdown can dominate payload size.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At four words it is terse rather than concise: for an 8-parameter tool with non-obvious windowing and sorting behavior, this is under-specification, not efficient front-loading. There is nothing to front-load because no substantive content is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool takes 8 parameters, paginates via cursor, and aggregates across UTC-day topic-discovery runs, yet the description explains none of this. There is no output schema, so return-value context also falls to the description, and none is provided; only the unusually detailed input schema compensates.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema itself is unusually rich on the non-trivial parameters (sort_by directions, window aggregation behavior, include_evaluation_criteria payload impact). The description contributes no parameter semantics of its own, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a verb (Get) and a resource (Agent Conversation Topics), so it is not a pure tautology of the name, and the added word 'Conversation' narrows the kind of topics. However, it never explains what a 'topic' actually is, what the topic set represents, or that this returns ranked/grouped topic analytics, so an agent gets only a vague sense of purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no exclusions, and no mention of how this differs from adjacent tools like get_agent_summaries_route, run_conversation_analysis, or get_agent_route. The agent must infer the appropriate context entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_widget_routeC
Read-onlyIdempotent

Get Agent Widget Config

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.
conversation_signatureNoAn expiring token that enables a websocket conversation to start. These can be generated for an agent using the /v1/convai/conversation/get_signed_url endpoint

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive, and openWorld, so safety is covered externally. The description adds no behavioral context such as auth requirements, return shape, or why conversation_signature might be needed, even for a read operation that may involve an expiring token.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is only four words, but the problem is under-specification rather than efficient conciseness. The single sentence does not earn its place because it conveys almost nothing beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with full schema descriptions and strong annotations, the description still needs to distinguish this from sibling get_agent_* routes. It fails to say what the widget config is or when this route is preferable, leaving an agent unable to select confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both agent_id and conversation_signature are already documented in the input schema. The description adds no parameter meaning, but the baseline for full schema coverage is 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

Only restates the tool name with 'Config' swapped for 'route'; it does not explain what an Agent Widget Config contains or how it differs from the many get_agent_* siblings. A minimal verb+resource is present but adds no differentiating detail beyond the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisite information, and no named alternative such as get_agent_route. The description gives the agent nothing to decide whether this tool is the right one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_assetC
Read-onlyIdempotent

Get Asset

ParametersJSON Schema
NameRequiredDescriptionDefault
asset_idYesID of the asset.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that — no note about what an "asset" is, whether lookup fails on unknown IDs, or what is returned, so with the low bar set by annotations it still contributes zero additional context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short but this is under-specification rather than conciseness. Two words that duplicate the tool name and title convey no information an agent could not already derive from the identifier.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-ID read tool with full annotation coverage and a complete schema, the structural fields carry most of the burden, but the description still leaves the concept of an "asset" and the failure/lookup behavior undefined, which is more relevant here than for many siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Single parameter with 100% schema description coverage, so the schema already documents asset_id fully. Baseline 3 applies; the description adds no format, source, or semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Asset" merely restates the tool name and title. It does imply retrieval of a single asset by ID (contrasted with list_assets/upload_asset/delete_asset_endpoint in the sibling set), but gives no verb detail, no scope, and no distinguishing information beyond the name itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no mention of alternatives (list_assets for enumeration, upload_asset for creation), and no prerequisites. The agent must infer usage entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_assignable_users_routeC
Read-onlyIdempotent

Get Agent Conversation Ticket Assignable Users

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is fully covered. The description adds nothing behavioral beyond that — no note on permissions, scope, or what 'assignable' means in practice, so it contributes no value over the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no wasted words and is front-loaded. But its brevity is under-specification rather than disciplined conciseness — it earns its place only minimally.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a required parameter, no output schema, and a non-obvious resource concept ('assignable users'), the description omits what is returned and how agent_id shapes the result. Annotations cover safety, but the semantic gaps remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required agent_id parameter, and the description does not explain what the agent_id identifies or how it scopes the returned users. With one undocumented input parameter, the description fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description expands the abbreviated name into a resource statement: it clarifies that 'assignable users' belong to an Agent Conversation Ticket domain, which the raw name does not convey. However, it stops at restating the name and gives no verb-level detail about what 'getting' them yields or how it differs from siblings like get_conversation_users_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus alternatives such as get_conversation_users_route or get_agent_conversation_ticket_route. The agent is left to infer the use case purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_audio_from_sampleC
Read-onlyIdempotent

Get Audio From Sample Returns audio/* bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
sample_idYesSample ID to be used, you can use GET https://api.elevenlabs.io/v1/voices/{voice_id} to list all the available samples for a voice.
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, openWorldHint=true, so the safety profile is covered. The description adds that it returns audio/* bytes and that output_path saves them, but doesn't disclose additional behavioral traits like whether it needs authorization, rate limits, or what happens if the sample doesn't exist. With annotations covering safety, the description should add more context, but it falls short.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise: one sentence that covers the return type and the output_path option. It's front-loaded with the primary action. However, it could be structured better with a clear separation between what it does and how to use output_path.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 3 parameters (2 required), no output schema, and annotations that cover safety. The description mentions the return format and output_path, but doesn't explain the relationship between voice_id and sample_id, nor does it provide any guidance on error conditions or authentication. Given the complexity of the domain and the many siblings, more context would be helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters in detail. The description mentions output_path and implies inline base64 when omitted, which slightly reinforces but doesn't add new meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: 'Get Audio From Sample Returns audio/* bytes'. It's clear what the tool does. However, it doesn't differentiate from siblings like get_pvc_sample_audio or get_speaker_audio, which might also return audio from samples. The name itself is somewhat awkward, but the description clarifies the return type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit guidance on when to use this tool versus alternatives. With many sibling tools related to audio retrieval (e.g., get_pvc_sample_audio, get_speaker_audio), an agent needs to know the specific context. The description only mentions output_path behavior, not usage scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_audio_full_from_speech_history_itemB
Read-onlyIdempotent

Get Audio From History Item Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
history_item_idYesHistory item ID to be used, you can use GET https://api.elevenlabs.io/v1/history to receive a list of history items and their IDs.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds the concrete return format (audio/mpeg bytes) and the inline-vs-saved behavior, which is useful context not present in the annotations, but it omits any mention of size limits, auth needs, or pagination. Adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the purpose and then the output mechanism. It is efficient, though the run-on phrasing ('Get Audio From History Item Returns audio/mpeg bytes') is slightly grammatically awkward.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of describing the return value, and it does so by naming the MIME type (audio/mpeg) and the inline-vs-file behavior. For a simple two-parameter binary fetch this is largely sufficient, though the relationship to sibling download tools remains unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (history_item_id and output_path) are already fully documented in the schema, including the ELEVENLABS_OUTPUT_DIR resolution and base64 fallback. The description repeats output_path's role without adding syntax beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (Get) and resource (audio from a speech history item) and states the return type (audio/mpeg bytes). It does not, however, distinguish itself from closely related siblings like download_speech_history_items or get_speech_history_item_by_id, leaving the agent to infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Only implied usage is present: it notes that passing output_path saves the bytes, which hints at a write-to-disk workflow. There is no statement of when to use this over alternatives like download_speech_history_items, and no prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_audio_isolation_historyB
Read-onlyIdempotent

Get Audio Isolation History

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNoPage number for search pagination (1-based). Only used when search is provided.
searchNoOptional search term used for filtering audio isolation history (title/text).
page_sizeNoHow many history items to return at maximum. Defaults to 100.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe, idempotent read operation. The description adds no behavioral context beyond the name, such as pagination behavior or return format, but with annotations covering the safety profile, a baseline 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient phrase that directly states the tool's function. It is appropriately sized and front-loaded, but it lacks structure beyond the title-like statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool that is read-only and idempotent, the description is minimally adequate. However, without an output schema, it should ideally explain what the history items contain or how pagination works, which it omits. It is complete enough to call the tool but not to fully understand its output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, fully documenting the optional 'page', 'search', and 'page_size' parameters. The description adds no additional parameter meaning, but since the schema is self-sufficient and the description does not need to compensate, the baseline 4 for 0 required parameters with complete schema documentation is warranted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (get) and resource (audio isolation history), making the tool's purpose clear. It is significantly better than its tautological title since it names the actual resource. However, it does not differentiate this list tool from the sibling 'delete_audio_isolation_history_item' or other history-related tools, leaving some ambiguity about scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no explicit guidance on when to use this tool versus alternatives like 'get_speech_history' or 'get_audio_isolation_history_item'. It provides no context about prerequisites, what happens if there is no history, or how to retrieve a specific item instead of the list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_audio_native_project_settings_endpointC
Read-onlyIdempotent

Get Audio Native Project Settings

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesThe ID of the Studio project.

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered by structured data. The description adds zero behavioral context — no mention of auth requirements, what 'settings' contains, or what happens for a missing project_id. With annotations carrying the load, a low-but-not-floor 2 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no wasted words, but it is under-specified rather than genuinely concise — there is nothing to front-load because no information exists beyond the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool whose annotations and schema cover safety and parameters, the description still omits what settings are returned. With no output schema present, the description should at least hint at the return shape; it provides nothing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (project_id is documented as 'The ID of the Studio project'), so the single parameter is fully described by the schema. Per the baseline rule, 3 applies when the schema does the heavy lifting and the description adds nothing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is essentially a restatement of the tool name — 'Get Audio Native Project Settings' for get_audio_native_project_settings_endpoint. It names a verb and resource but adds no scope, no distinguishing information from siblings like create_audio_native_project or audio_native_project_update_content_endpoint, functioning as a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance whatsoever about when to call this tool versus the many sibling get_* and audio_native_* endpoints. No preconditions, no alternatives, no context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_batch_callC
Read-onlyIdempotent

Get A Batch Call By Id.

ParametersJSON Schema
NameRequiredDescriptionDefault
batch_idYes

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false. The description adds no behavioral context beyond the obvious read operation, such as authorization requirements, rate limits, or return behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. However, it is sparse rather than appropriately complete for a tool definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter GET tool with rich annotations, the description is minimally sufficient to invoke the tool. It still omits routing guidance relative to batch-call siblings and gives no indication of expected return content.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the only parameter, batch_id, is not described in the schema. The description says 'By Id', which minimally signals that the parameter is an identifier, but it does not explain the ID's format, source, or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get A Batch Call By Id' essentially restates the tool name and annotation title, adding only 'By Id'. It does not distinguish this tool from siblings such as get_workspace_batch_calls or explain the resource scope beyond the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, when not to use it, or which sibling tools handle related operations. An agent must infer usage solely from the tool name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_branches_routeC
Read-onlyIdempotent

List Agent Branches

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoHow many results at most should be returned
agent_idYesThe id of an agent. This is returned on agent creation.
include_archivedNoWhether archived branches should be included
include_commit_statusNoWhether to compute how far each branch has diverged from main (commits_ahead/commits_behind). This walks the version DAG of every branch, so it is slow on agents with long histories and is off by default, leaving those fields null.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, covering the safety profile. The description adds no behavioral context beyond the verb 'List'—it omits that archived branches are excluded by default, that include_commit_status is slow on long histories, and any pagination or rate-limit details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase with no wasted words, which is efficient. However, it is arguably too terse to be structurally helpful, reading more like a title than a definition, and lacks the front-loaded detail that would guide invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list tool with four parameters—including one (include_commit_status) whose slow behavior and default-null result are non-obvious—the description is incomplete. With no output schema, the definition should explain return behavior, defaults, and trade-offs, but it provides none of that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all four parameters. The description adds no meaning beyond the parameter names; it does not clarify the impact of include_commit_status or the default for limit and include_archived. Baseline 3 is appropriate when the schema carries the full load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Agent Branches' states a clear verb and resource, making the tool's basic purpose immediately understandable. However, it does not differentiate from sibling tools like get_branch_route (single branch) or other branch operations, so it lacks the specificity needed for a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get_branch_route or the various branch mutation tools. The description provides no context, prerequisites, or exclusions, leaving the agent to infer usage solely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_branch_routeC
Read-onlyIdempotent

Get Agent Branch

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.
branch_idYesUnique identifier for the branch.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false. The description adds no behavioral context beyond those annotations, such as what the returned branch data contains, whether it can fail for missing branches, or any authorization requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but it is under-specified rather than usefully concise. It does not front-load any actionable information beyond the bare verb-noun phrase.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two required ID parameters and no output schema, the description is nearly empty. It leaves an agent unable to tell what a branch represents, how it differs from other branch-related tools, or what invoking it returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both required parameters (agent_id and branch_id) are documented in the schema. The description adds no additional meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Agent Branch' essentially restates the tool name and title with only the addition of 'Agent'. It gives no detail about what a branch is or what is returned, and it does not distinguish this singular get from sibling tools like get_branches_route or get_agent_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives such as get_branches_route, create_branch_route, or update_branch_route. The description provides no context, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_chapter_by_id_endpointC
Read-onlyIdempotent

Get Chapter

ParametersJSON Schema
NameRequiredDescriptionDefault
chapter_idYesThe ID of the chapter.
project_idYesThe ID of the Studio project.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is fully covered without the description. The description adds nothing beyond that — no note on auth requirements, missing-ID behavior, or what happens for unknown project/chapter combinations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words are not conciseness but under-specification; there is no front-loaded scope, no constraint, and no second sentence to earn its place. It is too small to convey anything an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with two required parameters, no output schema, and no annotation gaps, the description should at minimum identify the resource being fetched and its relationship to get_chapters. Nothing beyond the bare name is provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both required parameters (project_id, chapter_id) individually documented, so the schema carries the parameter burden. The description contributes no additional meaning, which is the expected baseline when the schema is complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Chapter" merely restates the tool name/title (get_chapter_by_id_endpoint / "Get Chapter By Id Endpoint") without adding scope. It does not distinguish this tool from close siblings like get_chapters, get_chapter_snapshot_endpoint, or get_chapter_snapshots, leaving the agent unable to tell which retrieval variant applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative guidance is given. With siblings such as get_chapters and get_chapter_snapshot_endpoint in the same namespace, the absence of any routing hint forces the agent to guess from names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_chaptersC
Read-onlyIdempotent

List Chapters

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesThe ID of the Studio project.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, and openWorldHint=true, covering the safety profile. The description adds no further behavioral context such as scoping, return format, or pagination behavior, so it provides no value beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

"List Chapters" is extremely short and front-loaded, but it is under-specified rather than optimally concise. For a tool with many chapter-related siblings, this brevity leaves critical context out.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and its annotations and schema cover safety and the required parameter, but the description does not state that chapters are listed for a specific project. It also fails to clarify its relationship to the other chapter endpoints, leaving an agent with insufficient context for reliable selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single project_id parameter is fully documented in the schema itself. The description adds no parameter meaning, but with high schema coverage the baseline score is 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"List Chapters" essentially restates the tool name get_chapters and title Get Chapters, making it tautological. It gives no project scope or sibling differentiation, so an agent cannot distinguish it from get_chapter_by_id_endpoint or get_chapter_snapshots using the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of alternatives, and no indication of when not to use this tool. The agent is left to infer usage from the schema and name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_chapter_snapshot_endpointC
Read-onlyIdempotent

Get Chapter Snapshot

ParametersJSON Schema
NameRequiredDescriptionDefault
chapter_idYesThe ID of the chapter.
project_idYesThe ID of the Studio project.
chapter_snapshot_idYesThe ID of the chapter snapshot.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, covering the safety profile. The description contributes nothing further — no note on what a snapshot contains, whether it is immutable, or what happens if the ID is stale. With annotations this is not a contradiction, but it is a missed opportunity to add value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is only four words with no padding, so there is nothing to trim, but it is also so minimal that 'conciseness' is achieved by omission rather than by efficient information density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter read tool with no output schema, the description should at least indicate what the snapshot returns and how it relates to chapter/project hierarchy. Nothing about return shape, snapshot semantics, or sibling routing is provided, leaving a clear gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three required parameters (project_id, chapter_id, chapter_snapshot_id) are documented in the schema itself. Baseline 3 applies because the description adds no syntax, format, or relationship information beyond the structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is effectively the tool name restated: 'Get Chapter Snapshot' adds no verb nuance, scope, or resource detail beyond the identifier. It does not distinguish this from siblings like get_chapter_snapshots (list) or get_project_snapshot_endpoint (project-level analog).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, no mention of alternatives. An agent cannot tell from the description whether to call this for metadata, audio, or a downloadable archive, nor when get_chapter_snapshots should be preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_chapter_snapshotsC
Read-onlyIdempotent

List Chapter Snapshots

ParametersJSON Schema
NameRequiredDescriptionDefault
chapter_idYesThe ID of the chapter.
project_idYesThe ID of the Studio project.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description contributes nothing further — no mention of ordering, pagination, or what a snapshot represents relative to a project snapshot.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no waste and no verbosity, but it is under-specified rather than tight — there is simply nothing beyond the name, so the brevity reflects missing content rather than disciplined writing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations gap, the description still needs to explain what a chapter snapshot is and how the listing behaves (ordering, pagination, or relationship to project snapshots). An agent cannot tell from this text what it will receive or how to apply the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters (project_id, chapter_id) are documented at 100% schema coverage, so the schema carries the parameter semantics. The description adds no additional meaning, which is the baseline-3 case for a fully covered schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is essentially a restatement of the tool name and title: 'list' + 'Chapter Snapshots'. It does not distinguish this from close siblings such as get_chapter_snapshot_endpoint (singular) or get_project_snapshots, nor indicate what a snapshot contains or how the list is scoped.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool instead of get_chapter_snapshot_endpoint or get_chapter_by_id_endpoint. No prerequisites, no exclusions, no alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_conversation_audio_routeB
Read-onlyIdempotent

Get Conversation Audio Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
conversation_idYesThe id of the conversation you're taking the action on.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered without the description. The description's real addition is the return MIME type (audio/mpeg), which matters because there is no output schema, plus the save-vs-inline behavior; however it says nothing about size limits, permissions, or persistence. Useful but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short and front-loads the core action before the output_path note. It is slightly marred by a run-on construction ('Get Conversation Audio Returns audio/mpeg bytes') that reads as two merged sentences, costing a point on structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only two fully-described parameters, no output schema, and annotations covering the safety profile, the description supplies the one thing structured fields do not: the returned content type. What remains missing is return-size or error behavior, which keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already fully documented in the schema, including the ELEVENLABS_OUTPUT_DIR resolution and base64-inline fallback. The description's 'pass output_path to save them' repeats the schema's meaning without adding new semantics; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('Conversation Audio') and adds the return type ('audio/mpeg bytes'), so the agent knows what it retrieves and in what form. It does not differentiate itself from adjacent sibling tools like get_conversation_signed_link or get_audio_full_from_speech_history_item, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites or auth, and no pointer to alternative retrieval tools. The only usage hint is the mechanical 'pass output_path to save them', which is invocation detail rather than selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_conversation_histories_routeC
Read-onlyIdempotent

Get Conversations

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
searchNoFull-text or fuzzy search over transcript messages
tag_idsNoFilter conversations by conversation tag IDs assigned via the conversation-tags endpoints.
user_idNoFilter conversations by the user ID who initiated them.
agent_idNoAgent id (agent_…) or speech engine external id (seng_), resolved to the same underlying resource.
branch_idNoFilter conversations by branch ID.
page_sizeNoHow many conversations to return at maximum. Can not exceed 100, defaults to 30.
text_onlyNo
topic_idsNoFilter conversations by topic IDs assigned during topic discovery.
rating_maxNoMaximum overall rating (1-5).
rating_minNoMinimum overall rating (1-5).
tool_namesNoFilter conversations by tool names used during the call.
version_idNoFilter conversations by version ID.
summary_modeNoWhether to include transcript summaries in the response.
main_languagesNoFilter conversations by detected main language (language code).
sort_directionNoThe direction to sort conversations by call start time. Defaults to descending (newest first).
call_successfulNoThe result of the success evaluation
guardrail_typesNoFilter to conversations where a guardrail of any of these types triggered (metadata.triggered_guardrails.guardrail_type). Repeat param to match any of several.
exclude_statusesNoExclude conversations with the given statuses. Useful for hiding in-progress / processing conversations from list views.
evaluation_paramsNoEvaluation filters. Repeat param. Format: criteria_id:result. Example: eval=value_framing:success
visited_agent_idsNoFilter conversations where any of these agents participated. Can not exceed 50 values.
tool_names_erroredNoFilter conversations by tool names that had errored calls.
data_collection_idsNoData collection field IDs to include in each conversation summary. Repeat param. When omitted, data_collection_results is not returned.
termination_reasonsNoFilter conversations by their stored termination_reason (metadata.termination_reason). Repeat param to match any of several.
has_feedback_commentNoFilter conversations with user feedback comments.
call_start_after_unixNoUnix timestamp (in seconds) to filter conversations after to this start date.
tool_names_successfulNoFilter conversations by tool names that had successful calls.
call_duration_max_secsNoMaximum call duration in seconds.
call_duration_min_secsNoMinimum call duration in seconds.
call_start_before_unixNoUnix timestamp (in seconds) to filter conversations up to this start date.
custom_guardrail_namesNoFilter to conversations where a custom guardrail with any of these names triggered (metadata.triggered_guardrails.guardrail_name). Only custom guardrails carry a name. Repeat param to match any of several.
data_collection_paramsNoData collection filters. Repeat param. Format: id:op:value where op is one of eq|gt|gte|lt|lte|missing.
parent_conversation_idNoFilter conversations by parent conversation ID for subagent conversations.
dynamic_variable_paramsNoDynamic variable filters. Repeat param. Format: name:op:value where op is one of eq|gt|gte|lt|lte. Comparison operators require a numeric value. Names containing ':' cannot be expressed.
evaluation_criteria_idsNoEvaluation criteria IDs to include in each conversation summary. Repeat param. When omitted, evaluation_criteria_results is not returned.
triggered_procedure_idsNoFilter conversations where any of these procedures were triggered. Can not exceed 50 values.
visited_agent_branch_idsNoFilter conversations where any of these agent branches participated. Can not exceed 50 values.
workflow_node_entered_idNoFilter conversations to only those that entered the given node.
conversation_product_typeNoRestrict results to a single conversation product surface.
include_invalid_tool_callsNoAlso match tool calls that never ran.
conversation_initiation_sourceNo

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is covered by structured data. The description adds nothing on top: not that results are cursor-paginated, not that summary_mode/data_collection_ids change payload shape, not the <=100 page cap behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is not conciseness but under-specification. For a 41-parameter listing endpoint, the description is sized far below what the task requires; nothing is front-loaded because nothing is said.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is one of the most complex list endpoints in the set (41 optional filters, pagination, response-shape toggles) with no output schema to fall back on. A two-word description leaves the agent with no notion of filtering semantics, pagination, or result shape beyond what it must reconstruct from the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 95%, so the schema itself documents nearly all 41 parameters including formats like criteria_id:result and id:op:value. With coverage this high the baseline of 3 applies; the description contributes no additional parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Conversations" is essentially a restatement of the tool name and title, with no verb nuance, no scope statement, and no differentiation from near-identical siblings like get_conversation_history_route, get_conversation_summary_route, or list_workspace_conversation_tickets_route. It conveys only the bare resource, and only because the plural noun appears in the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this paginated/filterable list route versus the singular get_conversation_history_route or the message-search routes (smart_search_conversation_messages_route, text_search_conversation_messages_route). No prerequisites, no exclusions, no mention that results are paginated via cursor.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_conversation_history_routeC
Read-onlyIdempotent

Get Conversation Details

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoResponse format. Defaults to 'json'. Set to 'opentelemetry' for an OTLP-compatible trace payload using the same structure as the post-call webhook.
conversation_idYesThe id of the conversation you're taking the action on.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true. The description adds no behavioral context beyond those annotations, such as pagination, response shape, or scoping, so it does not earn credit for supplementing structured data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At three words, it is technically short but under-specified rather than concise. It lacks front-loaded scope or useful structure for an agent deciding whether to call it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a retrieval tool with no output schema, the description should clarify what is returned and how it differs from sibling conversation endpoints. It omits that entirely, leaving only the annotations and schema to carry the definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters have detailed schema descriptions, including the format enum and conversation_id. With the schema doing the semantic work, the baseline of 3 applies; the description adds no further parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Conversation Details' gives a verb and broad resource but does not specify what kind of details or distinguish the tool from many siblings such as get_conversation_histories_route, get_conversation_summary_route, or get_conversation_audio_route. It is essentially a restatement of a generic conversation lookup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites, and no exclusions. The description provides no context that would help an agent choose among sibling conversation retrieval endpoints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_conversation_sip_messagesC
Read-onlyIdempotent

Get Sip Messages For A Conversation

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
page_sizeNo
conversation_idYesThe id of the conversation you're taking the action on.

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true. The description adds no behavioral context beyond the annotations, such as whether results are paginated, ordered, or limited to a specific message type.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. It is appropriately concise, though it is arguably too minimal for a tool with three parameters and multiple sibling getters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the large sibling set and the presence of list_sip_messages, the description is incomplete: it does not clarify the selection boundary, return scope, or pagination behavior. Annotations cover safety, but the description leaves key invocation context missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 67%: conversation_id and cursor have schema descriptions, but page_size has none. The description adds no parameter meaning and does not compensate for the undocumented page_size parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get') and resource ('Sip Messages For A Conversation'), making the basic operation clear. However, it does not distinguish this tool from the sibling list_sip_messages or explain the scoping behavior that differentiates the two.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as list_sip_messages, nor does it mention prerequisites or the pagination-oriented context implied by the cursor parameter. Usage is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_conversation_summary_routeC
Read-onlyIdempotent

Get Conversation Summary

ParametersJSON Schema
NameRequiredDescriptionDefault
max_messagesNoMaximum number of chat message turns to include inline. When the conversation has more than this, the messages are omitted and messages_omitted is set.
conversation_idYesThe id of the conversation you're taking the action on.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that: it does not say whether summaries are pre-generated or computed on demand, whether there is a cost/latency implication, or how the summary relates to the underlying messages.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single four-word phrase contains no waste, but it is under-specification rather than genuine conciseness. It is front-loaded only in the trivial sense of being the entire description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and only a title-level phrase, an agent cannot tell what the summary contains, its format, or how max_messages truncation surfaces. For a read tool with two parameters the safety annotations cover part of the burden, but the description is still too thin to be complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so conversation_id and max_messages (including the messages_omitted behavior) are already fully documented in the schema. Baseline 3 applies because the description contributes no additional parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name ('get_conversation_summary_route' -> 'Get Conversation Summary') and adds no distinguishing detail. It does not differentiate this tool from close siblings such as get_conversation_history_route, get_conversation_histories_route, or run_conversation_analysis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many conversation-related siblings, nor any prerequisites or exclusions. The agent is left to infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_conversation_tag_routeC
Read-onlyIdempotent

Get Conversation Tag

ParametersJSON Schema
NameRequiredDescriptionDefault
tag_idYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already state readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond the title, such as whether it requires specific permissions, what happens for an unknown tag_id, or whether missing IDs return null or an error. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only three words. While short, it is under-specified rather than genuinely concise: an agent gets no useful information beyond the name and title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter read tool with rich annotations, the description is still incomplete because it omits the meaning of the required tag_id and any sense of what the call returns. It relies entirely on annotations and the schema name for even basic semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one required parameter, tag_id, with 0% description coverage. The description does not explain whether tag_id is a UUID, slug, or internal identifier, nor does it compensate for the missing schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Conversation Tag' restates the tool name and title rather than specifying what is retrieved or how it differs from siblings like list_conversation_tags_route, update_conversation_tag_route, or delete_conversation_tag_route. It names the resource but adds no distinguishing scope or return semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. The sibling set contains multiple conversation-tag operations, yet the description gives no conditions, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_conversation_users_routeC
Read-onlyIdempotent

Get Conversation Users

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
searchNoSearch/filter by user ID (exact match).
sort_byNoThe field to sort the results by. Defaults to last_contact_unix_secs.
agent_idNoAgent id (agent_…) or speech engine external id (seng_), resolved to the same underlying resource.
branch_idNoFilter conversations by branch ID.
page_sizeNoHow many users to return at maximum. Defaults to 30.
sort_directionNoThe direction to sort the results
call_start_after_unixNoUnix timestamp (in seconds) to filter conversations after to this start date.
call_start_before_unixNoUnix timestamp (in seconds) to filter conversations up to this start date.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, and the description adds no behavioral context beyond them. It does not mention pagination, default limits, filtering behavior, or what 'conversation users' means.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only four words and is under-specified rather than effectively concise. For a tool with nine optional filtering and pagination parameters, it does not front-load enough information to guide invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although annotations cover safety and the schema covers parameters, the description omits critical context for a 9-parameter conversation-user listing tool: scope, intended use cases, and differentiation from siblings. No output schema exists to compensate for the missing return-value context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all nine parameters are documented in the schema. The description adds no parameter meaning, but the baseline is 3 when the schema carries the full parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Conversation Users' is a tautological restatement of the tool name/title, adding no scope, actor, or distinction from siblings such as get_assignable_users_route or get_workspace_members.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives. With many user-listing and conversation-related siblings, the absence of routing context leaves selection ambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_crawl_job_routeC
Read-onlyIdempotent

Get Crawl Job Details

ParametersJSON Schema
NameRequiredDescriptionDefault
crawl_job_idYesThe id of the crawl job to retrieve

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered by structured data. The description adds nothing beyond that - no note on error behavior for a missing job id, no indication of what 'details' includes, no auth or rate-limit context. It neither contradicts nor enriches the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded phrase with zero redundancy or filler. It is concise to a fault - nothing extraneous, though also nothing substantive beyond the heading.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (one required parameter, read-only, fully annotated), so the description carries a light burden and the annotations plus schema cover safety and input semantics. However, with no output schema, an agent gets no hint of what 'details' contains (status? config? errors?), and there is no differentiation from the list/cancel siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% - crawl_job_id is documented in the schema as 'The id of the crawl job to retrieve'. With only one well-documented parameter, the baseline of 3 applies; the description contributes no additional meaning about id format or provenance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Crawl Job Details' is essentially the tool name/title restated with the words reordered; it names no scope, return shape, or distinguishing feature. It does not differentiate this tool from siblings such as list_crawl_jobs_route (listing) or cancel_crawl_job_route (cancellation). An agent can guess it fetches one job, but only from the identifier in the name, not from the description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives. With list_crawl_jobs_route and cancel_crawl_job_route in the sibling set, the description should at minimum state that this retrieves a single job by id rather than enumerating or cancelling jobs. Nothing is offered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dashboard_settings_routeB
Read-onlyIdempotent

Get Convai Dashboard Settings

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds no behavioral context beyond restating that it is a read operation, such as auth requirements, rate limits, or return characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short, front-loaded phrase with no wasted words. It is appropriately concise for a zero-parameter read tool, though it could benefit from slightly more specificity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only, zero-parameter tool with rich annotations, the description is minimally adequate. It omits what dashboard settings include and how it relates to the similarly named get_settings_route and update_dashboard_settings_route tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document. Schema coverage is 100% and the empty object schema is self-explanatory.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Get') and resource ('Convai Dashboard Settings'), so an agent can tell it retrieves dashboard settings. However, it does not distinguish itself from sibling tools such as get_settings_route or update_dashboard_settings_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. It does not mention prerequisites, exclusions, or related sibling tools like get_settings_route or update_dashboard_settings_route.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_documentation_chunk_from_knowledge_baseC
Read-onlyIdempotent

Get Documentation Chunk From Knowledge Base

ParametersJSON Schema
NameRequiredDescriptionDefault
chunk_idYesThe id of a document RAG chunk from the knowledge base.
embedding_modelNoThe embedding model used to retrieve the chunk.
documentation_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered structurally. The description adds nothing on top of that — no note on retrieval scope, missing-chunk behavior, or how the embedding_model affects results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that merely echoes the title; it is short but does not earn its place since it carries no information beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with 100% schema coverage and full annotations, the description could still clarify the chunk-vs-document distinction and when to prefer this over sibling retrieval tools. That context is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with documentation_id, chunk_id, and embedding_model all documented in the schema, so the baseline is 3. The description contributes no additional meaning about how the required IDs relate to one another.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name ('Get Documentation Chunk From Knowledge Base') and adds no distinguishing detail. It does not separate this single-chunk retrieval from siblings like get_documentation_chunks_from_knowledge_base or get_documentation_from_knowledge_base.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the near-identical sibling tools. An agent cannot tell from the text why it would call this over the plural chunks route.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_documentation_chunks_from_knowledge_baseC
Read-onlyIdempotent

Get All Rag Chunks For A Document

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
page_sizeNoHow many documents to return at maximum. Can not exceed 100, defaults to 30.
embedding_modelYesThe embedding model used to retrieve the chunk.
documentation_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds little beyond implying enumeration of all chunks, and it does not mention pagination behavior or any rate limits, so it only marginally exceeds the annotation baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase with no padding, so it is concise, but the phrase is under-specified rather than front-loaded with useful information. It is not wasteful, just minimal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With four parameters, pagination semantics and a large set of sibling chunk/knowledge-base tools, the description should clarify scope and relationship to alternatives. It does not, and with no output schema it also says nothing about the return shape or pagination, leaving important context missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents cursor, page_size, embedding_model and documentation_id in detail. The description adds no parameter meaning, which is the baseline 3 when the schema does all the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb and resource (get chunks for a document), which is clearer than the name alone, but it does not distinguish this tool from close siblings like get_documentation_chunk_from_knowledge_base (singular) or query_agent_knowledge_base_rag_route. An agent cannot tell from the description alone which chunk-retrieval tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given and no alternatives are named. The description does not explain that this is for enumerating all chunks of a document versus performing a semantic query, nor does it mention prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_documentation_from_knowledge_baseC
Read-onlyIdempotent

Get Documentation From Knowledge Base

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idNo
documentation_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered by structured data. The description contributes nothing beyond that — no note on what the returned document contains, whether the id must come from add_documentation_to_knowledge_base, or any failure behavior for a stale id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short line, so there is no waste — but there is also no content. This is under-specification rather than conciseness, since the one sentence merely echoes the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with no output schema and only 50% parameter coverage, the description should carry real weight: what a 'documentation' object is, the role of agent_id, and how this differs from the chunk-level siblings. None of that is present, so an agent has almost nothing to act on beyond the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: documentation_id is documented as the id returned on document addition, but agent_id has no description anywhere. The prose adds no parameter meaning at all, so the undocumented half stays undocumented and the description fails to compensate for the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name and annotation title ('Get Documentation From Knowledge Base'), which is the definition of a tautology. It does not distinguish this tool from near-identical siblings like get_documentation_chunk_from_knowledge_base or get_documentation_chunks_from_knowledge_base, leaving an agent to guess which granularity to use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this tool versus retrieving chunks, listing the knowledge base, or searching its content. The only implied usage is the trivial 'retrieve one document', with no prerequisites, no exclusions, and no named alternatives among the dozens of knowledge-base siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dubbed_fileB
Read-onlyIdempotent

Get Dubbed File Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
dubbing_idYesID of the dubbing project.
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
language_codeYesID of the language.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds the useful behavioral detail that the tool returns audio/mpeg bytes, but it omits auth requirements, large-file/base64 caveats, and pagination or streaming behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short and front-loaded, with no wasted padding. The opening sentence runs 'Get Dubbed File Returns audio/mpeg bytes' together without punctuation, which is slightly awkward but not enough to impede use.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must name the return type, which it does (audio/mpeg bytes), and the schema fully documents the inputs. For a simple read/retrieve tool whose safety profile is covered by annotations, this is nearly complete, with only usage and auth context missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so dubbing_id, language_code, and output_path are all documented in the schema. The description only echoes output_path's save behavior, adding no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (Dubbed File) plus the return media type (audio/mpeg bytes), so the agent knows it fetches the rendered audio. It does not, however, distinguish itself from close siblings like get_dubbed_transcript_file, get_dubbed_metadata, or get_dubbing_resource, which all sound like plausible ways to 'get' dubbing output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says nothing about when to use this tool versus the many dubbing siblings (metadata, transcript file, dubbing resource). There are no prerequisites, no exclusions, and no routing guidance, leaving the agent to infer the choice from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dubbed_metadataC
Read-onlyIdempotent

Get Dubbing

ParametersJSON Schema
NameRequiredDescriptionDefault
dubbing_idYesID of the dubbing project.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond those annotations—it does not say what metadata is returned, whether authentication is required, or any rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short but under-specified rather than concise. It lacks the necessary detail to make the tool usable and leaves critical information missing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a simple read-only tool with one documented parameter and annotations covering safety, the description is still incomplete: it does not explain what 'metadata' includes or how this differs from similar dubbing retrieval tools. With no output schema, the description should describe the return value but does not.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter dubbing_id is fully documented in the schema as 'ID of the dubbing project.' The description adds no parameter meaning, but the baseline for high schema coverage with no additional param info is 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Dubbing' essentially restates part of the tool name without specifying that it retrieves metadata, and it fails to distinguish this tool from close siblings like get_dubbing_resource, get_dubbed_file, or get_dubbed_transcript_file. The agent cannot tell what resource is actually returned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is provided, no alternatives are named, and no conditions for selecting this tool over the many other dubbing-related getters are given. The description gives no clue about appropriate context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dubbed_transcript_fileC
Read-onlyIdempotent

Get Dubbed Transcript Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
dubbing_idYesID of the dubbing project.
format_typeNoFormat to return transcript in. For subtitles use either 'srt' or 'webvtt', and for a full transcript use 'json'. The 'json' format is not yet supported for Dubbing Studio.
language_codeYesISO-693 language code to retrieve the transcript for. Use 'source' to fetch the transcript of the original media.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior, so the safety profile is covered. The description adds only the fragment 'Deprecated upstream', which vaguely signals lifecycle status but does not say whether the endpoint still returns data or what the deprecation means for the caller.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is short, but it is not concise so much as incomplete and poorly formed — a sentence fragment trailing after the resource name. There is no front-loaded statement of what the tool returns or how it should be invoked.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter retrieval tool with no output schema, the description should at least explain what is returned (a transcript file in the requested format) and the implications of deprecation. Neither is provided, so an agent cannot confidently decide to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: dubbing_id, language_code, and the format_type enum are all documented in the schema itself, including the 'source' convention and format meanings. The description adds no parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description essentially restates the tool name ('Get Dubbed Transcript') rather than stating a distinct purpose, and the trailing fragment 'Deprecated upstream' is ungrammatical and ambiguous. It does not distinguish this tool from close siblings such as get_dubbing_transcripts, dubbing_transcript_get, or get_dubbed_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, when not to, or which sibling to prefer. The only guidance-adjacent text, 'Deprecated upstream', does not tell the agent whether to avoid the tool or what to call instead, leaving it unusable for routing decisions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dubbing_resourceB
Read-onlyIdempotent

Get The Dubbing Resource For An Id. Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
dubbing_idYesID of the dubbing project.

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds one useful behavioral fact beyond annotations: that the tool is deprecated upstream. It does not cover auth requirements, rate limits, or return behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no wasted words. The core purpose is front-loaded and the deprecation warning, while terse, earns its place by flagging an important operational caveat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with rich annotations and full schema coverage, the description is minimally sufficient: an agent knows to call it with a dubbing ID. However, it does not clarify what kind of dubbing resource is returned, which matters given the many dubbing siblings and the absence of an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the sole parameter dubbing_id is already documented in the schema as "ID of the dubbing project." The description adds no parameter meaning beyond that, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb and resource ("Get The Dubbing Resource For An Id"), but "dubbing resource" is ambiguous among many dubbing-related siblings such as dubbing_project_get, dubbing_transcript_get, and get_dubbed_metadata. It does not distinguish this tool from those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage signal is "Deprecated upstream," which warns that the tool is outdated but does not name a replacement or state when to use this tool versus alternatives. No explicit when/when-not guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dubbing_transcriptsC
Read-onlyIdempotent

Retrieve A Transcript

ParametersJSON Schema
NameRequiredDescriptionDefault
dubbing_idYesID of the dubbing project.
format_typeYesFormat to return transcript in. For subtitles use either 'srt' or 'webvtt', and for a full transcript use 'json'. The 'json' format is not yet supported for Dubbing Studio.
language_codeYesISO-693 language code to retrieve the transcript for. Use 'source' to fetch the transcript of the original media.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond those annotations—no auth requirements, rate limits, output characteristics, or format-handling nuances.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The four-word description is not concise but under-specified; it restates the tool name without earning its place. Brevity here reflects missing information rather than efficient structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three required parameters and many sibling transcript tools, the description omits what makes this tool distinct and when to use it. It fails to carry the contextual burden despite rich annotations and schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema fully documents all three required parameters, including the format_type enum and language_code semantics. The description adds no parameter meaning, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Retrieve A Transcript" restates the tool name minus 'dubbing,' adding no specific resource or scope. Among dozens of transcript-related siblings (dubbing_transcript_get, get_transcript_by_id, dubbing_target_transcript_get, etc.), the description provides no differentiation, leaving an agent unable to tell which transcript tool to select.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance, no prerequisites, and no alternatives. With many sibling transcript tools, this absence of routing information is a clear gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_environment_variableC
Read-onlyIdempotent

Get Environment Variable

ParametersJSON Schema
NameRequiredDescriptionDefault
env_var_idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that—no mention of what is returned, error conditions, or scope. With annotations doing the heavy lifting, the description could still add value but doesn't.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase with no waste, but it's too terse to be useful. It is appropriately sized for what it says, but what it says is minimal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter with one required parameter and no output schema, the description is incomplete. It doesn't explain what the tool returns, whether it's by ID or name, or how it relates to other environment variable tools. The agent lacks sufficient context to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 1 parameter (env_var_id) with 0% description coverage. The description doesn't mention the parameter at all, so it fails to compensate for the lack of schema documentation. The agent must guess what env_var_id represents and its format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Environment Variable' states a verb ('Get') and resource ('Environment Variable'), which is clear enough. However, it simply restates the tool name and title verbatim, providing no additional differentiation from siblings like list_environment_variables. It's minimally viable but not informative beyond the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as list_environment_variables or update_environment_variable. The description offers no context about when retrieval is appropriate, leaving the agent to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_finetuneC
Read-onlyIdempotent

Get Music Finetune

ParametersJSON Schema
NameRequiredDescriptionDefault
finetune_idYes

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered by structured data. The description adds nothing beyond that — no mention of auth requirements, retrieval scope, or behavior on a missing/invalid finetune_id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no wasted words and is fully front-loaded, but the brevity reflects under-specification rather than tight, information-dense writing. Nothing beyond the bare minimum is conveyed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of explaining the return value, and it does not. For a retrieval tool with an openWorld hint and one undocumented parameter, an agent learns nothing about what it will get back or what a 'music finetune' is.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter finetune_id is undocumented in both schema and description. The description does not clarify the expected ID format (e.g., where a finetune_id comes from, as returned by get_finetunes or create_finetune).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Music Finetune" is essentially a restatement of the tool name with the word "Music" added. It names a verb and resource but gives no detail about what a finetune object contains or how this differs from the sibling get_finetunes (list) tool. This is close to tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this tool versus get_finetunes, create_finetune, update_finetune, or delete_finetune, all of which are close siblings. No prerequisites, no context of use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_finetunesC
Read-onlyIdempotent

Get Music Finetunes

ParametersJSON Schema
NameRequiredDescriptionDefault
sortNoSort by field (created_at or name)
cursorNoUsed for fetching the next page. Cursor is returned in the response.
page_sizeNoHow many finetunes to return. Max 150, default 50.
created_byNoFilter by creator. 'self' returns finetunes you created; 'workspace' returns finetunes created by workspace teammates; 'elevenlabs' returns ElevenLabs curated finetunes. Omit to return finetunes from all creators.
visibilityNoFilter by visibility. 'private' returns private finetunes; 'workspace' returns workspace-shared finetunes; 'public' returns public finetunes, which are currently ElevenLabs curated finetunes. Omit to return all accessible finetunes.
sort_directionNoSort direction (asc or desc)

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, covering the safety profile fully. The description adds nothing behavioral—no mention of pagination, result scope, or auth context—so it provides no value beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At three words it is maximally short and front-loaded, with no wasted sentences. However, the brevity stems from under-specification rather than well-chosen economy, so it does not earn a top score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a six-parameter, paginated listing tool with no output schema, and the description says nothing about it being a paginated list, what the cursor returns, or how results are scoped. Given the moderate complexity and the rich annotation set, the description leaves key context unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: all six parameters (sort, cursor, page_size, created_by, visibility, sort_direction) are thoroughly documented in the schema, including enum meanings. The description adds no parameter meaning beyond that, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Music Finetunes" essentially restates the tool name and title, adding only the word "Music." It does not distinguish this list tool from siblings like get_finetune (singular), create_finetune, update_finetune, or delete_finetune. An agent learns only that it retrieves something related to finetunes, with no scope or filtering detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus get_finetune, the update/delete/create finetune siblings, or any other listing tool. No prerequisites or exclusions are given, so the agent must infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_groups_endpointC
Read-onlyIdempotent

Get All Groups

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is fully carried by structured data. The description adds nothing behavioral (no pagination, no auth scope, no return shape), but with annotations covering the read/idempotent/foreign-world traits, a repository-covered baseline of 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is two words, so there is no structural bloat. It is front-loaded and free of filler, though the extreme brevity is under-specification rather than exemplary conciseness, which keeps it out of the top score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no annotations burden, but the description is so thin that an agent gets no sense of what a 'group' is, whether the list is scoped to a workspace, or how it relates to search_groups. For a listing tool in a large sibling set, it should do more.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which is the baseline-4 case for this dimension. Nothing about parameters needs explaining, so the description cannot be faulted here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get All Groups' is a tautology that largely restates the tool name get_groups_endpoint (which already encodes the get_groups action). It names no specific resource semantics beyond 'groups', and gives no hint of scope (which groups endpoint, what a group is) that would distinguish it from sibling listers like get_agents_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance is provided. The sibling search_groups is the obvious alternative for lookup by criteria, yet the description never mentions it or any condition choosing between listing all groups and searching them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_image_generationC
Read-onlyIdempotent

Get Image Generation

ParametersJSON Schema
NameRequiredDescriptionDefault
generation_idYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is fully covered externally. The description adds nothing on top of that: no statement about what a lookup returns, failure behavior for an unknown id, or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words is brief, but this is under-specification rather than conciseness: there is no front-loaded purpose statement that earns its place. The brevity costs the definition all useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, one undocumented parameter, and a tautological description, an agent cannot know what the response contains or how the id is obtained. Annotations cover the safety profile but the descriptive burden for retrieval semantics is entirely unmet.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required parameter generation_id, so the schema gives no meaning beyond the type. The description does not compensate at all, leaving the agent to guess that this is an identifier of an image generation resource.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Get Image Generation" merely restates the tool name and title, adding no verb-resource specificity beyond what the identifier already conveys. An agent learns only that it retrieves something called an image generation, with no detail on scope or what is returned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus siblings such as list_image_generations or create_image_generation, nor any prerequisites or exclusions. The sibling list makes the get/list/create distinction inferable, but the description itself provides nothing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_knowledge_base_bulk_dependent_agents_routeC

Get Dependent Agents For Multiple Documents

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
page_sizeNoHow many documents to return at maximum. Can not exceed 100, defaults to 30.
document_idsYesThe ids of documents or folders from the knowledge base.
dependent_typeNoType of dependent agents to return.

TDQS

C2.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description says 'Get', which implies a read-only retrieval, but the annotations declare readOnlyHint=false. This is a direct contradiction in the safety signal. The description also adds no behavioral context such as pagination, auth requirements, or what dependent_type returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is a single front-loaded sentence with no wasted words. Its brevity is appropriate structurally, though it contributes to gaps in other dimensions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite 100% schema coverage, the description omits how this bulk route differs from the singular dependent-agents tool and what 'dependent agents' actually returns. With no output schema and a contradictory read-only annotation, the definition is not complete enough for confident tool selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with clear descriptions for cursor, page_size, document_ids, and dependent_type. The description adds no parameter meaning beyond the schema, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and resource ('Dependent Agents') with a clear scope ('For Multiple Documents'). It implies the bulk variant, though it does not explicitly name the sibling singular tool get_knowledge_base_dependent_agents to sharpen differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this bulk tool versus the singular get_knowledge_base_dependent_agents or get_tool_dependent_agents_route. The scope 'For Multiple Documents' weakly implies usage but provides no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_knowledge_base_contentC
Read-onlyIdempotent

Get Document Content

ParametersJSON Schema
NameRequiredDescriptionDefault
documentation_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is fully covered structurally. The description adds nothing beyond that – no indication of what the returned content looks like, whether it is paginated, or any size/auth caveats.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no waste, which is fine structurally, but the brevity reflects under-specification rather than efficiency. There is nothing misleading, just very little substance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining what "content" is returned (raw document text, chunks, file URL?), and it does not. For a retrieval tool sitting next to several similar knowledge-base getters, this leaves an agent guessing about the return payload.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter documentation_id is well documented in the schema (its origin is even stated: "returned on document addition"). The description adds no meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Document Content" is essentially a restatement of the tool name and title (get_knowledge_base_content / "Get Knowledge Base Content"). It does not distinguish this tool from adjacent siblings such as get_documentation_from_knowledge_base, get_documentation_chunks_from_knowledge_base, or search_knowledge_base_content_route, so an agent cannot tell which one fetches raw content versus metadata or chunks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisite (e.g. that a documentation_id must first come from add_documentation_to_knowledge_base), and no mention of alternatives. The only hint at usage comes from the schema, not the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_knowledge_base_dependent_agentsC
Read-onlyIdempotent

Get Dependent Agents List

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
page_sizeNoHow many documents to return at maximum. Can not exceed 100, defaults to 30.
dependent_typeNoType of dependent agents to return.
documentation_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is fully covered structurally. The description adds nothing beyond that — no note on pagination behavior, no scope of 'dependent' (direct vs transitive vs all), and no hint about what a dependent agent even is.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is maximally short and front-loaded, but the brevity comes from under-specification rather than efficient packing. Four words that only echo the title are less useful than a sentence that earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter retrieval tool with no output schema and multiple near-duplicate siblings, the description does not explain what dependent agents are, what the returned list contains, or how this differs from the bulk variant. It is too thin to guide correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with four well-documented parameters (cursor, page_size, dependent_type enum, documentation_id), so the schema carries the semantics. The description contributes no additional parameter meaning, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Get Dependent Agents List" essentially restates the tool name (get_knowledge_base_dependent_agents) without adding a verb+resource specificity beyond it. It does not distinguish this tool from close siblings like get_knowledge_base_bulk_dependent_agents_route or get_tool_dependent_agents_route, so an agent cannot tell them apart from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. Given that sibling tools cover bulk dependent agents and tool dependent agents, the absence of any routing guidance leaves the agent to guess which retrieval is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_knowledge_base_list_routeD
Read-onlyIdempotent

Get Knowledge Base List

ParametersJSON Schema
NameRequiredDescriptionDefault
typesNoIf present, the endpoint will return only documents of the given types.
cursorNoUsed for fetching next page. Cursor is returned in the response.
searchNoIf specified, the endpoint returns only such knowledge base documents whose names start with this string.
sort_byNoThe field to sort the results by
page_sizeNoHow many documents to return at maximum. Can not exceed 100, defaults to 30.
folders_firstNoWhether folders should be returned first in the list of documents.
sort_directionNoThe direction to sort the results
parent_folder_idNoIf set, the endpoint will return only documents that are direct children of the given folder.
ancestor_folder_idNoIf set, the endpoint will return only documents that are descendants of the given folder.
created_by_user_idNoFilter documents by creator user ID. When set, only documents created by this user are returned. Takes precedence over show_only_owned_documents. Use '@me' to refer to the authenticated user.
show_only_owned_documentsNoIf set to true, the endpoint will return only documents owned by you (and not shared from somebody else). Deprecated: use created_by_user_id instead.

TDQS

D1.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds zero behavioral context—no mention of pagination, result ordering, rate limits, or what the response contains. It is effectively missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single terse phrase that is under-specified rather than concise. For an 11-parameter list endpoint with filtering, sorting, and pagination, this is far too minimal to be structurally adequate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (11 parameters, pagination, filtering) and no output schema, the description is completely inadequate. It omits any explanation of how to use the filters, how pagination works, or what the return shape looks like.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 11 parameters are fully documented in the schema itself. The description provides no additional meaning, which meets the baseline of 3 for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Knowledge Base List' is a tautology that simply restates the tool name and title without adding any distinguishing information. It does not clarify scope, filtering, or how it differs from other knowledge base list endpoints like get_agent_knowledge_base_summaries_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. It doesn't mention pagination, filtering, or any prerequisite context that would help an agent select it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_knowledge_base_source_file_urlC
Read-onlyIdempotent

Get Document Source File Url

ParametersJSON Schema
NameRequiredDescriptionDefault
documentation_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is fully covered by structured data. The description adds nothing beyond those hints — it does not mention whether the returned URL is signed, time-limited, or requires the document to be of a particular source type.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no wasted words and is trivially front-loaded, but the brevity comes from under-specification rather than disciplined editing. There is nothing to trim and equally nothing that earns its place beyond the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of explaining what is returned, yet it never states that the result is a URL, whether it expires, or how it relates to the knowledge-base document lifecycle. For a lookup tool in a dense family of retrieval siblings, this leaves an agent with too little to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single parameter documentation_id is well described in the schema itself as the document id returned on document addition. The description adds no parameter meaning, so the baseline of 3 applies when the schema does the explanatory work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Document Source File Url' is essentially a title-case restatement of the tool name, with no added specificity about scope, source type, or what distinguishes it from siblings like get_documentation_from_knowledge_base or get_knowledge_base_content. It conveys only a generic verb+resource pairing that the tool name already provides.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many neighboring knowledge-base retrieval tools (get_documentation_from_knowledge_base, get_documentation_chunks_from_knowledge_base, get_knowledge_base_content). No prerequisites or context are stated, leaving the agent to infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_library_voicesC
Read-onlyIdempotent

Get Voices

ParametersJSON Schema
NameRequiredDescriptionDefault
ageNoAge used for filtering
pageNo
sortNoSort criteria. Must be one of: created_date, usage_character_count_1y, trending, cloned_by_count.
accentNoAccent used for filtering
genderNoGender used for filtering
localeNoLocale used for filtering
searchNoSearch term used for filtering
categoryNoVoice category used for filtering
featuredNoFilter featured voices
languageNoLanguage used for filtering
owner_idNoFilter voices by public owner ID
page_sizeNoHow many shared voices to return at maximum. Can not exceed 100, defaults to 30.
use_casesNoUse-case used for filtering
descriptivesNoSearch term used for filtering
reader_app_enabledNoFilter voices that are enabled for the reader app
include_custom_ratesNoInclude/exclude voices with custom rates
include_live_moderatedNoInclude/exclude voices that are live moderated
min_notice_period_daysNoFilter voices with a minimum notice period of the given number of days.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that — no indication of pagination behavior, result ordering, or that these are shared/library voices.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short only because it is under-specified, not because it is efficient. A two-word fragment carries no useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 18-parameter listing tool with no output schema, the definition should clarify what resource is being listed, that results are paginated (page/page_size exist), and how it differs from sibling voice endpoints. The read-only annotations cover safety but not this scope, leaving the definition substantially incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 94%, so the 18 parameters are already well documented by the schema itself. The description contributes no additional clarification, but the high coverage establishes the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Voices' merely restates the tool name with no verb-level specificity about scope, source, or filtering. It does not distinguish this tool from close siblings such as get_voices, get_user_voices_v2, get_similar_library_voices, or get_voice_by_id.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance whatsoever on when to use this tool versus the many other voice-retrieval siblings. An agent has nothing to route on beyond the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_live_countC
Read-onlyIdempotent

Get Live Count

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idNoThe id of an agent to restrict the analytics to.
agent_idsNoRestrict analytics to the union of the given agents. Takes precedence over `agent_id` when both are supplied.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds no behavioral context beyond those annotations, such as what is counted or whether the count is real-time.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short but under-specified rather than concise. It is front-loaded only in the sense that it is the title repeated, with no useful content to structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should explain what the returned count represents, but it does not. For a read tool with two optional filtering parameters, the definition is not complete enough to reliably guide invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so agent_id and agent_ids are fully documented in the schema itself. The description provides no parameter meaning, but the baseline is 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Live Count' is essentially a restatement of the tool name/title. It does not specify what entity or metric the live count belongs to, and it gives no differentiation from the many sibling 'get_*' tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, what context it applies to, or which alternatives exist. An agent cannot infer usage conditions from the description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_livekit_tokenC
Read-onlyIdempotent

Get Webrtc Token

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesAgent id (agent_…) or speech engine external id (seng_), resolved to the same underlying resource.
branch_idNoThe ID of the branch to use
version_idNoThe ID of the version to use
environmentNoThe environment to use for resolving environment variables (e.g. 'production', 'staging'). Defaults to 'production'.
participant_nameNoOptional custom participant name. If not provided, user ID will be used
debug_events_requestNoWhether to enable debug events. Only available for users with editor access to the agent.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, covering the basic safety profile. The description adds no behavioral context beyond this, such as token scope, expiry, authentication requirements, or what enabling debug events changes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only three words and functions as a bare fragment rather than a structured statement. While it is not verbose, it is under-specified for a tool with six parameters and does not front-load enough actionable context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that issues a LiveKit/WebRTC token and takes six parameters, the description omits nearly all contextual detail. The schema covers parameter meanings and annotations cover safety, but the description still does not explain the token's purpose, audience, or usage context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema itself already documents all six parameters thoroughly. The description adds no meaning beyond the schema, making the baseline score of 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Get Webrtc Token" essentially restates the tool name and title with a synonym, providing no additional specificity about what the token is for or which resource it grants access to. It does not distinguish this token from siblings such as get_single_use_token.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites, and no mention of the required agent_id or optional branch/version/environment context. The agent receives no routing information at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_mcp_routeC
Read-onlyIdempotent

Get Mcp Server

ParametersJSON Schema
NameRequiredDescriptionDefault
mcp_server_idYesID of the MCP Server.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is fully covered elsewhere. The description adds nothing beyond that: no mention of what a 'route' returns, whether the read hits a live remote server, or any error/not-found behavior. No contradiction with annotations, but no added behavioral context either.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At three words the text is not concise but under-specified; it lacks the information an agent would need rather than trimming redundancy. There is no front-loaded purpose statement, usage note, or anything that earns its place beyond the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only single-parameter getter this is on the lighter side of acceptable, but the description still fails to say what is returned (server config? tool list?) or how it relates to list_mcp_servers_route and get_mcp_tool_config_override_route. With no output schema and a dense sibling namespace, more explanation is warranted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is exactly one parameter and schema description coverage is 100%, so the schema already documents mcp_server_id as 'ID of the MCP Server.' The description contributes no additional meaning (e.g., ID format or where to obtain it), so the baseline of 3 for full schema coverage applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Mcp Server' essentially restates the tool name (get_mcp_route) without adding a verb-resource-object framing or any scope detail. It does not distinguish this single-server retrieval from the many sibling read tools such as list_mcp_servers_route, get_mcp_tool_config_override_route, or get_tools_route. This is close to a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives. An agent cannot tell from the description whether this fetches one server by ID (as the schema implies) or returns a collection, nor when to prefer it over list_mcp_servers_route. Only the required mcp_server_id parameter weakly implies the intent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_mcp_tool_config_override_routeC
Read-onlyIdempotent

Get Mcp Tool Configuration Override

ParametersJSON Schema
NameRequiredDescriptionDefault
tool_nameYesName of the MCP tool to retrieve config overrides for.
mcp_server_idYesID of the MCP Server.

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered by structured data. The description adds nothing beyond that: no mention of authentication needs, scope of the override lookup, behavior when no override exists, or error semantics. For a read tool whose annotations carry the burden, this is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single fragment restating the title. Brevity here reflects under-specification rather than efficient structure: no sentence earns its place because none conveys information beyond the name. A short but informative description would be superior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should describe what an override record contains or at least that it returns the override configuration, but it does not. Combined with no usage guidance and no sibling differentiation, the definition is insufficient for an agent to call it confidently amid many similar MCP config tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (mcp_server_id, tool_name) carry their own descriptions, so the schema does the heavy lifting. The tool description adds no additional syntax, format, or constraint information beyond what the schema already states, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (MCP tool config override) with an implied read action ('Get'), so the object of retrieval is identifiable. However, it is nearly a verbatim restatement of the tool name and title, and it does nothing to distinguish this tool from its close siblings (add_mcp_tool_config_override_route, update_mcp_tool_config_override_route, remove_mcp_tool_config_override_route). It is minimally viable rather than genuinely descriptive.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which alternative applies. An agent must infer from the tool name alone that this reads an existing override rather than creating, updating, or removing one. With four near-identical sibling tools in the list, the absence of any routing guidance is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_modelsC
Read-onlyIdempotent

Get Models

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond those structured fields — no scope, no result shape, no rate-limit or auth context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is not conciseness but under-specification; the text is too short to earn its place as a description. There is no front-loaded information beyond the title itself.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless read tool with no output schema, the description still needs to disambiguate the resource, and it does not. With many 'get/list' siblings in the catalog, an agent cannot determine what this returns or when to prefer it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and schema description coverage is 100%, so there is nothing for the description to clarify. Per the baseline for parameterless tools, a 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Models" restates the tool name/title with no added specificity: it does not say what kind of models (LLM models, speech models, fine-tuned models), what scope, or what is returned. In a sibling list containing list_available_llms, get_speech_engine, get_finetunes and get_tools_route, the agent cannot tell which resource 'models' refers to.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no usage guidance of any kind — no when-to-use, no when-not-to-use, and no mention of any alternative sibling. The agent is left to guess entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_or_create_rag_indexesC

Compute Rag Indexes In Batch

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesList of requested RAG indexes. Minimum 1, maximum 100 items.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is partly covered. The description adds only "In Batch," which is also implied by the schema's 1–100 item constraint, and says nothing about side effects, non-idempotent behavior, or what gets created or modified. It does not contradict the annotations, but it adds essentially no behavioral context beyond them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four words and front-loaded, but it is under-specified rather than effectively concise. It omits the core get-or-create distinction needed to select and invoke the tool correctly. Brevity here leaves the agent with unresolved ambiguity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should compensate by explaining get-or-create behavior, batch semantics, and likely return shape. Instead it omits all of that, despite the tool's non-read-only, non-idempotent, open-world annotations. The schema covers the parameter list, but the overall definition remains incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single parameter items is already documented as "List of requested RAG indexes. Minimum 1, maximum 100 items." The description adds no parameter syntax, item shape, or usage detail beyond what the schema provides. This aligns with the baseline 3 when the schema already carries parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase gives a verb and resource ("Compute Rag Indexes") plus batch scope, so it is not a bare tautology. However, it is vague about the actual get-or-create semantics of the tool and does not distinguish it from siblings like get_rag_indexes, get_rag_index_overview, rag_index_status, or delete_rag_index. An agent cannot tell from this description whether it fetches existing indexes, creates new ones, or does both.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool instead of get_rag_indexes, get_rag_index_overview, rag_index_status, delete_rag_index, or query_agent_knowledge_base_rag_route. No prerequisites, exclusions, or contextual triggers are stated. The description gives only an action label, leaving the agent to infer usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_phone_number_routeC
Read-onlyIdempotent

Get Phone Number

ParametersJSON Schema
NameRequiredDescriptionDefault
phone_number_idYesThe phone number ID. This is returned when a phone number is imported.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is fully covered structurally. The description adds no behavioral context beyond those annotations, such as what the retrieval returns or whether any permissions are required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two words and is under-specified rather than usefully concise. It neither front-loads nor communicates enough for reliable invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter getter with full schema coverage and safety annotations, the description is still too thin to distinguish it from list_phone_numbers_route or to confirm it retrieves a single phone number by ID. No output schema exists, but the description does not compensate with basic return or scope context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single phone_number_id parameter is already fully documented in the schema. The description adds no additional meaning or syntax beyond the schema, which makes 3 the appropriate baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Phone Number" restates the tool name and title without adding scope or distinguishing it from siblings such as list_phone_numbers_route, create_phone_number_route, update_phone_number_route, or delete_phone_number_route. It is a tautological purpose statement rather than a specific description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no when-to-use guidance, no exclusions, and no alternative tools. An agent receives no signal about why it should call this instead of list_phone_numbers_route or another sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_procedure_draft_routeC
Read-onlyIdempotent

Get Procedure Draft

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesAgent ID to get the procedure draft from
branch_idYesBranch ID to get the procedure draft from
procedure_idYesThe procedure ID

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond that — no indication of what a 'draft' contains, whether unpublished edits are visible, or any auth/scope context. With annotations doing all the work, the description is empty of added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The four-word phrase is technically free of waste, but this is under-specification rather than conciseness. There is no structure or front-loading of useful information because there is no information to front-load.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool requiring three IDs and returning a 'procedure draft' with no output schema, the description provides no context about what a draft is, how it relates to a published procedure, or what the returned data represents. The annotations cover safety but the description leaves a meaningful conceptual gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each of the three required parameters (agent_id, branch_id, procedure_id) is documented in the schema itself. The description adds no meaning beyond that, so the baseline of 3 for high-coverage schemas applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is essentially a restatement of the tool name 'get_procedure_draft_route' with no added specificity. It names a verb and resource, but does nothing to distinguish this from the numerous siblings such as get_procedure_route, update_procedure_draft_route, or delete_procedure_draft_route. This is tautological rather than clarifying.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites (agent/branch/procedure scoping), and no routing to alternatives like get_procedure_route for published procedures. The agent is left to infer everything from the name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_procedure_routeD
Read-onlyIdempotent

Get Procedure

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesAgent ID to get the procedure draft from
branch_idYesBranch ID to get the procedure draft from
version_idNoThe version ID to retrieve. If omitted, returns the version at branch HEAD.
procedure_idYesThe procedure ID
agent_version_idNoThe agent version ID to retrieve the procedure for.

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. However, the description adds nothing beyond the annotations—it does not explain what 'Procedure' means in this routing context or what the return behavior is. The description fails to supplement the structured data with any useful behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short at only two words, but this is under-specification rather than effective conciseness. It lacks front-loaded context or any structural elements that would help an agent understand the tool's role.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters and no output schema, the description is completely inadequate. It provides no information about the procedure retrieval process, the route context, or how the returned procedure differs from other procedure-related tools, leaving the agent without the necessary context to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all 5 parameters including agent_id, branch_id, procedure_id, version_id, and agent_version_id. With complete schema coverage, the baseline is 3, and the description adds no additional parameter meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Procedure' restates the tool name and adds no independent information about what the tool actually does. It does not distinguish this tool from siblings like get_procedure_draft_route or list_procedures_route, leaving the agent unable to tell them apart based on the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as get_procedure_draft_route, list_procedures_route, or compile_procedures_route. The agent receives no context about prerequisites or appropriate scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_project_by_idC
Read-onlyIdempotent

Get Studio Project

ParametersJSON Schema
NameRequiredDescriptionDefault
share_idNoThe share ID of the project
project_idYesThe ID of the Studio project.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, covering the safety and idempotency profile. The description adds nothing beyond this—no detail on what happens if the project_id is invalid, whether share_id bypasses auth, or what the response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short and has zero waste, but it is arguably under-specified rather than concise. For a one-line description, brevity is fine, but it leaves too much unsaid.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a one-line description, the agent has insufficient context about what a 'Studio Project' is, how it differs from get_projects, and what return shape to expect. Annotations cover safety but not operational context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both project_id and share_id fully documented in the schema. The description adds no parameter semantics beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Studio Project' essentially restates the tool name get_project_by_id in a slightly reworded form. It does not specify what a 'Studio Project' is or what distinguishes it from get_projects, edit_project, delete_project, or get_project_snapshot_endpoint, all present among siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to call this tool versus get_projects, get_project_snapshot_endpoint, or any other project-related sibling. An agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_project_muted_tracks_endpointC
Read-onlyIdempotent

Get Project Muted Tracks

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesThe ID of the Studio project.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered structurally. The description adds nothing beyond that — no mention of auth requirements, pagination, ordering, or what 'muted tracks' actually contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but brevity here reflects under-specification rather than disciplined conciseness — a single fragment with no front-loaded detail, scope, or return information to justify its size.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read endpoint with no output schema, the description should at least hint at what is returned (a list of muted tracks, keys, pagination). None of that is present, leaving an agent unable to anticipate the response shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single project_id parameter and 100% schema description coverage, the schema fully documents the parameter. The description contributes no additional meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Project Muted Tracks' merely restates the tool name/title verbatim without adding scope, format, or sibling differentiation. An agent learns nothing beyond what the identifier already conveys.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool rather than related endpoints such as get_project_by_id or get_chapters, and no prerequisites or context are given. Usage must be fully inferred from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_projectsB
Read-onlyIdempotent

List Studio Projects

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered by structured data. The description adds nothing beyond that — no pagination behavior, no result ordering, no scope of 'Studio Projects' against other project types. With annotations doing all the work, the description contributes no additional behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single three-word phrase that is front-loaded and wastes nothing. It is arguably too terse rather than too long, but as conciseness it is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool with rich annotations and no output schema, the description is minimally adequate. It still leaves open whether results are paginated, how many projects are returned, or which project types count as 'Studio Projects', which an agent would want before calling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there is no parameter information the description needs to supply and none is missing. No syntax, format, or filtering detail is warranted for a parameterless call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a clear verb ('List') and a specific resource ('Studio Projects'), so an agent knows exactly what entity it operates on. It does not, however, distinguish itself from the many sibling list/get tools (e.g. get_project_by_id, list_assets), nor does it state the scope of the listing. Clear but undifferentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus get_project_by_id or the other project-related endpoints, and no prerequisites or exclusions are mentioned. Usage must be inferred entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_project_snapshot_endpointC
Read-onlyIdempotent

Get Project Snapshot

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesThe ID of the Studio project.
project_snapshot_idYesThe ID of the Studio project snapshot.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive, so the safety profile is covered. The description adds nothing beyond that — no mention of what a snapshot contains, whether it includes archive/audio links, or any retrieval constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words with no padding, so it is concise, but its brevity comes at the cost of usefulness rather than being tight and informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with no output schema and full annotation coverage, the description still should convey what the snapshot represents and how it relates to the plural listing tool. Nothing here helps the agent decide or interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both project_id and project_snapshot_id documented in the schema itself. The description adds no parameter meaning beyond that, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description merely restates the tool name ('Get Project Snapshot') without elaborating on what a snapshot is retrieved for or how it differs from siblings. With siblings like get_project_snapshots (plural list) and get_chapter_snapshot_endpoint, the description misses an easy chance to distinguish itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus get_project_snapshots or get_project_by_id. The agent is left to infer that this fetches a single snapshot by ID.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_project_snapshotsB
Read-onlyIdempotent

List Studio Project Snapshots

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesThe ID of the Studio project.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds no behavioral context beyond that, such as pagination, ordering, or authentication requirements, and it does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase that is front-loaded and contains no filler. It is appropriately sized for a simple list operation, though its brevity contributes to gaps in other dimensions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with full schema coverage and rich annotations, the description is minimally adequate. It does not explain return behavior or pagination, but no output schema exists and the annotations already cover safety, so the remaining gap is moderate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema documents the single project_id parameter at 100% coverage with a clear description, so the schema already carries parameter semantics. The tool description adds no additional meaning about the parameter, which is acceptable given the high schema coverage baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('Studio Project Snapshots'), so an agent can tell it is a plural list operation. However, it does not distinguish this tool from close siblings such as get_project_snapshot_endpoint or get_chapter_snapshots, which could cause selection ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. The only implied usage is that the tool lists snapshots for a project, but nothing tells the agent when this is preferable to a single-snapshot or chapter-snapshot tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pronunciation_dictionaries_metadataC
Read-onlyIdempotent

Get Pronunciation Dictionaries

ParametersJSON Schema
NameRequiredDescriptionDefault
sortNoWhich field to sort by, one of 'created_at_unix' or 'name'.
cursorNoUsed for fetching next page. Cursor is returned in the response.
page_sizeNoHow many pronunciation dictionaries to return at maximum. Can not exceed 100, defaults to 30.
sort_directionNoWhich direction to sort the voices in. 'ascending' or 'descending'.
include_archivedNoWhether to include archived pronunciation dictionaries in the response.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered elsewhere. The description adds nothing beyond that - no mention of pagination behavior, default page size, or archive handling, despite these being operationally relevant for a list tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but the brevity comes from under-specification rather than efficiency - there is no front-loaded scope, constraint, or routing information. A three-word restatement of the name does not earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter list endpoint with no output schema, the description should at least sketch what is returned and how paging works. None of that is present, leaving the agent to infer everything from the parameter names.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters (sort, cursor, page_size, sort_direction, include_archived) are documented in the schema itself. The description adds no additional meaning beyond the schema, which is the baseline-3 expectation when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Pronunciation Dictionaries' merely restates the tool name and adds no specificity about scope, filters, or return content. It fails to distinguish itself from the closely named sibling get_pronunciation_dictionary_metadata (singular), which an agent could easily confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this list endpoint versus the singular get_pronunciation_dictionary_metadata or update_pronunciation_dictionaries. No prerequisites, pagination context, or filtering guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pronunciation_dictionary_metadataC
Read-onlyIdempotent

Get Metadata For A Pronunciation Dictionary

ParametersJSON Schema
NameRequiredDescriptionDefault
pronunciation_dictionary_idYesThe id of the pronunciation dictionary

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, giving a full safety profile, so the description bears a lower burden. However, the description adds nothing beyond the name – no note on whether the ID must reference an existing dictionary or what happens on a miss. This is the baseline acceptable level given annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is essentially a Title Case restatement of the tool name. It is concise in length but wastes that brevity on redundancy rather than useful front-loaded information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description should explain what metadata is returned, yet it says nothing about the return shape. For a read tool sitting among many similarly named dictionary endpoints, it fails to give an agent enough to call it correctly versus its siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter 'pronunciation_dictionary_id' is documented in the schema. The description adds no syntax, format, or source guidance for the ID, so it cannot exceed the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name in title case, providing a generic verb+resource ('get metadata for a pronunciation dictionary') without distinguishing it from close siblings like get_pronunciation_dictionaries_metadata (plural) or get_pronunciation_dictionary_version_pls. It is not misleading, but it adds no differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus the clearly related sibling get_pronunciation_dictionaries_metadata or get_pronunciation_dictionary_version_pls. The agent is left to infer context from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pronunciation_dictionary_version_plsB
Read-onlyIdempotent

Get A Pls File With A Pronunciation Dictionary Version Rules

ParametersJSON Schema
NameRequiredDescriptionDefault
version_idYesThe id of the pronunciation dictionary version
dictionary_idYesThe id of the pronunciation dictionary

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish that this is a safe, idempotent, read-only operation, so the safety profile is covered. The description adds one piece of genuinely useful behavioral context — that the return artifact is a PLS file (lexicon format) rather than JSON — but says nothing about size, encoding, or how rules are represented. Adequate but thin beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no filler, which is appropriate for a simple getter. The choppy capitalization ('Get A Pls File') slightly hurts readability but costs no tokens.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only fetch with full schema coverage and safety annotations already present, the description covers the essential missing piece: the nature of the returned artifact (a PLS file). Nothing critical about invocation is absent, though a note on file size or rule structure would have fully closed the gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (both dictionary_id and version_id are documented in the schema), so the baseline is 3. The description adds no further meaning about parameter format or expected values, and it does not need to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Get') and resource (a PLS file containing a pronunciation dictionary version's rules), so the basic operation is inferable. However, it does not differentiate from siblings such as get_pronunciation_dictionary_metadata or get_pronunciation_dictionaries_metadata, and the phrasing is awkward enough to read as a restatement of the name rather than a distinct clarification.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when this tool should be used instead of the metadata siblings, nor any prerequisites (e.g., that a valid dictionary_id/version_id pair must already exist). Given roughly a dozen pronunciation-dictionary-related siblings, this absence leaves the agent to guess.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_public_llm_expected_cost_calculationC

Calculate Expected Llm Usage

ParametersJSON Schema
NameRequiredDescriptionDefault
rag_enabledYesWhether RAG is enabled.
prompt_lengthYesLength of the prompt in characters.
number_of_pagesYesPages of content in PDF documents or URLs in the agent's knowledge base.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description adds nothing beyond those structured hints: it does not say what side effects may occur, what authentication or rate limits apply, or whether the calculation is cached. It does not contradict the annotations, but it is uninformative.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only four words, but it is drastically undersized for a calculation tool with three required inputs and no output schema. Brevity here reflects missing information rather than efficient structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should explain what the calculation returns and what distinguishes it from the agent-scoped sibling. Instead it provides only a vague phrase, leaving the agent without enough context to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the three required parameters are already fully documented in the schema. The description adds no additional meaning such as valid ranges, units, or interaction between prompt_length, number_of_pages, and rag_enabled. Baseline 3 is appropriate when the schema carries the full parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the title as a bare imperative phrase and does not distinguish this public calculation tool from the sibling get_agent_llm_expected_cost_calculation. It also drops the word 'cost' from the tool name, replacing it with the vaguer 'usage', so an agent cannot confidently tell what is being calculated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool, when not to use it, or which alternatives exist. The sibling get_agent_llm_expected_cost_calculation is never mentioned, leaving the agent with no routing signal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pvc_sample_audioC
Read-onlyIdempotent

Retrieve Voice Sample Audio

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
sample_idYesSample ID to be used
remove_background_noiseNoIf set will remove background noise for voice samples using our audio isolation model. If the samples do not include background noise, it can make the quality worse.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds no behavioral context beyond that, such as whether it requires authentication, returns a URL or binary, or has rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The four-word phrase is concise and front-loaded, but it is under-specified for a tool with three parameters and no output schema. It earns no place beyond repeating the title, so it is too minimal for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description does not explain what the returned audio is, its format, or how it relates to sibling tools. Annotations and schema cover safety and parameters, but the description leaves the agent without enough context to know when this tool is the right choice.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all three parameters including voice_id, sample_id, and remove_background_noise. The description adds no parameter meaning beyond what the schema provides; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb (Retrieve) and resource (Voice Sample Audio), but essentially restates the tool name with no additional detail or sibling differentiation. An agent cannot distinguish it from get_audio_from_sample or get_pvc_sample_visual_waveform based on this description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance or alternatives are provided. The description does not mention when to choose this tool over get_audio_from_sample or any other audio retrieval sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pvc_sample_speakersC
Read-onlyIdempotent

Retrieve Speaker Separation Status

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
sample_idYesSample ID to be used

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive and openWorld behavior, so the safety profile is covered. The description adds essentially nothing beyond the title, not explaining what 'speaker separation status' contains (e.g., in-progress/complete, speaker count) or whether the sample must first be processed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded phrase with no wasted words, but it is under-specified rather than genuinely concise – the brevity comes at the cost of missing information rather than distilling it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with two required params and no output schema, the description should clarify what the returned status represents and any interplay with sample processing. It provides none of this, and the name/description mismatch leaves the tool's behavior ambiguous.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both voice_id and sample_id are documented in the schema (voice_id even links the voices listing endpoint). The description adds no parameter meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a plausible verb+resource (retrieve speaker separation status), but it does not match the tool name get_pvc_sample_speakers cleanly, leaving the agent unsure whether it returns speaker data or a job status. It also does not differentiate from siblings like get_pvc_sample_audio or start_speaker_separation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to call this tool versus alternatives such as start_speaker_separation (to launch separation) or get_pvc_sample_audio (to fetch audio). No prerequisites or context are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pvc_sample_visual_waveformC
Read-onlyIdempotent

Retrieve Voice Sample Visual Waveform

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
sample_idYesSample ID to be used

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, covering the safety profile. The description adds nothing on top of that — no mention of what the waveform is (image bytes, URL, JSON points), rate limits, or auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with zero padding, so it is not verbose. However, it is under-specified rather than efficient — the brevity comes from omission, not from disciplined writing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only retrieval tool with no output schema, the description should at minimum indicate what a 'visual waveform' return looks like (image, data array, URL). None of that is present, leaving the agent unable to anticipate the response shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with both voice_id and sample_id documented in the schema, so the baseline of 3 applies. The description contributes no additional parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Retrieve Voice Sample Visual Waveform' is essentially a restatement of the tool name get_pvc_sample_visual_waveform, adding no verb or scope beyond the title. It does implicitly distinguish itself from sibling get_pvc_sample_audio by the word 'visual', but the agent gains no explanatory content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative-tool guidance is provided. With siblings like get_pvc_sample_audio and get_pvc_sample_speakers, the agent is left to infer which retrieval variant to call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pvc_voice_captchaC
Read-onlyIdempotent

Get Pvc Voice Captcha

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds nothing about auth requirements, rate limits, captcha expiration, or what the returned captcha token is meant for.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four words that merely echo the name, with no front-loaded purpose statement. This is under-specification rather than conciseness — nothing useful is conveyed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must explain the return value, but it does not (is it a challenge, a token, a session ID from request_pvc_manual_verification?). For a cryptically named read operation, the definition leaves the agent without enough to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single voice_id parameter is fully documented in the schema, including a pointer to the voices listing endpoint. Baseline 3 applies; the description adds no further parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name with no added explanation. It never says what a 'PVC voice captcha' is, what it is for, or how it differs from the sibling verify_pvc_voice_captcha.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or prerequisite information is given. The closely related verify_pvc_voice_captcha sibling is never mentioned or contrasted, leaving the agent to guess which to call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_rag_indexesC
Read-onlyIdempotent

Get Rag Indexes Of The Specified Knowledgebase Document.

ParametersJSON Schema
NameRequiredDescriptionDefault
documentation_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds nothing further — no note on what the indexes represent, whether an empty result is possible, or how it relates to index creation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no wasted words. It is efficient, though the title-case phrasing and lack of any routing clause keep it from being exemplary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read tool with no output schema and strong annotations, this is minimally adequate. It leaves the return shape (a list of RAG indexes) and the distinction from get_or_create_rag_indexes unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single documentation_id parameter is fully documented in the schema. The description's phrase 'Specified Knowledgebase Document' corroborates the parameter but adds no syntax, format, or sourcing detail beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get') and resource ('RAG Indexes') scoped to a knowledgebase document, which is clearer than a bare name restatement. However, it does not differentiate itself from near-identical siblings such as get_or_create_rag_indexes, get_rag_index_overview, or rag_index_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of when a caller should prefer this over get_or_create_rag_indexes or rag_index_status. The agent must infer selection purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_rag_index_overviewC
Read-onlyIdempotent

Get Rag Index Overview.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that: no mention of what the overview summarizes, aggregation scope, cost, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is extremely short, which is concise, but the brevity reflects under-specification rather than efficient communication. There is no front-loaded value beyond the name itself.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should explain what the 'overview' returns, yet it says nothing. For a discovery tool sitting among many RAG-related siblings, this leaves the agent without enough context to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and a schema description coverage of 100%, so there are no parameter semantics to document. The baseline of 4 applies for a no-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a near-tautological restatement of the name and title, 'Get Rag Index Overview.' It does not clarify what an 'overview' contains or how it differs from siblings like get_rag_indexes, rag_index_status, or get_or_create_rag_indexes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance is provided. With closely named siblings such as get_rag_indexes and rag_index_status, the agent has no signal for selecting this tool over the alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_resource_metadataC
Read-onlyIdempotent

Get Resource

ParametersJSON Schema
NameRequiredDescriptionDefault
resource_idYesThe ID of the target resource.
resource_typeYesResource type of the target resource.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, covering the safety profile. The description adds no behavioral context beyond those annotations, such as what metadata is returned for each resource type or any pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The two-word description is not concise in a helpful sense; it is under-specified for a tool with two required parameters and broad applicability. It lacks front-loaded, actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of fetching metadata across many resource types and the absence of an output schema, the description should clarify what metadata is returned or any type-specific variations. It does not, leaving a significant gap in contextual completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including a detailed enum for resource_type, so the schema fully documents both parameters. The description adds no additional meaning, which establishes the baseline score of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Resource' is nearly a tautology of the tool name 'get_resource_metadata' and omits the crucial 'metadata' aspect. It does not distinguish this generic resource-fetching tool from the many other get_* siblings that retrieve specific resource types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives, nor any prerequisites or context. The description offers no usage instructions at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_secret_dependencies_routeC
Read-onlyIdempotent

Get Secret Dependencies By Type

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
page_sizeNoHow many dependency items to return per page.
secret_idYes
resource_typeYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond the name — no mention of pagination behavior, that results are paginated via cursor, or what a dependency item represents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is a single short phrase with zero waste, but it is under-specified rather than genuinely concise — it reads as a title fragment, not an informative description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 4 parameters, two of them undocumented, no output schema, and no sibling routing, the description leaves significant gaps. An agent would not know how to correctly interpret the resource_type filter or paginate results from the description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: cursor and page_size are documented in the schema, but secret_id and resource_type have no schema descriptions. The description's 'By Type' weakly implies resource_type, but it does not explain the enum values (tools/agents/phone_numbers) or that secret_id identifies the secret whose dependencies are listed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a verb and resource ('Get Secret Dependencies') and adds a scope qualifier ('By Type'), which hints at the resource_type filter. However, it largely restates the tool name and gives no differentiation from close siblings such as get_secret_route, get_secrets_route, or get_tool_dependent_agents_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like get_secret_route (a single secret) or the other '*dependent_agents*' tools. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_secret_routeC
Read-onlyIdempotent

Get Convai Workspace Secret

ParametersJSON Schema
NameRequiredDescriptionDefault
secret_idYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false. The description adds only the workspace scope implied by the resource name and does not disclose any additional behavior such as sensitivity of the returned value, required permissions, or whether the secret value is revealed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loaded, with no filler. However, it is a sentence fragment rather than a structured sentence, and its extreme brevity comes at the cost of useful context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a secret-retrieval tool, the description omits the required secret_id semantics and does not explain what the tool returns or any access considerations. The annotations cover safety traits, but the description is not complete enough to ensure correct invocation beyond a basic guess.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one required parameter, secret_id, with 0% description coverage, and the description does not mention the parameter or explain its meaning. Because low schema coverage requires the description to compensate, this is a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Get Convai Workspace Secret'. This is clear enough to distinguish a single-secret retrieval from the plural get_secrets_route, but it does not explicitly differentiate from nearby siblings such as get_secret_dependencies_route or explain that retrieval is by secret ID.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. It does not mention get_secrets_route for listing secrets or get_secret_dependencies_route for dependency inspection, nor does it state any prerequisite or context for retrieving a secret.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_secrets_routeC
Read-onlyIdempotent

Get Convai Workspace Secrets

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
searchNoIf specified, returns only secrets whose names start with this string.
page_sizeNoHow many documents to return at maximum. Can not exceed 100. If not provided, returns all secrets.
dependency_limitNoMaximum number of dependent resources (tools, agents, phone numbers) to return per secret. Can not exceed 100.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered externally. The description itself adds nothing beyond the annotations - it does not mention pagination, the cursor-based listing behavior, or the fact that it returns secrets rather than a single secret.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence, so it is certainly concise and front-loaded. However, that brevity comes at the cost of underspecification, and the tool title merely repeats the description verbatim in the annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool is a paginated listing operation with four parameters, a near-duplicate sibling (get_secret_route), and no output schema, the description is far too thin. It should at minimum clarify that it lists secrets and how it relates to the single-secret sibling, neither of which it does.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the four parameters (cursor, search, page_size, dependency_limit) are already fully documented in the schema. The description adds no parameter meaning beyond this. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Convai Workspace Secrets' names a resource but the verb 'Get' is ambiguous between fetching a single secret and listing all secrets. It does not distinguish itself from its close sibling get_secret_route, which presumably fetches one secret. An agent cannot tell these apart without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all. With near-duplicate siblings such as get_secret_route and get_secret_dependencies_route in the list, the description gives the agent no way to choose this tool over those.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_service_account_api_keys_routeD
Read-onlyIdempotent

Get Service Account Api Keys Route

ParametersJSON Schema
NameRequiredDescriptionDefault
service_account_user_idYes

TDQS

D1.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=true, openWorldHint=true, idempotentHint=true, destructiveHint=false). The description adds zero behavioral context beyond what annotations already declare. The only hint is the word 'Get', which corroborates readOnlyHint, but this is already explicit in structured data. With annotations shouldering the burden, the description earns a low score for adding no unique value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is concise to a fault: a single tautological sentence with no information. Brevity without substance does not earn credit, as every word merely restates the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a required parameter, 0% schema description coverage, no output schema, and no meaningful description, the definition is inadequate. An agent lacks sufficient context to invoke the tool correctly or understand its behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and there is 1 required parameter (service_account_user_id). The description does not even mention the parameter, let alone explain its format, source, or constraints. The description is completely silent on parameter semantics, failing to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim repeat of the tool name and title: 'Get Service Account Api Keys Route'. This is a tautology that conveys no additional purpose information. An agent cannot distinguish this from siblings like get_secret_route or get_tool_route based on the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when or why to use this tool versus alternatives. There are many 'get_*' route tools in the sibling list, and the description provides no routing condition or context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_settings_routeC
Read-onlyIdempotent

Get Convai Settings

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that — no indication of what settings are returned, whose settings they are, or any auth/scope constraints, which matters given multiple settings-related siblings.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short phrase with no padding, so nothing is wasted, but the brevity reflects under-specification rather than efficient information density. It is minimally acceptable in size.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no parameters, the description carries the burden of explaining what this endpoint returns and which settings scope it covers. Given the crowded family of settings tools (get_dashboard_settings_route, update_settings_route, patch_agent_settings_route), omitting the scope leaves the agent unable to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the calibration the baseline is 4. The description cannot add parameter meaning because there is nothing to parameterize.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name: 'Get Convai Settings' says only that it retrieves some unspecified settings. It does not specify scope (workspace, dashboard, agent, convai account) and gives no basis to distinguish it from siblings like get_dashboard_settings_route or update_settings_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus the related settings tools, no prerequisites, and no exclusions. An agent must guess based on the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_signed_url_deprecatedC
Read-onlyIdempotent

Get Signed Url Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesAgent id (agent_…) or speech engine external id (seng_), resolved to the same underlying resource.
branch_idNoThe ID of the branch to use
version_idNoThe ID of the version to use
environmentNoThe environment to use for resolving environment variables (e.g. 'production', 'staging'). Defaults to 'production'.
debug_events_requestNoWhether to enable debug events. Only available for users with editor access to the agent.
include_conversation_idNoWhether to include a conversation_id with the response. If included, the conversation_signature cannot be used again.

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description's only addition is the 'Deprecated' marker, which is a genuinely useful signal, but it omits the replacement path, expiry behavior of the signed URL, or any auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence with no padding, but its brevity comes from under-specification rather than economy: every meaningful word is already in the tool name. There is no front-loaded explanation of what the caller receives.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Six parameters, a required agent_id, environment/branch/version resolution, and a conversation-id side effect ('cannot be used again') go entirely unexplained, and there is no output schema to compensate. For a deprecated endpoint whose semantics an agent cannot infer, this is inadequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all six parameters, so the schema fully documents agent_id, branch_id, version_id, environment, debug_events_request, and include_conversation_id. The description adds nothing further about these parameters, which matches the baseline of 3 when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name almost verbatim ("Get Signed Url Deprecated upstream") without saying what resource a signed URL is produced for or what 'upstream' means. It never distinguishes itself from siblings like get_conversation_signed_link, get_knowledge_base_source_file_url, or get_single_use_token. This is tautology rather than purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this tool, when not to, or which alternative signed-URL/token tool to prefer. The word 'Deprecated' hints an agent should look elsewhere, but the description never names the replacement or the condition that selects this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_similar_library_voicesC

Get Similar Library Voices Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
top_kNoNumber of most similar voices to return. If similarity_threshold is provided, less than this number of voices may be returned. Values range from 1 to 100.
audio_file_pathNoFile for "audio_file". Local path.
audio_file_base64NoBase64 contents for "audio_file". Use this when the server cannot read your local disk.
audio_file_filenameNoFilename to send for "audio_file". Some endpoints infer the audio format from it.
similarity_thresholdNoThreshold for voice similarity between provided sample and library voices. Values range from 0 to 2. The smaller the value the more similar voices will be returned.

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, covering the safety profile. The description adds that the tool spends ElevenLabs credits, a meaningful cost behavior not present in annotations. It does not explain how many credits are consumed, whether the cost is per call, or any auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loads the purpose, but it reads as a run-on fragment ('Get Similar Library Voices Spends ElevenLabs credits') with missing punctuation and structure. It is concise but not well-crafted, and the second idea is not cleanly separated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should clarify what is returned and what input is required, but it does neither. It fails to explain that an audio sample is needed (via audio_file_path, audio_file_base64, or audio_file_filename) or how top_k and similarity_threshold affect results. Annotations and schema cover some context, but key invocation details are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are documented in the schema. The description adds no parameter detail beyond the tool name, and provides no explanation of the audio input options or similarity threshold beyond what the schema already states. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a general verb and resource ('Get Similar Library Voices'), but it does not specify what the voices are similar to (a provided audio sample) or differentiate itself from sibling tools like get_similar_voices_for_speaker. The second clause about credits is a cost note, not purpose, leaving the purpose minimally viable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions. The only contextual detail is a cost warning ('Spends ElevenLabs credits'), which is a behavioral note rather than usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_similar_voices_for_speakerB
Read-onlyIdempotent

Search The Elevenlabs Library For Voices Similar To A Speaker. Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
dubbing_idYesID of the dubbing project.
speaker_idYesID of the speaker.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety and idempotency profile is covered. The description adds a genuinely useful behavioral fact — 'Deprecated upstream' — but fails to say what to use instead, leaving the agent unable to act on the warning.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with no padding; the core purpose leads and the deprecation note trails. Slightly cryptic phrasing of the deprecation notice keeps it from being a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only tool with full schema coverage and no output schema, the definition is minimally adequate. However, a deprecated tool should name its replacement or state the migration path, and that critical context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both dubbing_id and speaker_id are already documented in the schema. The description adds no extra semantics about how the speaker relates to the dubbing project or how results are scoped. Baseline 3 applies when the schema carries the parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource: searching the ElevenLabs voice library for voices similar to a given speaker. An agent can grasp the intent immediately, though it does not differentiate itself from the close sibling get_similar_library_voices, leaving some ambiguity about which 'similar voices' tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus alternatives like get_similar_library_voices or get_library_voices. The 'Deprecated upstream' note signals caution but stops short of naming the replacement or saying whether it should be avoided entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_single_use_tokenD

Create Single Use Token

ParametersJSON Schema
NameRequiredDescriptionDefault
token_typeYes

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, and openWorldHint=true, so the write/open-world nature is covered there. The description adds nothing beyond that: no token lifetime, no auth requirements, no note that repeated calls mint distinct tokens. A security-sensitive credential-minting tool warrants more, but with annotations covering the safety profile this is a modest gap rather than a severe one.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four words with no front-loaded constraint or context. This is under-specification rather than conciseness — the sentence ends before conveying anything an agent could act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A token-minting tool with one undocumented enum parameter and no output schema needs the description to carry nearly everything: what the token is for, its lifetime, its scope, and how the three token types differ. None of that is present, so the definition is unusable for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the single token_type parameter and it does not. The enum values realtime_scribe, batch_scribe, and tts_websocket are self-evident only at a glance; the description never explains which token to request for which use case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Create Single Use Token" restates the tool's own resource without adding scope, and it names no audience, format, or distinguishing trait versus the many other token/credential tools in the sibling list. It is effectively a tautology of the title, and it even conflicts with the tool name's "get_" prefix, which will confuse an agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this instead of alternatives such as get_livekit_token or create_service_account_api_key, nor any prerequisites or call-order guidance. The agent is left to infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_speaker_audioC
Read-onlyIdempotent

Retrieve Separated Speaker Audio

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
sample_idYesSample ID to be used
speaker_idYesSpeaker ID to be used, you can use GET https://api.elevenlabs.io/v1/voices/{voice_id}/samples/{sample_id}/speakers to list all the available speakers for a sample.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds nothing about the retrieval workflow, what the returned audio represents, or how it relates to start_speaker_separation, so it contributes no behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no wasted words, which is structurally clean, but its brevity comes at the cost of under-specification rather than intentional economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-required-parameter retrieval tool with no output schema, the description should at minimum explain that it fetches the output of speaker separation and name any prerequisite call. Instead it provides only a title-like phrase, leaving a meaningful gap for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter (voice_id, sample_id, speaker_id) is documented, including URLs for listing valid values. The description adds no parameter meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Retrieve Separated Speaker Audio' gives a verb and a resource, but it is essentially a restatement of the tool name and title. It does not distinguish this tool from close siblings such as get_audio_from_sample, get_pvc_sample_audio, or start_speaker_separation, leaving the agent to infer which one actually returns speaker-separated audio.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, and no mention of a prerequisite workflow (e.g., start_speaker_separation must run first) even though such a sibling exists. The agent gets no context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_speech_engineC
Read-onlyIdempotent

Get Speech Engine

ParametersJSON Schema
NameRequiredDescriptionDefault
speech_engine_idYesThe speech engine ID (accepts seng_ or agent_ prefix)

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered elsewhere. The description adds no behavioral context of its own, such as error behavior for a missing or invalid engine ID.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is short, but it is under-specification rather than conciseness: three words that carry no information beyond the title. There is nothing to front-load because nothing substantive is said.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter read tool the annotations cover safety and the schema covers the parameter, but with no output schema the description could have said what a speech engine record returns. As written it leaves the agent with only the tool name.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter is documented in the schema, including the accepted seng_/agent_ prefix forms. The description contributes nothing beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description only restates the tool name/title, 'Get Speech Engine'. It does not specify what a speech engine record contains or how retrieval differs from siblings like list_speech_engines, create_speech_engine, update_speech_engine, or delete_speech_engine.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites, and no reference to the sibling tools. An agent must infer from the name alone that this fetches one engine by ID rather than listing engines.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_speech_historyC
Read-onlyIdempotent

List Generated Items

ParametersJSON Schema
NameRequiredDescriptionDefault
searchNosearch term used for filtering
sourceNoSource of the generated history item
model_idNoModel ID to filter history items by.
voice_idNoVoice ID to be filtered for, you can use GET https://api.elevenlabs.io/v1/voices to receive a list of voices and their IDs.
page_sizeNoHow many history items to return at maximum. Can not exceed 1000, defaults to 100.
sort_directionNoSort direction for the results.
date_after_unixNoUnix timestamp to filter history items after this date (inclusive).
date_before_unixNoUnix timestamp to filter history items before this date (exclusive).
start_after_history_item_idNoAfter which ID to start fetching, use this parameter to paginate across a large collection of history items. In case this parameter is not provided history items will be fetched starting from the most recently created one ordered descending by their creation date.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered elsewhere. The description adds nothing on top of that — no mention of default ordering, pagination behavior, result caps, or what the returned items contain.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase with no wasted words and is trivially front-loaded, but its brevity reflects under-specification rather than disciplined concision. It neither misleads nor informs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nine-parameter listing/filtering tool with no output schema, the description should at minimum clarify that it returns a filtered, paginated set of speech-history items and note the default sort. Given many closely named siblings, this is insufficient for correct selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 100%, the schema fully documents all nine filter parameters (search, source, model_id, voice_id, pagination, date bounds), so the baseline of 3 applies. The description contributes no additional parameter meaning such as filter interaction or defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"List Generated Items" names a verb and a resource, but "Generated Items" is so generic it does not resolve to the speech-generation history the tool name implies, nor does it distinguish this from siblings like list_text_to_speech_generations or list_image_generations. It is close to a restatement that carries no disambiguating information.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this listing tool versus get_speech_history_item_by_id, download_speech_history_items, or the other list_* generation tools. Nothing tells the agent which context selects this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_speech_history_item_by_idC
Read-onlyIdempotent

Get History Item

ParametersJSON Schema
NameRequiredDescriptionDefault
history_item_idYesHistory item ID to be used, you can use GET https://api.elevenlabs.io/v1/history to receive a list of history items and their IDs.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so safety is covered. The description adds no behavioral context beyond those annotations, such as return contents, auth requirements, or whether the item includes metadata versus audio.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loaded, but it is under-specified rather than efficiently concise. Like a minimal placeholder, it lacks enough structure to orient an agent beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-ID tool, the schema and annotations cover much of the need, but with no output schema the description should at least hint at what a speech history item contains or what the caller gets back. 'Get History Item' leaves that unclear and does not fully compensate for the missing output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single parameter history_item_id is fully documented in the schema, including how to obtain IDs. The description adds no further parameter meaning, so the baseline of 3 applies when the schema carries the semantic load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get History Item' essentially restates the tool name/title without adding specificity. It names a verb and a generic resource, but does not clearly distinguish this from siblings like get_speech_history, download_speech_history_items, or get_audio_full_from_speech_history_item.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives. It does not mention when-not to use it, prerequisites, or relevant sibling tools, leaving selection entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_test_invocation_routeC
Read-onlyIdempotent

Get Test Invocation

ParametersJSON Schema
NameRequiredDescriptionDefault
test_invocation_idYesThe id of a test invocation. This is returned when tests are run.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, covering the safety profile. The description adds no behavioral context beyond those annotations—no mention of what is returned, authentication needs, rate limits, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The two-word phrase is concise but under-specified rather than efficient; it omits any useful detail that would help an agent invoke the tool correctly. This matches the calibration example where extreme brevity without substance scores low on conciseness/structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple schema and annotation coverage, the description is still too incomplete to route the agent effectively among many test-related siblings. It does not explain what a test invocation is, when to use this get-by-id operation, or how it relates to list_test_invocations_route.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter test_invocation_id is fully documented in the schema itself. The description adds no meaning beyond the schema, so the baseline score of 3 applies when the schema carries the entire parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Test Invocation' simply restates the tool name and title without adding scope, distinguishing it from siblings like list_test_invocations_route, or clarifying what resource is retrieved. This is a tautology rather than a substantive purpose statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as list_test_invocations_route, run_agent_test_suite_route, or other test-related siblings. No prerequisites, context, or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_text_to_speech_generationD
Read-onlyIdempotent

Get Speech Generation

ParametersJSON Schema
NameRequiredDescriptionDefault
generation_idYes

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing further: no note on retrieval semantics, authorization, error behavior when the id is absent, or whether the response includes audio artifacts vs metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The phrase is short and front-loaded, but it is under-specified rather than concise — it is a fragment carrying almost no information. Brevity here reflects missing content, not efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter lookup tool with no output schema, the description should at minimum explain that generation_id comes from a prior generation call or a listing endpoint. Annotations cover the safety profile, but the description leaves an agent without enough to call this correctly in a crowded sibling set.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description provides no meaning for generation_id — not its accepted format, nor how to obtain one (e.g., from create_text_to_speech_generation or list_text_to_speech_generations). With one undocumented parameter, the description fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Speech Generation' is a near-tautological restatement of the tool name and annotation title. It conveys the verb 'get' and the resource only because the name already does; it adds no scope, no distinguishing detail, and does not differentiate it from siblings like get_speech_history_item_by_id or list_text_to_speech_generations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no named alternative (e.g., list_text_to_speech_generations for enumeration). Usage is only faintly implied by the 'Get' verb and the required generation_id.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tool_dependent_agents_routeC
Read-onlyIdempotent

Get Dependent Agents List

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
tool_idYesID of the requested tool.
page_sizeNoHow many documents to return at maximum. Can not exceed 100, defaults to 30.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds no behavioral context such as pagination behavior, what 'dependent agents' means, or any limits, so it contributes nothing beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase that is front-loaded but under-specified rather than concise. It lacks the information an agent needs to understand scope and return behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list endpoint with a required tool_id and pagination, the description does not explain what dependent agents are, what the return contains, or how this differs from knowledge-base dependent-agent tools. With no output schema, this gap is significant, though the annotations cover the safety profile.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with cursor, tool_id, and page_size all documented in the input schema, so the description is not required to carry parameter meaning. The description adds nothing beyond the schema, which is the baseline expectation when coverage is high.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Dependent Agents List' essentially restates the tool name/title with the word 'tool' removed. It does not state that the list is of agents dependent on a specific tool, and it fails to distinguish this tool from siblings like get_knowledge_base_dependent_agents and get_knowledge_base_bulk_dependent_agents_route. This is nearly a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, no prerequisites, and no mention of alternatives. The agent must infer usage entirely from the name and required tool_id parameter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tool_executions_routeC
Read-onlyIdempotent

Get Tool Executions

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
tool_idYesID of the requested tool.
agent_idNoFilter by agent ID.
end_timeNoFilter executions until this Unix timestamp (inclusive).
is_errorNoFilter by error status. If not provided, returns all executions.
branch_idNoFilter by agent branch ID.
page_sizeNoHow many documents to return at maximum. Can not exceed 100, defaults to 30.
start_timeNoFilter executions from this Unix timestamp (inclusive).

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is covered. The description contributes nothing further — no mention of pagination behavior, filtering defaults, or result ordering — despite being a paginated list tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single phrase, but this reflects under-specification rather than efficient concision. There is nothing to front-load because no meaningful content is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter paginated retrieval tool with no output schema, the description leaves filtering semantics, pagination, and result shape entirely to the schema. It does not contradict anything, but it is substantially incomplete as an instruction source.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (tool_id, cursor, agent_id, branch_id, is_error, start_time, end_time, page_size) is already documented in the schema. Baseline 3 applies; the description adds no semantic detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Tool Executions' merely restates the tool name/title, providing no scope, no distinction from siblings such as get_tools_route, get_tool_route, or get_test_invocation_route. It names a verb and a resource, but adds nothing an agent could not infer from the identifier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no mention of prerequisites or alternatives. The agent must infer from the name alone whether this is for listing execution history of a specific tool versus the sibling retrieval tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tool_routeD
Read-onlyIdempotent

Get Tool

ParametersJSON Schema
NameRequiredDescriptionDefault
tool_idYesID of the requested tool.
environmentNoEnvironment whose values are used when the MCP server URL, headers, or auth connection reference environment variables. Mirrors the environment a conversation would run in; defaults to production.

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and non-destructive behavior, so safety is covered, but the description adds zero context of its own. It says nothing about what a tool route is, whether missing IDs error, or what the environment parameter affects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is under-specification rather than conciseness; there is no waste but also no content to front-load. The fragment fails to earn its place as a usable definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description would need to explain the returned route shape, yet it says nothing. In an ecosystem with hundreds of siblings, a stub description leaves the agent unable to confirm it has selected the right tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters (tool_id and environment), so the schema already carries the semantic load. The description adds nothing beyond it, which is the baseline case when the schema documents itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Tool" is a bare restatement of the tool name and title, giving no verb-resource specificity beyond the literal words. It does not distinguish this route from siblings like get_tools_route or get_tool_executions_route. This is the classic tautology case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no pointer to alternatives such as get_tools_route (list) versus this route (single fetch by ID). An agent must infer everything from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tools_routeC
Read-onlyIdempotent

Get Tools

ParametersJSON Schema
NameRequiredDescriptionDefault
typesNoIf present, the endpoint will return only tools of the given types.
cursorNoUsed for fetching next page. Cursor is returned in the response.
searchNoIf specified, the endpoint returns only tools whose names start with this string.
sort_byNoThe field to sort the results by
page_sizeNoHow many documents to return at maximum. Can not exceed 100, defaults to 30.
sort_directionNoThe direction to sort the results
created_by_user_idNoFilter tools by creator user ID. When set, only tools created by this user are returned. Takes precedence over show_only_owned_documents. Use '@me' to refer to the authenticated user.
show_only_owned_documentsNoIf set to true, the endpoint will return only tools owned by you (and not shared from somebody else). Deprecated: use created_by_user_id instead.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered structurally. The description adds nothing beyond that - no mention of pagination behavior, filtering semantics, or that results are scoped to tools in the workspace. With annotations doing the work, a low score is warranted for zero added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is not conciseness but under-specification; the description is too short to carry any signal. This mirrors the 'Process' calibration case, where extreme brevity was scored as inadequate rather than efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an eight-parameter, paginated, filterable listing endpoint, the description conveys none of the operational context an agent needs (what a 'tool' means here, that results are paginated, how cursoring works). The 100%-covered schema mitigates some of this, and the read-only annotations cover safety, but the description itself is essentially absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all eight parameters (types, cursor, search, sort_by, page_size, sort_direction, created_by_user_id, show_only_owned_documents) are already documented in the schema. The description adds no parameter meaning whatsoever, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Tools" restates the tool's own name and title almost verbatim, adding no scope, verb nuance, or resource qualification. It does not distinguish this from siblings such as get_tool_route (singular), list_mcp_server_tools_route, or get_tool_dependent_agents_route. An agent gets no more information from the description than from the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling to prefer. With many similarly named tool-listing/getter endpoints in the sibling set, some routing guidance is exactly what is missing. The description offers no guidance at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_transcript_by_idC
Read-onlyIdempotent

Get Transcript By Id

ParametersJSON Schema
NameRequiredDescriptionDefault
transcription_idYesThe unique ID of the transcript to retrieve

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive and openWorld, so the safety profile is covered without the description. The description adds nothing beyond those annotations — no note on permissions, ID format, or what happens for an unknown/invalid ID. With annotations present the bar is lower, but zero contribution warrants a low score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no wasted words, so it is not bloated. However, its brevity reflects under-specification rather than disciplined conciseness — it carries no substance to front-load.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with full annotation coverage the description still omits what is returned (no output schema exists) and how a transcript relates to the many sibling transcript/dubbing tools. An agent can call it, but only by inferring everything from the name.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter transcription_id is clearly documented in the schema itself, so the baseline of 3 applies. The description contributes nothing extra about the parameter (e.g., where the ID comes from or whether it is a workspace-scoped identifier).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is the tool name re-spaced with no added information: 'Get Transcript By Id'. It does not distinguish this from siblings like dubbing_transcript_get or delete_transcript_by_id, nor does it say what a transcript is in this system or what 'Id' refers to. A verb+resource is technically present, but only because the name was restated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus the many other transcript-related siblings (dubbing_transcript_get, get_dubbing_transcripts, transcribe). No prerequisites, no scoping, no alternatives named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_user_infoC
Read-onlyIdempotent

Get User Info

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is fully covered. The description adds nothing beyond that — no note on auth requirements, scope of the returned user, or sourcing — so it contributes no behavioral value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single three-word phrase with no filler, so nothing is wasted, but it is also so minimal that it functions as a label rather than a specification. Brevity here reflects under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters, no output schema, and full annotation coverage, the structural burden on the description is modest — yet the one thing that matters, whose user info is returned, is left entirely ambiguous. For a tool sitting among dozens of sibling get_* calls, that omission leaves the definition incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is nothing for the description to clarify about arguments; the only residual gap is what 'user' implicitly refers to, which the schema cannot express.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get User Info" merely restates the tool name and title; it names no distinguishing scope or resource beyond the obvious. It gives no hint whether this returns the authenticated user, a workspace member, or something else, so it cannot be told apart from siblings like get_user_subscription_info or get_workspace_members.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative-tool guidance at all. The agent receives no signal about which user is fetched or when this is preferable to the many other get_* tools nearby.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_user_subscription_infoC
Read-onlyIdempotent

Get User Subscription Info

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is fully covered without the description. The description adds nothing beyond that — no mention of whose subscription is returned, permission requirements, or scoping behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single phrase contains no waste, but it is under-specification rather than conciseness — it says nothing beyond the title. It is not front-loaded with any usable information an agent could act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param read tool with annotations covering the safety profile, the description is the only place to convey scope (e.g. current user vs. workspace, what subscription attributes are returned). None of that is present, so the definition is too thin to guide correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema is trivially complete, so the baseline is 4. There is no parameter meaning for the description to add or omit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is an exact restatement of the tool name and title ('Get User Subscription Info'), which the rubric classifies as tautology rather than a specific purpose statement. It names a verb and resource, but does not distinguish this tool from closely named siblings such as get_user_info or get_settings_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this tool, what prerequisites or authentication context it requires, or how it differs from get_user_info. The agent is left to infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_user_voices_v2D
Read-onlyIdempotent

Get Voices V2

ParametersJSON Schema
NameRequiredDescriptionDefault
ageNoAge used for filtering, based on the voice's 'age' label.
sortNoWhich field to sort by, one of 'created_at_unix' or 'name'. 'created_at_unix' may not be available for older voices.
accentNoAccent used for filtering, based on the voice's 'accent' label.
genderNoGender used for filtering, based on the voice's 'gender' label.
searchNoSearch term to filter voices by. Searches in name, description, labels, category.
categoryNoCategory of the voice to filter by. One of 'premade', 'cloned', 'generated', 'professional'
languageNoLanguages used for filtering, based on the voice's 'language' label. Voices matching any of the given languages are returned.
page_sizeNoHow many voices to return at maximum. Can not exceed 100, defaults to 10. Page 0 may include more voices due to default voices being included.
use_casesNoUse cases used for filtering, based on the voice's 'use_case' label. Voices matching any of the given use cases are returned.
voice_idsNoVoice IDs to lookup by. Maximum 100 voice IDs.
voice_typeNoType of the voice to filter by. One of 'personal', 'community', 'default', 'workspace', 'non-default', 'non-community', 'saved'. 'non-default' is equal to all but 'default'. 'non-community' is equal to 'personal' and 'workspace' combined (excludes library copies). 'saved' is equal to non-default, bu
high_qualityNoWhen true, only return studio-quality voices (those whose category is 'high_quality').
collection_idNoCollection ID to filter voices by.
sort_directionNoWhich direction to sort the voices in. 'asc' or 'desc'.
next_page_tokenNoThe next page token to use for pagination. Returned from the previous request. Use this in combination with the has_more flag for reliable pagination.
fine_tuning_stateNoState of the voice's fine tuning to filter by. Applicable only to professional voices clones. One of 'draft', 'not_verified', 'not_started', 'queued', 'fine_tuning', 'fine_tuned', 'failed', 'delayed'
include_total_countNoWhether to include the total count of voices found in the response. NOTE: The total_count value is a live snapshot and may change between requests as users create, modify, or delete voices. For pagination, rely on the has_more flag instead. Only enable this when you actually need the total count (e.
include_custom_ratesNoWhether to include voices that have a custom sharing rate. Defaults to including them.
include_live_moderatedNoWhether to include voices that have live moderation enabled. Defaults to including them.
min_notice_period_daysNoFilter to voices whose sharing notice period is at least the given number of days.

TDQS

D1.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety and read-only nature is covered. However, the description adds no behavioral context such as pagination quirks, default inclusions, or differences from V1, which is a significant gap for a tool with 20 parameters and no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only three words and lacks any structure or front-loading of key information. It is not overly verbose, but it is severely under-specified for a tool of this complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool has 20 parameters, complex filtering options, and no output schema, the description is completely inadequate. It provides no information about return values, pagination, default behaviors, or distinctions from similar tools, leaving the agent to rely entirely on the schema and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 20 parameters in detail. The description adds no parameter meaning beyond what the schema provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Voices V2' is a tautology that merely restates the tool name. It does not specify a verb or resource meaning beyond the name, and it does not distinguish this tool from siblings like get_voices, get_library_voices, or get_voice_by_id.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get_voices or get_library_voices. The agent must infer usage from the name alone, which is insufficient given the large number of voice-related siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_version_metadata_routeC
Read-onlyIdempotent

Get Agent Version Metadata

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.
version_idYesUnique identifier for the version.

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is fully covered. The description adds no behavioral context beyond that — nothing about what metadata is returned, whether versions can be missing, or any auth/rate-limit notes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded phrase with no wasted words. It is appropriately terse for a simple getter, though it is so minimal that it borders on under-specification rather than genuine conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity read-only tool with two fully documented required params and rich annotations, no output schema, and no nested objects, the description is minimally sufficient. It omits any hint of the returned metadata's shape, but the tool's simplicity makes that a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both agent_id and version_id documented in the schema itself. The description adds no meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb ('Get') and a resource ('Agent Version Metadata'), and adds slight scope over the bare name by specifying 'Agent Version'. However, it largely restates the tool name and says nothing about what the metadata contains or how it differs from siblings like get_resource_metadata or get_agent_route. It is adequate but thin.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, nor any mention of prerequisites (e.g., needing a valid version_id from agent creation). The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_video_generationD
Read-onlyIdempotent

Get Video Generation

ParametersJSON Schema
NameRequiredDescriptionDefault
generation_idYes

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, covering the basic safety profile. The description adds nothing beyond this, offering no insight into return behavior, error cases, or async/polling needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but it is under-specified rather than appropriately concise. It provides no front-loaded functional detail and fails to earn its place as a useful tool description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read tool with annotations covering safety, the description is still missing essential context: it does not differentiate from list_video_generations, explain the generation_id parameter, or hint at output. The title-only description leaves important gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for the single required parameter generation_id, and the description does not explain what the parameter represents, its format, or where to obtain it. The description fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Video Generation' simply restates the tool name and title without adding any specificity. It does not distinguish the tool from sibling list_video_generations or clarify that it retrieves a single generation by ID.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like list_video_generations. The agent receives no context about prerequisites beyond the required generation_id parameter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_voice_accentsD
Read-onlyIdempotent

Get Voice Accents

ParametersJSON Schema
NameRequiredDescriptionDefault
languageNoIf provided, only accents for this language code are returned.
model_idNoIf provided, returns the accents available for this model. Defaults to the most complete accent list when omitted.

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint, openWorldHint, idempotentHint, destructiveHint=false) already cover the safety profile, so the description doesn't need to restate that. However, it adds nothing about return format, filtering behavior beyond the schema, or limitations. It contributes no behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short and front-loaded, but it is under-specified rather than concise. It fails to convey any useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a retrieval tool with two optional filtering parameters and no output schema, the description is completely inadequate. It leaves the agent unable to determine what the tool does or when to use it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents both optional parameters (language and model_id). The description adds no parameter meaning; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Voice Accents' merely restates the tool name with no additional detail about what a 'voice accent' is, what it returns, or how it relates to voice tooling. It is effectively a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to call this tool versus siblings like get_voices, get_models, or get_library_voices. An agent has no guidance on context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_voice_by_idC
Read-onlyIdempotent

Get Voice

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
with_settingsNoThis parameter is now deprecated. It is ignored and will be removed in a future version.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the description carries a lower burden — but it adds nothing at all on top of them. It does not mention that this fetches a single voice by ID (versus listing), nor what happens with the deprecated with_settings parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is not conciseness but under-specification — there is essentially no content to be front-loaded or structured. Nothing is wasted, but nothing is communicated either.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-resource lookup with no output schema, the description should at least indicate that it returns one voice's details given an ID. With two parameters, one deprecated, and a crowded sibling namespace, the definition is too thin to guide correct selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3; the voice_id description even points at the listing endpoint and the deprecated parameter is explicitly explained in the schema. The tool description itself contributes no additional parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Voice' merely restates the tool name in slightly different form and gives no scope, return content, or distinguishing detail. It does not differentiate this tool from near siblings such as get_voices, get_library_voices, get_user_voices_v2, get_voice_settings, or get_similar_voices_for_speaker.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many other voice-retrieval tools in the sibling list. No prerequisites, no exclusions, no routing hints are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_voicesC
Read-onlyIdempotent

List Voices

ParametersJSON Schema
NameRequiredDescriptionDefault
show_legacyNoIf set to true, legacy premade voices will be included in responses from /v1/voices

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is fully covered by structured data. The description adds nothing beyond that strip: no mention of pagination, result ordering, scope of the listing, or the legacy-voice behavior controlled by show_legacy.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

'List Voices' is two words with no structural waste, but it is under-specified rather than genuinely concise. There is no front-loaded scope or qualifying information, so the brevity comes at the cost of usefulness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only listing tool whose annotations cover safety and whose single parameter is fully documented, the description is still too thin: it does not say what set of voices is returned, which matters given sibling tools like get_library_voices and get_user_voices_v2. No output schema exists, but the listing behavior itself is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is only one optional parameter, show_legacy, which the schema documents explicitly. The description adds no further meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ('List') and resource ('Voices'), so the basic action is unambiguous. However, it offers no differentiation from the many sibling voice tools (get_voice_by_id, get_library_voices, get_user_voices_v2, get_similar_library_voices, get_voice_accents), leaving scope unclear. It is essentially a restatement of the title 'Get Voices'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus the numerous alternative voice-listing tools. The agent must infer from the name alone whether this returns workspace voices, library voices, or something else. No exclusions or prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_voice_settingsC
Read-onlyIdempotent

Get Voice Settings

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe, repeatable read. The description adds nothing beyond that, but with annotations covering the safety profile, a 3 is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short and front-loaded, but it is arguably under-specified rather than concise; still, it contains no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a read-only tool with no output schema and a mere two-word description, the definition is incomplete: it does not clarify what settings are returned, whether there are permission requirements, or how it relates to the many sibling voice tools. An agent would need to open the schema and guess the rest.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter 'voice_id' is fully documented in the schema, including a link to list voices. The description adds no parameter meaning, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Voice Settings' is a tautology: it merely restates the tool name with no additional specificity about what 'voice settings' encompasses (stability, similarity, style, etc.) or how it differs from siblings like get_voice_settings_default or edit_voice_settings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as get_voice_settings_default (for default settings) or edit_voice_settings (for mutation). The description gives no context or exclusions, leaving the agent to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_voice_settings_defaultB
Read-onlyIdempotent

Get Default Voice Settings.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description adds no behavior beyond that — it does not say whether defaults are workspace-scoped, user-scoped, or what they contain — so it contributes little.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no waste, and the outcome is front-loaded. It is tight, though it is essentially the title restated rather than an addition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read tool with annotations covering safety, the description is minimally adequate. With no output schema, it would help to know what the returned defaults represent or how they relate to get_voice_settings, which is the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to clarify; baseline for a parameterless tool is 4. Schema coverage is reported at 100% and the empty object schema is self-consistent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get Default Voice Settings'), so the operation is unambiguous. However, it does not differentiate itself from close siblings such as get_voice_settings or edit_voice_settings, leaving the agent to guess whether this returns a static default template or defaults for a specific voice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites, and no reference to the sibling get_voice_settings that appears to return per-voice settings. The agent must infer context entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_whatsapp_accountC
Read-onlyIdempotent

Get Whatsapp Account

ParametersJSON Schema
NameRequiredDescriptionDefault
phone_number_idYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description contributes nothing beyond that – no indication of what is returned, whether the phone_number_id must exist, or any error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but this is under-specification rather than conciseness – the single fragment carries no useful information and leaves the agent with nothing to act on beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with full annotation coverage this is near the minimum, but the description still fails to clarify the resource being fetched or the parameter's expected form, leaving the definition incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description adds no meaning for the sole parameter phone_number_id. Its name is somewhat self-describing, but neither schema nor description states the expected format (e.g., numeric ID vs E.164) or lookup behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Whatsapp Account' essentially restates the tool name and title with no additional specificity. It does not distinguish this single-account fetch from the sibling list_whatsapp_accounts or explain what 'an account' contains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus list_whatsapp_accounts (or update_whatsapp_account / delete_whatsapp_account). The agent must infer usage purely from the name, with no conditions or alternatives stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workspace_audit_logsC
Read-onlyIdempotent

Get Workspace Audit Logs

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of entries per page
cursorNoCursor for the next page (from previous response)
actor_uidNoFilter by actor user ID
class_nameNoFilter by OCSF event class name (e.g. Account Change)
activity_nameNoFilter by audit activity name (e.g. Subscription Creation)
time_to_unix_msNoOnly include entries at or before this time (ms since epoch)
time_from_unix_msNoOnly include entries at or after this time (ms since epoch)

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=true, covering the safety profile. The description adds nothing beyond the name – no mention of pagination, event formats, or that it returns audit entries. With annotations carrying the safety burden, the description contributes almost no behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is extremely short and front-loaded, but that brevity comes from under-specification rather than efficient conciseness. It fails to earn its place because it provides no information beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter read tool with pagination and filtering, the description is completely inadequate. It omits any mention of pagination, filtering, return format, or the existence of cursor-based traversal – all critical for correct invocation. No output schema exists, so the description should carry this burden and fails to do so.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all seven parameters (limit, cursor, actor_uid, class_name, activity_name, time_from_unix_ms, time_to_unix_ms) are fully documented in the schema. The description adds nothing beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Workspace Audit Logs' restates the tool name with no added specificity. It states a verb (Get) and resource (Workspace Audit Logs), so it's not fully tautological, but it gives no scope, filtering, or format detail to distinguish it from siblings like get_workspace_members or get_workspace_batch_calls.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool, no conditions, no exclusions, and no references to alternatives. An agent has no basis for choosing this over other get_workspace_* tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workspace_batch_callsC
Read-onlyIdempotent

Get All Batch Calls For A Workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
agent_idNoFilter batch calls to a single agent.
last_docNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so safety is covered structurally. The description adds nothing beyond that — no mention of pagination (implied by limit/last_doc), default limits, or ordering. With annotations doing the heavy lifting, the description still contributes no extra behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded sentence with no waste, which is structurally fine, but the brevity comes from under-specification rather than disciplined concision — there is simply very little content to organize.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a listing tool with three parameters (two undocumented) and no output schema, the description should at minimum describe pagination behavior and what the returned batch calls represent. None of that is present, leaving the agent to guess how to page through results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%: agent_id is documented in the schema, while limit and last_doc are bare. The description says nothing about these parameters, and last_doc in particular reads as a pagination cursor that an agent would need explained; the description does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb and resource ('Get All Batch Calls For A Workspace'), which distinguishes it at a high level from singular get_batch_call and from mutating siblings like create_batch_call or cancel_batch_call. However, it never explains the workspace-scoping or batch-specific framing beyond restating the name, so an agent must infer the distinction from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this rather than get_batch_call (single batch call) or any of the other batch-call siblings. Nothing about prerequisites, workspace context, or result scope is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workspace_membersC
Read-onlyIdempotent

Get Workspace Members

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that: no note on pagination, result size, or whether the listing is workspace-scoped or filtered by the caller's permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no wasted words, but that brevity comes from under-specification rather than tight editing, so it only reaches the minimum viable level.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is low-complexity (no parameters, read-only, annotations cover safety), so a short description is defensible. Still, with no output schema, the agent has no idea what a member entry looks like or whether results are paginated, leaving a real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies; there is nothing for the description to clarify beyond what the empty schema already communicates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Workspace Members' merely restates the tool name and annotation title verbatim, adding no scope, no distinction from siblings such as get_workspace_service_accounts or get_assignable_users_route, and no indication of what a 'member' record contains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of alternatives (e.g. get_assignable_users_route for users eligible to be assigned, or get_workspace_service_accounts for non-human principals), and no preconditions. The agent must infer everything from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workspace_service_accountsC
Read-onlyIdempotent

Get Workspace Service Accounts

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered by structured data. The description adds nothing beyond that — no mention of pagination, filters, or what the listing contains (e.g., whether API keys or secrets are exposed). It neither contradicts nor enriches the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single short phrase is trivially concise and front-loaded, but it is under-specified rather than efficiently worded — there is no content beyond the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no parameters, the description is the only source of information about the return value, yet it says nothing about what the service-account listing includes. For a workspace-scoped list tool it is too thin to let an agent confirm the call produces the data it needs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which is the baseline-4 case per the rubric. The empty schema leaves nothing for the description to compensate for, though it also offers no extra semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name ('Get Workspace Service Accounts'), with no elaboration on what a service account is or what is returned. It gives no differentiation from related siblings such as get_service_account_api_keys_route or create_service_account. This is tautology rather than genuine purpose definition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites, and no routing to alternatives like get_service_account_api_keys_route for keys. The agent must infer usage purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workspace_webhooks_routeC
Read-onlyIdempotent

List Workspace Webhooks

ParametersJSON Schema
NameRequiredDescriptionDefault
include_usagesNoWhether to include active usages of the webhook, only usable by admins

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive behavior, so safety is covered structurally. The description adds nothing further — no note about the admin-only nature of include_usages, pagination, or return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words, front-loaded with the verb, zero waste. It is terse to the point of being minimal but nothing is extraneous.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with annotations covering the safety profile and full schema coverage on the sole parameter, this is adequate. Missing only minor context such as the admin restriction (which lives in the schema) or any result overview.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single parameter (include_usages, admin-only) is fully documented in the schema itself. The description adds no parameter meaning beyond that, which fits the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('Workspace Webhooks'), which cleanly distinguishes it from the create_/delete_/edit_workspace_webhook_route siblings. Clear but does not explicitly reference those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no when-to-use context, no mention of the sibling create/delete/edit webhook tools, and no prerequisites or exclusions. The agent must infer usage from the verb alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handle_exotel_outbound_callD

Handle An Outbound Call Via Exotel

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes
to_numberYes
agent_phone_number_idYes
telephony_call_configNo
conversation_initiation_client_dataNo

TDQS

D1.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate this is a non-read-only, open-world, non-idempotent operation, implying it initiates an external call with side effects. However, the description adds nothing beyond that: it doesn't disclose authentication requirements, rate limits, cost implications, or what happens to the call after handling begins. For a tool that triggers real outbound telephony, the behavioral disclosure is near-zero, though the annotations do carry the basic safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single nine-word sentence that merely restates the tool name. While it is not verbose, it is under-specified rather than concise — every word duplicates the name without adding information. There is no front-loaded context, purpose, or usage detail to earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters (3 required), nested objects, no output schema, and 0% schema description coverage, the description is completely inadequate. It provides no purpose elaboration, no parameter guidance, no behavioral context, and no differentiation from sibling call-handling tools. An agent would have to guess at the tool's function and all parameter semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so five parameters — including a nested telephony_call_config object and a complex conversation_initiation_client_data structure — are undocumented in both schema and description. The description provides no information about any parameter, not even which are required or what agent_id, agent_phone_number_id, and to_number mean. This is a severe gap for a tool with nested objects and zero parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name and title verbatim: "Handle An Outbound Call Via Exotel." It does not clarify the specific action — whether it initiates a call, manages an active call's lifecycle, or routes an existing call — leaving the agent to infer the purpose from the name alone. Siblings like handle_twilio_outbound_call and handle_sip_trunk_outbound_call show this is one of several provider-specific call entry points, but the description offers no distinguishing detail about what Exotel handling entails.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The presence of handle_twilio_outbound_call, handle_sip_trunk_outbound_call, and whatsapp_outbound_call as siblings makes provider selection the critical decision, yet the description gives no criteria for choosing Exotel over these alternatives. There is no when-to-use or when-not-to-use context whatsoever.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handle_sip_trunk_outbound_callC

Handle An Outbound Call Via Sip Trunk

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes
to_numberYes
agent_phone_number_idYes
telephony_call_configNo
conversation_initiation_client_dataNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, openWorldHint=true, so the agent knows this is a non-idempotent, real-world action. The description adds nothing beyond that — it doesn't say a live outbound phone call is placed, that credentials/telephony setup are required, or that duplicate invocations may create separate calls. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, so nothing wastes space, but it is under-specified rather than truly concise — the brevity comes from omitting information the agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a mutation tool with nested optional config objects, no output schema, and open-world side effects, yet the description supplies none of the setup context, side-effect framing, or return expectations an agent needs. For this complexity level the definition is largely incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Top-level schema description coverage is 0% for 5 params, and the description names none of them (agent_id, agent_phone_number_id, to_number). Some nested optional fields carry their own descriptions, but the description under review contributes no parameter meaning at all, leaving the required params' roles to be inferred from their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name almost verbatim ('Handle An Outbound Call Via Sip Trunk' vs handle_sip_trunk_outbound_call). It does carry the 'via SIP trunk' qualifier that distinguishes it from the sibling handle_twilio_outbound_call and handle_exotel_outbound_call, so it is not purely a tautology, but the verb 'handle' is vague about what actually happens (a real call is placed).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites (e.g., agent and phone number must be configured), and no reference to the sibling provider alternatives that would tell an agent when to pick SIP trunk over Twilio or Exotel. Usage is only implied by the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handle_twilio_outbound_callD

Handle An Outbound Call Via Twilio

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes
to_numberYes
agent_phone_number_idYes
telephony_call_configNo
call_recording_enabledNo
conversation_initiation_client_dataNo

TDQS

D1.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false, so the safety profile is partly covered. The description adds nothing beyond that: no note that this actually dials a real external number with cost/side effects, no auth requirements, no async/return behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence, but this is under-specification rather than conciseness — nothing here earns its place because it only echoes the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutation tool with nested objects, a required-parameter set, zero schema description coverage, and no output schema, the description is wholly inadequate. It gives an agent no basis for calling it correctly or safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 6 parameters, three of which are required (agent_id, agent_phone_number_id, to_number), and the description supplies no parameter meaning at all. The nested telephony_call_config and conversation_initiation_client_data objects are entirely undocumented in the prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Handle An Outbound Call Via Twilio" mostly restates the tool name/title. It names the provider (Twilio) and the action class (outbound call), but "Handle" is vague and it does nothing to distinguish itself from the sibling handle_exotel_outbound_call, handle_sip_trunk_outbound_call, or whatsapp_outbound_call beyond the provider word already in the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus any alternative, no prerequisites (e.g. that agent_id and a provisioned Twilio number must exist), and no mention of the sibling outbound-call tools that an agent could easily confuse it with.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

invite_userD

Invite User

ParametersJSON Schema
NameRequiredDescriptionDefault
emailYesThe email of the customer
group_idsNo
seat_typeNoSeat types for workspace members.
usage_limitNo
workspace_permissionNo

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate a non-read-only, open-world, non-idempotent operation, but the description doesn't elaborate on side effects, required permissions, or what happens on duplicate invites. For a mutation tool with annotations, some behavioral context is provided by annotations, but description adds nothing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, overly terse phrase that fails to convey necessary information. It is concise but at the cost of being unhelpful; structure is minimal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters, 1 required, 40% schema coverage, and no output schema, the description is inadequate. It does not explain the purpose, usage, or parameters, leaving the agent with insufficient information to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 40% (only email and seat_type have descriptions), so the description must compensate for undocumented parameters like group_ids, usage_limit, and workspace_permission. It provides no parameter meaning at all, leaving significant gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Invite User' merely restates the tool name without adding a verb-resource clarification or scope. It doesn't differentiate from sibling invite tools like invite_users_bulk or delete_invite. Tautological.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus invite_users_bulk or other invite-related tools. No context or alternatives mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

invite_users_bulkD

Invite Multiple Users

ParametersJSON Schema
NameRequiredDescriptionDefault
emailsYesThe email of the customer
group_idsNo
seat_typeNoSeat types for workspace members.
usage_limitNo

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is covered elsewhere. The description adds nothing behavioral — no mention that invitations send emails, whether seats are consumed, or whether re-inviting is safe.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but that brevity is under-specification rather than efficiency — the one sentence conveys no more than the tool name already does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-parameter, non-idempotent, open-world write action with no output schema and incomplete parameter documentation, the description provides none of the context needed: no seat/role implications, no group assignment semantics, no result information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50% (group_ids and usage_limit are undocumented, and emails has a mismatched description saying 'the email of the customer' for an array). The description compensates with zero parameter information, leaving half the inputs unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Invite Multiple Users' largely restates the tool name invite_users_bulk, adding only the word 'Multiple'. It does not distinguish this tool from the sibling invite_user, which is exactly the distinction an agent needs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given, and the obvious alternative (invite_user, for single invitations) is never mentioned. The agent must infer bulk-vs-single selection purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_agent_conversation_tickets_routeC
Read-onlyIdempotent

List Agent Conversation Tickets

ParametersJSON Schema
NameRequiredDescriptionDefault
labelNoFilter tickets by an exact label.
cursorNoUsed for fetching next page. Cursor is returned in the response.
statusNoFilter tickets by status.
sourcesNoFilter tickets by how they were raised (qa, agent, manual). Repeat the parameter to filter by multiple sources.
agent_idYes
page_sizeNoHow many agent conversation tickets to return. Can not exceed 100.
issue_typeNoFilter clusters by issue type.
owner_user_idNoFilter tickets by creator. Use 'agent' for agent-raised tickets.
conversation_idNoFilter tickets by conversation id.
assignee_user_idNoFilter tickets by assignee. Use 'unassigned' for tickets with no assignee.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, non-destructive, openWorld. The description adds nothing beyond that — no pagination behavior, no note on the required agent_id scoping, no hints about result cardinality. With annotations carrying the safety profile the bar is lower, but a one-line restatement adds no behavioral context at all.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no padding and is front-loaded, but for a 10-parameter list tool this brevity is under-specification rather than useful conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter listing tool with no output schema and a nearly identically named workspace-level sibling, the description omits the differentiation and scoping context an agent needs to pick correctly. The rich schema partially compensates, but the description itself is inadequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 90%, so the schema itself documents nearly all 10 parameters (filters, cursor, page_size limit of 100). The description adds no parameter meaning, so baseline 3 applies since structured data does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Agent Conversation Tickets' merely restates the tool name and title rather than adding a verb+resource statement with scope. It does not distinguish this tool from the sibling list_workspace_conversation_tickets_route, which an agent could easily confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus list_workspace_conversation_tickets_route or get_agent_conversation_ticket_route. Nothing states the filtering prerequisites (e.g., that agent_id is required) or when-not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_assetsC
Read-onlyIdempotent

List Assets

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoToken from a previous response's `next_cursor`. Omit to fetch the first page.
searchNoOptional free-text search filter over asset names.
page_sizeNoNumber of assets to return.

TDQS

C2.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile and read-only nature are clear. With annotations covering this, the bar is lower; the description adds nothing beyond the name, but it does not contradict the annotations. A baseline 3 is appropriate given annotations carry the behavioral load.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At two words, it is concise, but it is under-specified rather than efficient. It fails to front-load any meaningful information, and 'List Assets' earns no place beyond the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description does not explain return values, pagination behavior, or the shape of the asset list. Given the tool has three parameters and pagination (cursor/next_cursor), the description is inadequate for an agent to understand what to expect beyond what the schema provides.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already defines cursor, search, and page_size with clear descriptions. The description adds no parameter information, but per the rule, when coverage is high, the baseline is 3 even with no param info in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Assets' is a tautology that restates the tool name without adding any scope, verb nuance, or resource specificity. It does not distinguish from siblings like get_asset, upload_asset, or delete_asset_endpoint beyond the obvious list vs retrieve/single distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus alternatives such as get_asset or upload_asset. There is no context, no prerequisites, and no exclusions mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_auth_connectionsB
Read-onlyIdempotent

Get Workspace Auth Connections

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds only the workspace-level scope and does not disclose return format, pagination, or auth requirements, but it is consistent with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short phrase, front-loaded and free of filler. It could be slightly more informative without becoming bloated, but it is appropriately concise for a zero-parameter list tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only list tool with rich annotations but no output schema, the description is minimally sufficient to invoke correctly. It omits whether all connections are returned or how results are structured, though the low complexity limits the harm.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters and schema description coverage is 100%, so the baseline is 4. The description has nothing additional to explain for parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase names the operation and resource ('Get Workspace Auth Connections'), so an agent can infer it retrieves the workspace's auth connections. However, it is terse and largely restates the name/title, without distinguishing it from the create/update/delete auth-connection siblings beyond the verb.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or alternatives are provided. An agent can infer it is the list operation from the name, but the description itself gives no selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_available_llmsC
Read-onlyIdempotent

List Available Llms

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds no additional behavioral context such as the source of the list, caching, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase, which is concise but under-specified. It is not verbose, but it also lacks any front-loaded value beyond the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that lists available LLMs in a large API surface, the description is incomplete. It does not explain what 'available' means, whether it lists all models or only those for the current user, or how it relates to get_models. With no output schema, the description should do more to set expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema coverage is 100%. Per the rules, zero params yields a baseline of 4; there are no parameters to describe, and the description adds nothing but also has no gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Available Llms' identifies a verb (list) and resource (available LLMs). However, it is a tautological restatement of the tool name and title, offering no scope, filtering, or differentiation from siblings like get_models.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get_models or public_get_available_languages. The description provides no context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_chat_response_tests_routeC
Read-onlyIdempotent

List Agent Response Tests

ParametersJSON Schema
NameRequiredDescriptionDefault
typesNoIf present, the endpoint will return only tests/folders of the given types.
cursorNoUsed for fetching next page. Cursor is returned in the response.
searchNoSearch query to filter tests and folders by name.
page_sizeNoHow many Tests to return at maximum. Can not exceed 100, defaults to 30.
sort_modeNoSort mode for listing tests. Use 'folders_first' to place folders before tests.
sharing_modeNoFilter test visibility. Use `shared_with_me` to return only tests/folders shared with the current user that they did not create.
include_foldersNoDeprecated. Use the `types` query param and include `folder` instead.
parent_folder_idNoFilter by parent folder ID. Use 'root' to get items in the root folder.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that—no mention of pagination via cursor, default page size, or filtering behavior—so it contributes no behavioral context of its own.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single phrase is front-loaded and free of waste, but its brevity reflects under-specification rather than efficiency. There is no structure beyond the bare title-like statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a listing endpoint with 8 optional filtering/pagination parameters and no output schema, the description should at least hint at filtering, pagination, and result shape. It leaves the agent to reconstruct all of that from the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and all 8 parameters (types, cursor, search, page_size, sort_mode, sharing_mode, include_folders, parent_folder_id) are documented in the schema itself. The description adds no parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Agent Response Tests' names a verb and resource, so the purpose is identifiable, but it essentially restates the tool name and adds nothing about scope, filtering, or pagination. It does not distinguish this listing endpoint from sibling readers like get_agent_response_test_route or get_agent_response_tests_summaries_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get_agent_response_tests_summaries_route, list_test_invocations_route, or the folder/test read endpoints. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_conversation_tags_routeC
Read-onlyIdempotent

List Conversation Tags

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
page_sizeNoHow many conversation tags to return. Can not exceed 100.

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered without description help. The description adds no behavioral context such as pagination semantics or whether tags are workspace-scoped, which is a gap but the annotations carry the load.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with zero waste. It is appropriately terse for a simple list operation, though it is so short that it barely counts as a description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, zero-required-parameter listing tool whose annotations and schema are rich, the description is minimally sufficient. It omits anything about the return contents or workspace scope, but with no output schema and a fully documented input schema, the burden on the description is lower.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema fully documents cursor and page_size including the 100-item cap and pagination meaning. The description adds nothing about parameters, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Conversation Tags' restates the tool name with a clear verb (list) and resource (conversation tags), so the purpose is understandable. However, it offers no differentiation from siblings like get_conversation_tag_route or list_* tools, so it lands at vague-but-adequate rather than specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no named alternatives, and no conditions distinguishing this listing tool from the get/assign/unassign tag routes. Nothing beyond the name implies the usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_crawl_jobs_routeB
Read-onlyIdempotent

List Ongoing And Recent Crawl Jobs Created By A User

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
page_sizeNoHow many documents to return at maximum. Can not exceed 100, defaults to 30.
include_job_idsNoIds of additional crawl jobs to retrieve

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is fully covered. The description adds only the temporal scope ('ongoing and recent'), which is useful but thin, and says nothing about result ordering, pagination limits, or whether all users' jobs are visible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single lean sentence with no wasted words and the resource front-loaded. The title-case phrasing and the vague 'Created By A User' clause are the only minor blemishes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, 3-parameter list tool with full schema coverage and no output schema, the description is minimally adequate. It leaves unclear whose crawl jobs are returned and how ongoing vs recent are defined, which an agent must infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with cursor, page_size (max 100, default 30) and include_job_ids all documented in the schema. The description adds no parameter-level meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb 'List' with resource 'Crawl Jobs' and a scope qualifier 'ongoing and recent'. An agent can distinguish it from get_crawl_job_route (single item) and cancel/create_crawl_job_route by inference, though the description never names those siblings. The trailing phrase 'Created By A User' is ambiguous about which user, slightly muddying the scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance, no mention of alternatives such as get_crawl_job_route for a single job, and no note about when to prefer this listing. Usage is only implied by the verb 'List'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_dubsC
Read-onlyIdempotent

List Dubs

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
order_byNoThe field to use for ordering results from this query.
page_sizeNoHow many dubs to return at maximum. Can not exceed 200, defaults to 100.
dubbing_modelsNoFilter by dubbing model generation.
dubbing_statusNoWhat state the dub is currently in.
order_directionNoThe order direction to use for results from this query.
creation_sourcesNoFilter by dubbing creation source.
dubbing_statusesNoFilter by dubbing status.
filter_by_creatorNoFilters who created the resources being listed, whether it was the user running the request or someone else that shared the resource with them.
target_language_codesNoFilter by target language code.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only, idempotent, non-destructive, and open-world safety profile. The description adds no behavioral context beyond those annotations, such as pagination behavior, filtering defaults, or access scope.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The phrase is front-loaded and has no wasted words, but it is under-specified rather than appropriately concise. It omits necessary information for a tool with ten optional filtering and pagination parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a listing tool with ten parameters and no output schema, the description is nearly empty. Annotations and the schema carry much of the burden, but the description still provides no usage context, filters overview, or scoping information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all ten parameters are documented directly in the schema. The description adds no parameter meaning, which leaves it at the baseline score for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Dubs' only restates the tool name and title. It does not clarify what a dub is, what is returned, or how this tool differs from siblings such as dubbing_project_list or get_dubbing_transcripts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is provided, and no alternatives are mentioned. The agent receives no indication of when listing dubs is preferable to retrieving dubbing resources or transcripts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_environment_variablesC
Read-onlyIdempotent

List Environment Variables

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNoFilter by variable type
labelNoFilter by exact label match
cursorNoPagination cursor from previous response
page_sizeNoNumber of items to return (1-100)
environmentNoFilter to only return variables that have this environment. When specified, the values dict in the response will only contain this environment.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered by structured fields. The description contributes nothing beyond that - it does not mention pagination behavior, the cursor/page_size interaction, or the special environment-scoping side effect on the response. No contradiction, but no added value either.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single phrase is not bloated, but it is under-specified rather than concise - there is simply nothing to trim because nothing was said. It is front-loaded by default but carries no payload.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter, paginated list tool with no output schema, the description should explain pagination via cursor, the meaning of page_size bounds, and the filter semantics. None of that appears, leaving the agent reliant entirely on the schema and annotations to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: all five parameters (type, label, cursor, page_size, environment) are documented in the schema, including the non-obvious note that 'environment' narrows the response values dict. Per the rubric, a fully-covered schema sets the baseline at 3; the description adds nothing on top of it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name and title, adding no distinguishing information. It does not differentiate from siblings such as get_environment_variable, create_environment_variable, or update_environment_variable, nor does it state scope (workspace/environment). This is the textbook 'tautology' case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus get_environment_variable (single fetch) or list_auth_connections. No prerequisites, no exclusions, no context about pagination-driven iteration. The agent must infer all routing decisions from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_image_generationsC
Read-onlyIdempotent

List Image Generations

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoPagination cursor: the `next_cursor` value of the previous page's response. Omit it for the first page.
statusNoOnly return generations with this lifecycle status.
model_idNoOnly return generations of this model.
page_sizeNoHow many generations to return per page.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, covering safety. The description adds nothing about pagination behavior, filtering semantics, or what 'Generations' encompasses. With annotations providing the safety profile, the description is purely tautological.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but not in a helpful way – it's under-specified rather than concise. A single tautological phrase fails to earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, and the description doesn't explain return values, pagination mechanics, or how to interpret the list. For a listing tool with four optional filters, more context about filtering and pagination would be expected.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (cursor, status, model_id, page_size) are fully documented in the schema. The description adds no parameter meaning, which is acceptable when schema coverage is high.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Image Generations' restates the tool name and title with no additional specificity. It does indicate a list operation on image generations, but doesn't distinguish from siblings like list_video_generations or list_text_to_speech_generations beyond the noun.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. There are multiple listing tools for different asset types, and the description provides no routing hints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_mcp_servers_routeC
Read-onlyIdempotent

List Mcp Servers

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered structurally. The description adds nothing beyond that — no mention of pagination, filtering, or what the returned server objects contain.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short fragment with no wasted words, which is technically concise. However, it is under-specified rather than economical in a useful way — nothing is front-loaded because nothing beyond the name is conveyed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless read-only list operation whose safety profile is fully carried by annotations, the description is minimally adequate. It still omits any hint of return contents, ordering, or pagination, but the structured fields cover most of what an agent needs to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema is empty, so there is nothing for the description to clarify. Baseline 4 applies for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"List Mcp Servers" is essentially the tool name restated with no added specificity about scope, format, or coverage. It does not distinguish this tool from the adjacent list_mcp_server_tools_route, which an agent could easily confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or sibling reference. The agent is left to infer that this enumerates MCP servers rather than the tools inside an MCP server.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_mcp_server_tools_routeC
Read-onlyIdempotent

List Mcp Server Tools

ParametersJSON Schema
NameRequiredDescriptionDefault
environmentNoEnvironment whose values are used when the MCP server URL, headers, or auth connection reference environment variables. Mirrors the environment a conversation would run in; defaults to production.
mcp_server_idYesID of the MCP Server.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, destructiveHint=false, covering the safety profile. The description adds nothing beyond that — no note on whether it contacts the remote MCP server (openWorld implies network reach, worth spelling out), whether results are cached, or how failures surface.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short and front-loaded, but it is a bare fragment with no structure because there is nothing substantive to structure. It is not verbose, yet it is under-specified rather than concise in the helpful sense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter listing tool with no output schema, the description should at least say what the listing contains (e.g., tool names/schemas exposed by that MCP server) and any scope limits. Without that, an agent cannot confidently predict the result or distinguish this from related MCP listing/config tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with mcp_server_id and the environment parameter (production default, env-var resolution) fully documented in the schema. The description contributes no additional parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is essentially a restatement of the tool name/title ('List Mcp Server Tools' → list_mcp_server_tools_route). It identifies verb and resource only implicitly and provides no scope information, such as whether it lists all tools registered on one MCP server or discovered tools from the server's connection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no mention of alternatives. Given siblings such as list_mcp_servers_route, get_mcp_route, get_mcp_tool_config_override_route, and get_tools_route, the agent gets no help deciding when this listing is the right call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_phone_numbers_routeC
Read-onlyIdempotent

List Phone Numbers

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idNoFilter by assigned agent ID
providerNoFilter by telephony provider
branch_idNoFilter by assigned branch ID

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so safety semantics are covered structurally. The description adds nothing beyond that — no mention of pagination, result limits, default ordering, or scope of the listing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single phrase is brief but not concise in the useful sense — it is underspecified rather than economical, providing no information an agent could act on beyond the name itself.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list endpoint with no output schema, the description should at least hint at return shape, filtering behavior, or pagination. Neither the description nor an output schema tells the agent what comes back, leaving a real gap despite the well-covered input schema and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each of the three optional filters (agent_id, provider, branch_id) documented inline, so the baseline of 3 applies. The description contributes no additional parameter meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name ('List Phone Numbers' = list_phone_numbers_route), which the rubric classifies as tautology territory. It does convey a read/list operation, but adds no scope, resource detail, or differentiation from siblings like get_phone_number_route or create_phone_number_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus get_phone_number_route (single lookup) or create/update/delete variants, nor any mention that the three filters can be combined. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_procedures_routeC
Read-onlyIdempotent

List Procedures

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesAgent ID to get the procedure draft from
branch_idYesBranch ID to get the procedure draft from
agent_version_idNoThe agent version ID to retrieve the procedure for.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, covering the safety profile. The description adds nothing beyond this — no mention of return shape, pagination, or scoping by agent/branch that would help an agent reason about behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two words, which is technically concise but reflects under-specification rather than efficiency. No front-loaded detail or useful structure is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter read tool with no output schema and many similarly named procedure/branch siblings, the description should clarify scope and distinguish it from related tools. It does neither, leaving routing ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each of the three parameters documented in the schema itself, so the baseline is 3. The description contributes no additional parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Procedures' is essentially a restatement of the tool name list_procedures_route, offering no verb+resource distinction beyond the name itself. It does not differentiate from close siblings such as get_procedure_route, get_procedure_draft_route, or compile_procedures_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, nor any conditions or prerequisites mentioned. The agent is left to infer usage entirely from the name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_sip_messagesC
Read-onlyIdempotent

Get Sip Messages For A Phone Number

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
page_sizeNo
phone_number_idYesThe phone number ID. This is returned when a phone number is imported.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true. The description adds no behavioral context beyond the annotations, such as pagination behavior, rate limits, or authentication needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It is appropriately sized for a simple list/get endpoint, though it could be more structured by mentioning pagination or filtering.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is low complexity and annotations carry the safety profile, but the description does not explain pagination or the missing page_size semantics, and there is no output schema to document return values. It is minimum viable but leaves gaps an agent might need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%, with page_size undocumented in the schema. The description implies the required phone_number_id but adds no clarifying meaning for it, cursor, or page_size beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Get Sip Messages') and scopes it to a phone number, which differentiates it from the conversation-level sibling get_conversation_sip_messages. It stops short of explicitly naming or excluding that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get_conversation_sip_messages or list_phone_numbers_route. The only usage context is implicit in the phrase 'For A Phone Number.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_speech_enginesC
Read-onlyIdempotent

List Speech Engines

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
searchNoSearch term to filter Speech Engines by name
sort_byNoThe field to sort the results by
page_sizeNoHow many Speech Engines to return at maximum. Can not exceed 100, defaults to 30.
sort_directionNoThe direction to sort the results

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that — no mention of pagination behavior via the cursor, default page size, or the search/sort capabilities that would help an agent call it correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but this is under-specification rather than conciseness. There is no front-loaded scope or filter information, so brevity comes at the cost of usefulness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter paginated list tool with no output schema, the description does not explain what a Speech Engine is, how results are ordered by default, or how pagination terminates. Annotations cover safety only, leaving the functional picture incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters (cursor, search, sort_by, page_size, sort_direction) are fully documented in the schema itself. Per the baseline rule, a 3 is appropriate since the description contributes no additional parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name/title ('List Speech Engines'), which is the definition of a tautology. It does not distinguish this tool from siblings like get_speech_engine, create_speech_engine, update_speech_engine, or delete_speech_engine.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no alternatives named, no prerequisites or conditions. The agent is left to infer everything from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_test_invocations_routeC
Read-onlyIdempotent

List Test Invocations

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
searchNoSearch query to filter tests and folders by name.
agent_idNoFilter by agent ID
page_sizeNoHow many Tests to return at maximum. Can not exceed 100, defaults to 30.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered without the description. The description adds nothing beyond that – no pagination behavior, no result ordering, no filtering semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but that brevity reflects under-specification rather than efficient front-loading. The single phrase carries no information an agent could not already get from the name, so it does not earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter paginated listing tool with no output schema, the description should at least explain what a 'test invocation' represents and how pagination/filtering combine. None of that is present; only the safety hints from annotations fill part of the gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: cursor, search, agent_id, and page_size are all documented in the input schema, including the page_size cap of 100 and default of 30. The description contributes no additional parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Test Invocations' is essentially a restatement of the tool name (list_test_invocations_route) and its title. It names a verb and resource but adds no scope, distinguishing detail, or sibling differentiation against the many other list_*/get_*_route tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get_test_invocation_route (single invocation) or other listing endpoints. No prerequisites, filters, or exclusions are mentioned; usage must be inferred entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_text_to_speech_generationsC
Read-onlyIdempotent

List Speech Generations

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoPagination cursor: the `next_cursor` value of the previous page's response. Omit it for the first page.
statusNoOnly return generations with this lifecycle status.
model_idNoOnly return generations of this model.
page_sizeNoHow many generations to return per page.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, open-world, and non-destructive behavior, so the description does not need to repeat safety traits. However, it adds no operational context such as pagination mechanics, ordering, or result shape beyond what 'List' implies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The phrase is extremely short and front-loaded, but it is under-specified rather than concisely informative. Like the LOW calibration example, this is minimal wording that omits necessary context instead of earning its brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should help explain what is returned and how pagination works. Given four optional filter/pagination parameters and no required parameters, the definition is too sparse for an agent to understand result behavior confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 even with no parameter details in the description. The description adds no additional meaning for cursor, status, model_id, or page_size beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Speech Generations' essentially restates the tool name/title without adding scope, filters, or return details. It does not distinguish this listing tool from siblings such as list_image_generations, list_video_generations, or get_text_to_speech_generation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives like get_text_to_speech_generation or other list_* generation tools. It gives no filtering context beyond the schema, and no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_video_generationsC
Read-onlyIdempotent

List Video Generations

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoPagination cursor: the `next_cursor` value of the previous page's response. Omit it for the first page.
statusNoOnly return generations with this lifecycle status.
model_idNoOnly return generations of this model.
page_sizeNoHow many generations to return per page.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds no behavioral context beyond annotations, such as pagination behavior, filtering semantics, or return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a three-word phrase that merely restates the title. Its brevity is not helpful because it conveys no new information and under-specifies the tool's use.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity and rich annotations/schema, the description still fails to provide any selection guidance among many sibling list tools. It omits pagination, filtering details, and any distinguishing context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema fully documents cursor, status, model_id, and page_size. The description adds no parameter syntax or meaning beyond what the schema already provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description exactly restates the tool name and title ('List Video Generations') with no additional verb-object detail or scope. It does not distinguish this tool from sibling list tools such as list_image_generations or list_text_to_speech_generations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of alternatives, and no prerequisites. An agent gets no signal about when to choose this over get_video_generation or other list endpoints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_whatsapp_accountsC
Read-onlyIdempotent

List Whatsapp Accounts

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idNoFilter by assigned agent ID

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that — no indication of what accounts are listed, whether results are paginated, or whether the agent filter changes scope.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words with no filler, but the brevity stems from under-specification rather than tight writing. There is no front-loaded detail because there is no detail at all.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with one optional, fully documented parameter, the description is nearly adequate but still omits what is returned and how the optional agent_id filter behaves. Annotations and schema carry the load, but the description itself is hollow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single agent_id parameter is documented in the schema as 'Filter by assigned agent ID'. The description adds no parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List Whatsapp Accounts' is essentially a restatement of the tool name and title, adding no distinguishing information. It does convey a verb and resource, but it does nothing to separate this from sibling tools like get_whatsapp_account or delete_whatsapp_account.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of alternatives (e.g., get_whatsapp_account for a single account), and no stated prerequisites or scope. The agent must infer all routing from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_workspace_conversation_tickets_routeC
Read-onlyIdempotent

List Workspace Conversation Tickets

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
statusNoFilter tickets by status.
page_sizeNoHow many agent conversation tickets to return. Can not exceed 100.
assignee_user_idNoFilter tickets by assignee. Use 'unassigned' for tickets with no assignee.

TDQS

C2.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond them — no pagination behavior, no rate-limit or auth notes — so it neither helps nor harms.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence with zero waste, but also zero substance — it is concise to the point of being uninformative rather than front-loading useful constraints.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list endpoint whose sibling set includes both 'list_agent_conversation_tickets_route' and 'list_conversation_tags_route', the description provides no scope or disambiguation. With no output schema and no added behavioral detail, an agent lacks what it needs to confidently pick and use this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each of the four parameters (cursor, status, page_size, assignee_user_id) explained in the schema itself. The description contributes no parameter semantics, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name almost verbatim ('List Workspace Conversation Tickets' vs 'list_workspace_conversation_tickets_route'). It conveys a verb (list) and a resource (conversation tickets), but adds no scope, no differentiation from the sibling 'list_agent_conversation_tickets_route', and nothing about filtering behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of the near-identical sibling 'list_agent_conversation_tickets_route', and no indication of when to prefer the filtered queries. The agent must guess which list endpoint to select.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

merge_branch_into_targetC

Merge A Branch Into A Target Branch

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoForce source branch changes onto the target, overriding timestamp-based conflict resolution
agent_idYesThe id of an agent. This is returned on agent creation.
source_branch_idYesUnique identifier for the source branch to merge from.
target_branch_idYesThe ID of the target branch to merge into.
archive_source_branchNoWhether to archive the source branch after merging

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true. The description adds no behavioral context beyond what annotations provide, and does not describe merge semantics, conflict resolution, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single phrase that is front-loaded but under-specified for a multi-parameter mutation tool. It is essentially a duplicate of the title/name and does not earn its place by adding information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given five parameters, two of which control merge behavior (force and archive_source_branch), and the absence of an output schema, the description is completely inadequate. It omits any explanation of merge semantics, conflict resolution, or the effect of optional flags.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are fully documented in the schema. The description provides no additional parameter meaning, but baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name ('Merge A Branch Into A Target Branch') with no additional scope or differentiation from siblings like merge_preview_route or rebase_branch_onto_main.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as merge_preview_route or rebase_branch_onto_main. The description provides no context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

merge_preview_routeC
Read-onlyIdempotent

Preview Merged Configuration

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoWhen true, source branch changes always win conflicts regardless of timestamps
agent_idYesThe id of an agent. This is returned on agent creation.
source_branch_idYesUnique identifier for the source branch to merge from.
target_branch_idYesThe ID of the target branch to merge into.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered structurally. The description adds nothing on top: no statement that the merge is not applied, no conflict-reporting behavior, no permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words with no waste, but this is under-specification rather than effective conciseness — the brevity comes at the cost of all routing and behavioral information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no description of what a "preview" returns (conflict list, diff, merged config shape), an agent cannot tell what to expect from the call. Annotations cover safety but not the preview semantics that make this tool distinct.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with 4 documented parameters (including the force conflict-resolution flag), so the schema carries the parameter semantics. The description adds no meaning beyond the schema, which is the baseline-3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Preview Merged Configuration" restates the tool name (merge_preview_route) and the annotation title almost verbatim, adding no verb+resource specificity beyond what the identifier already conveys. It does not distinguish this preview from siblings like merge_branch_into_target or rebase_preview_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance whatsoever. The critical distinction an agent needs — preview versus the actual merge (merge_branch_into_target), or versus rebase_preview_route — is left entirely unstated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

migrate_segmentsB

Move Segments Between Speakers Spends ElevenLabs credits. Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
dubbing_idYesID of the dubbing project.
speaker_idYes
segment_idsYes

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (non-read-only, non-idempotent, open-world), so the bar is lower. The description adds two genuinely non-obvious traits: it consumes ElevenLabs credits and is deprecated upstream, both of which materially affect whether an agent should call it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short and front-loaded, but the two clauses run together without punctuation ('...Spends ElevenLabs credits. Deprecated upstream.'), making the credit cost read as part of the action name. It is terse without being well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-required-param mutation tool with no output schema, the description covers cost and deprecation and annotations cover safety. However, it omits what happens to the moved segments and leaves two parameters undocumented, so it is only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (only dubbing_id documented), and the description says nothing about speaker_id or segment_ids. With two undocumented required parameters, the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: moving segments between speakers, which is distinct from siblings like dubbing_target_transcript_segment_update. The phrase is clear enough that an agent can identify the operation, though it does not explicitly name which sibling to prefer for segment edits.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage signal is 'Deprecated upstream,' which warns against use but names no replacement tool and gives no condition for when this is still appropriate. There is no guidance on when to use this versus the many other segment/dubbing tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

patch_agent_settings_routeD

Patches An Agent Settings

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
tagsNo
agent_idYesThe id of an agent. This is returned on agent creation.
workflowNo
branch_idNoThe ID of the branch to use
proceduresNo
platform_settingsNo
conversation_configNo
version_descriptionNo
enable_versioning_if_not_enabledNoDeprecated: all agents are versioned. This parameter is ignored.

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is partly covered. The description adds nothing beyond this: it doesn't explain that PATCH semantics mean only supplied fields change, whether omitted fields are preserved, or what versioning side effects (version_description) imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase, but this is under-specification rather than conciseness. There is no front-loaded information beyond a restated name, so brevity buys nothing here.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter mutation tool with nested objects, no output schema, and low schema coverage, the description is completely inadequate. An agent cannot determine what is being patched, how partial updates behave, or what fields matter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 10 parameters and only 30% schema description coverage, the description carries the burden of explaining parameters and contributes nothing. Fields like name, tags, workflow, procedures, platform_settings, and conversation_config are undocumented in both places, and the description does not compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Patches An Agent Settings' is essentially a title-case restatement of the tool name, adding no resource scope, no statement of what settings are affected, and no differentiation from siblings like update_tool_route or create_agent_route. It conveys only that a partial update happens to some agent-related settings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus alternatives (e.g., update_settings_route, update_branch_route, or other agent mutation routes). No prerequisites, no exclusions, no context are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

patch_pronunciation_dictionaryC

Update Pronunciation Dictionary

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoThe name of the pronunciation dictionary, used for identification only.
archivedNoWhether to archive the pronunciation dictionary.
pronunciation_dictionary_idYesThe id of the pronunciation dictionary

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, but the description adds no meaningful behavioral context beyond 'update'. It does not explain patch semantics, whether unspecified fields are preserved, what archiving does, or any auth/rate-limit considerations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only three words and functions as a title restatement rather than a structured explanation. It is concise but under-specified, with no front-loaded usage or behavioral detail to justify its brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with three parameters and no output schema, the description is too sparse. While annotations cover safety hints and the schema covers parameters, the description leaves out patch behavior, intended use, and any operational context an agent would need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for all three parameters, so the schema already documents name, archived, and pronunciation_dictionary_id. The description adds no additional parameter meaning beyond what the schema provides, matching the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb and resource: update a pronunciation dictionary. However, it is essentially a restatement of the tool title and does not distinguish this tool from the similar sibling update_pronunciation_dictionaries or clarify the partial-update patch semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as update_pronunciation_dictionaries, get_pronunciation_dictionary_metadata, or other pronunciation dictionary tools. The description merely says it updates without context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

post_agent_avatar_routeD

Post Agent Avatar

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.
avatar_file_pathNoAn image file to be used as the agent's avatar. Local path. Required for this call.
avatar_file_base64NoBase64 contents for "avatar_file". Use this when the server cannot read your local disk.
avatar_file_filenameNoFilename to send for "avatar_file". Some endpoints infer the audio format from it.

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=false, destructiveHint=false, openWorldHint=true, and idempotentHint=false, so the safety profile is partially covered externally. However, the description adds no behavioral context of its own, such as upload constraints, required permissions, side effects, or how the avatar is stored or replaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but this is under-specification rather than useful conciseness. It is front-loaded only in the sense that there is nothing else to read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool that uploads an avatar file with four parameters and no output schema, the description is completely inadequate. It omits purpose, usage context, behavioral details, and any parameter guidance, leaving the agent dependent entirely on structured fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, so all four parameters are already well documented there. The description adds no meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Post Agent Avatar" merely restates the tool name and title without adding any clarification about what posting an avatar means, what resource it affects, or how it differs from siblings such as post_agent_hold_audio_route. It is a tautological restatement rather than a useful purpose statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool, when not to use it, prerequisites, or alternatives. With dozens of sibling tools, an agent receives no routing information whatsoever.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

post_agent_hold_audio_routeD

Post Agent Hold Audio

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.
hold_audio_file_pathNoAn MP3 or WAV file played on loop to callers waiting in the agent's concurrency wait queue. Maximum size 40 MB, maximum duration 180 seconds. Local path. Required for this call.
hold_audio_file_base64NoBase64 contents for "hold_audio_file". Use this when the server cannot read your local disk.
hold_audio_file_filenameNoFilename to send for "hold_audio_file". Some endpoints infer the audio format from it.

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=false, openWorldHint=true and destructiveHint=false, so the write/side-effect profile is covered by structured data. The description adds nothing about behavior — not what is replaced, whether the previous audio is discarded, auth needs, or size/duration constraints — so beyond the annotations it contributes no context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is short, but that brevity reflects under-specification rather than efficient front-loading. There is no informational content to front-load, so it reads as a bare title rather than a working description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent, non-destructive write tool with four parameters and no output schema, the description provides nothing an agent needs: no statement of what is being set, whether it replaces existing hold audio, or what a successful call produces.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (agent_id, hold_audio_file_path, base64, filename) are already fully documented with constraints such as 40 MB / 180 s limits and the base64 fallback. Per the baseline rule for high coverage, a 3 is appropriate even though the description adds no parameter information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Post Agent Hold Audio' merely rephrases the tool name 'post_agent_hold_audio_route' without adding a verb+resource that clarifies what actually happens (e.g. setting/uploading an agent's hold audio file). It is essentially a tautology, and it gives no hint of how it differs from siblings like post_agent_avatar_route or delete_agent_hold_audio_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, when not to, or which alternatives exist. The sibling list contains several closely related audio/agent routes, and nothing here routes the agent to the right one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

post_conversation_feedback_routeD

Send Conversation Feedback

ParametersJSON Schema
NameRequiredDescriptionDefault
feedbackNo
conversation_idYesThe id of the conversation you're taking the action on.

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish the behavior profile (readOnlyHint=false, idempotentHint=false, openWorldHint=true), so the description is not required to repeat them. However, it adds no behavioral context at all: no statement that this mutates an existing conversation's feedback state, whether repeat calls are allowed, or who may submit feedback. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but the brevity reflects under-specification rather than efficiency — there is no scoping, no alternative routing, and no parameter hints. Four words cannot earn their place when they merely echo the tool title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, an undocumented enum parameter, and only annotation-level behavioral coverage, the description should explain what is being changed and the meaning of the feedback values. It provides none of this, leaving an agent able to guess only the roughest intent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: conversation_id is documented in the schema, but the feedback parameter (enum like/dislike) has no schema description, and the description supplies nothing about it. The description mentions "feedback" only as part of the tool name, so it fails to compensate for the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Send Conversation Feedback" restates the tool name almost verbatim rather than stating what sending feedback accomplishes (e.g., rating a conversation as like/dislike). There is no differentiation from siblings such as get_conversation_summary_route or run_conversation_evaluations, so an agent gains nothing beyond the name itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance whatsoever on when to call this tool, what prerequisites exist (e.g., the conversation must already exist), or when an alternative like run_conversation_evaluations would be more appropriate. The agent must guess entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

post_knowledge_base_bulk_delete_routeC

Bulk Delete Knowledge Base Documents

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoIf set to true, documents or folders will be deleted regardless of whether they are used by any agents and will be removed from the dependent agents. For non-empty folders, this will also delete all child documents and folders.
document_idsYesThe ids of documents or folders from the knowledge base.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, openWorldHint=true and destructiveHint=false, so the write/safety profile is mostly carried structurally. The description adds nothing on top: it does not warn that non-forced deletes fail or skip documents in use, that forcing detaches them from dependent agents, or that folder deletion cascades to children — all of which the schema's 'force' text implies but the description never surfaces.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short fragment that wastes no words, but it is essentially a restatement of the tool name/title rather than a front-loaded summary with content. It is under-specified rather than genuinely concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive bulk operation touching a knowledge base used by agents, the description should at minimum mention the force/non-force behavior and the cascading effect on dependent agents. With no output schema and no usage context, the description leaves the agent without the operational picture it needs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%; the 'force' parameter and 'document_ids' are both fully documented in the schema. The description contributes no additional parameter meaning, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Bulk Delete') and resource ('Knowledge Base Documents'), so the operation is unambiguous. However, it does nothing to distinguish itself from the sibling 'delete_knowledge_base_document' (singular delete) or 'post_knowledge_base_bulk_move_route', leaving the agent to infer that 'bulk' is the differentiating trait.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this bulk tool versus the single-document delete sibling, nor any prerequisite (e.g., checking dependent agents first with get_knowledge_base_bulk_dependent_agents_route). The agent gets only the purpose statement and must infer usage entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

post_knowledge_base_bulk_move_routeC

Bulk Move Entities To Folder

ParametersJSON Schema
NameRequiredDescriptionDefault
move_toNo
document_idsYesThe ids of documents or folders from the knowledge base.

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is largely covered. The description adds nothing beyond that — no note on what happens to existing folders, whether the move is reversible, or required permissions. No contradiction, but no added value either.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short phrase — no wasted words and front-loaded — but it reads as a title fragment rather than a purposeful description, and its brevity comes from under-specification rather than disciplined economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, non-idempotent bulk operation with no output schema and only 50% parameter coverage, the description is too thin. It omits the single-vs-bulk distinction, the meaning and format of move_to, and any behavioral context the annotations don't already carry.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: document_ids is documented in the schema, but move_to has no description in either place. The phrase 'To Folder' hints that move_to is the destination folder, offering marginal compensation, but its format (ID vs. name vs. null behavior) is left entirely to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Bulk Move Entities To Folder' gives a verb (move), a scope (bulk), and a destination (folder), which is more than a bare tautology. However, 'Entities' is vague — the schema reveals these are documents or folders — and it never distinguishes itself from the sibling post_knowledge_base_move_route (single move). Adequate but with clear ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, stated prerequisites, or reference to alternatives despite the existence of a single-move sibling (post_knowledge_base_move_route) and a bulk-delete sibling. The agent must infer the appropriate context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

post_knowledge_base_move_routeC

Move Entity To Folder

ParametersJSON Schema
NameRequiredDescriptionDefault
move_toNo
document_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the agent knows this is a mutating, non-idempotent operation. The description adds no behavioral context beyond that, such as whether the move is reversible, what happens if the target folder is missing, or what permissions are needed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loaded, but it is a fragment that omits essential context rather than a tight, complete statement. It avoids waste but is under-specified for a mutation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that this is a mutating tool, has 50% schema coverage, no output schema, and a closely related bulk-move sibling, the description is not complete. It fails to specify what entity is being moved, what a folder is in this context, or any return/error behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: document_id is documented in the schema, while move_to has no schema description. The description's 'To Folder' phrase weakly maps to move_to as the destination, but it does not clarify the expected format or valid values, so it only minimally compensates for the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a verb (Move) and a destination (To Folder), but the resource is stated vaguely as 'Entity' rather than specifying a knowledge base document or folder. It does not distinguish this singular move from the sibling bulk move tool post_knowledge_base_bulk_move_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites, and no exclusions. The description only restates the basic action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

public_create_orderC

Create Order Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
sandboxNoWhen true, creates a sandbox order that auto-progresses without producer intervention.

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-read-only, non-idempotent, open-world, non-destructive behavior, so the safety profile is covered. The description adds one genuinely new trait — that the call spends ElevenLabs credits — which is useful cost context, but nothing about reversibility, permissions, or what an order becomes after creation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no padding, which is appropriately sized for a one-parameter tool. However the second sentence is grammatically broken ('Create Order Spends ElevenLabs credits'), which slightly impedes parsing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter tool with a fully described optional parameter and no output schema, the bar is low and the annotation set covers safety. What remains missing is the tool's place in the order lifecycle relative to its numerous order siblings, which is the main unresolved gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single optional parameter and schema coverage is 100%, so the sandbox flag is fully documented in the schema. The description adds no meaning about what a sandbox order is beyond what the schema already says, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb+resource 'Create Order' is stated plainly, so the basic purpose is recoverable. But it is not distinguished from sibling order tools such as public_submit_order, public_upsert_order_item, or public_create/update variants, so an agent cannot tell which order-creation path to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, what prerequisites exist, or which alternative order endpoints exist. The only contextual hint is the credit-cost note, which is a consequence rather than guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

public_get_available_languagesC
Read-onlyIdempotent

Get Available Languages

ParametersJSON Schema
NameRequiredDescriptionDefault
order_item_kindYesThe kind of order item.

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare this as read-only, idempotent, and non-destructive, so the safety profile is covered. However, the description adds nothing beyond that—no information about return format, pagination, or what 'available' means in this context (e.g., availability per order_item_kind). With annotations carrying the safety burden, the lack of any additional behavioral context is a gap for a tool in a rich ecosystem.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise—a single short phrase. It's front-loaded and wastes no words, but that conciseness comes at the cost of under-specification. For a tool with no other context, this level of brevity is efficient but not optimally informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool has no output schema, the description should at least hint at what is returned (e.g., a list of language codes or names). It doesn't. Combined with vague purpose and no usage guidance, the description is incomplete for an agent trying to understand the tool's role in a complex API with many language-related siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% description coverage—the single parameter 'order_item_kind' is clearly documented with an enum and a description. The tool description adds no parameter semantics beyond the schema, which is acceptable when the schema fully documents the parameter. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is just the title restated: 'Get Available Languages' with no additional detail. It doesn't specify what kind of languages are returned (dubbing, subtitle, transcription), nor does it distinguish this tool from the vast array of language-related siblings like dubbing_language_list, get_voice_accents, or add_language.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidelines are provided. It doesn't state when to use this tool versus alternatives like dubbing_language_list or how it relates to order items. The agent receives no context on the tool's intended scenario.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

public_get_media_infoC
Read-onlyIdempotent

Get Media Info

ParametersJSON Schema
NameRequiredDescriptionDefault
media_idYesThe ID of the media file.
order_idYesThe ID of the order.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is covered externally. The description adds nothing beyond that — no note on auth requirements, whether the caller must own the order, or what happens for missing media.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but that shortness is under-specification rather than economy — a three-word fragment with no verb object detail or scoping. Nothing is front-loaded because there is nothing to load.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool requiring two IDs and returning unspecified output with no output schema, the description should at least indicate what media info is returned and how it ties to the order. The public_* sibling family implies an order-scoped lookup, but the description never states or confirms this.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both media_id and order_id documented inline, so the baseline of 3 applies. The description contributes no additional meaning about the two required IDs or their relationship.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Media Info" essentially restates the tool name with no added specificity — it does not say which media, in what context, or how it differs from siblings such as get_resource_metadata or public_get_order_deliverables. A reader learns nothing beyond the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this tool, when not to, or what alternatives exist (e.g., other public_get_* order endpoints). The description provides zero routing information for an agent choosing among hundreds of siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

public_get_orderC
Read-onlyIdempotent

Get Order

ParametersJSON Schema
NameRequiredDescriptionDefault
order_idYesThe ID of the order.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing on top of that: no note about auth/scope (public endpoint), no statement of what an order record contains, and no error behavior if the order_id is unknown.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is not conciseness here, it is under-specification; nothing is front-loaded because there is nothing beyond the title. The single phrase earns its place only as a label, not as usable instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must carry the burden of explaining what is returned, and it does not. For an openWorld public API endpoint with no other documentation, an agent has no way to know what a retrieved 'order' includes or how to handle missing records.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single required order_id parameter, so the schema fully documents the input. The description adds no additional semantics (e.g., ID format, scope of the lookup), so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Order" names a verb and a resource, but it is essentially a restatement of the tool name (public_get_order) with no added specificity. It does not distinguish this tool from close siblings such as public_get_order_deliverables, public_list_orders, or public_get_media_info, so an agent gets no signal about what is actually retrieved.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no reference to any alternative. Among many get_* and public_* siblings, the agent must guess when this is the right call versus public_list_orders or public_get_order_deliverables.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

public_get_order_deliverablesC
Read-onlyIdempotent

Get Order Deliverables

ParametersJSON Schema
NameRequiredDescriptionDefault
order_idYesThe ID of the order.

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond the title, such as return format, pagination, authentication, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase with no wasted words, but it is essentially the title restated and does not earn its place by adding useful information. It is concise yet under-specified for an agent trying to select the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read operation with rich annotations and full schema coverage, the description is minimally adequate for invocation. However, with no output schema, it provides no context about what 'deliverables' are returned, which weakens tool selection and result interpretation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the only parameter order_id is documented in the schema as 'The ID of the order.' The description adds no further meaning beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name/title as 'Get Order Deliverables' and adds no scope or distinguishing detail. It is a tautological summary of the resource rather than an explanation of what the tool returns or how it differs from siblings like public_get_order or public_list_orders.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, no prerequisites, and no exclusions. It leaves the agent to infer usage entirely from the name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

public_list_ordersC
Read-onlyIdempotent

List Orders

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetNoNumber of orders to skip for pagination.
statusNoFilter orders by one or more statuses.
end_dateNoFilter orders created on or before this date.
page_sizeNoMaximum number of orders to return per page.
start_dateNoFilter orders created on or after this date.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered structurally. The description adds nothing beyond that: no pagination behavior, no default ordering, no note on whether results are scoped to the caller's public account.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words are technically concise and front-loaded, but here brevity reflects under-specification rather than economy of language. The description is too small to earn its place as documentation for a five-parameter filtered list tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter listing endpoint with pagination and date-range filters and no output schema, the description says nothing about filtering, paging, date formats, or the public-scope semantics implied by the name. Only the annotations cover part of the picture, leaving the description materially incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with five documented parameters (offset, status, start_date, end_date, page_size), so the schema carries the full semantic load. Per the high-coverage baseline, a 3 is appropriate even though the description contributes no parameter information at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"List Orders" essentially restates the tool name public_list_orders and the annotation title "Public List Orders" without adding scope, filters, or the public/order-domain framing. It is a verb+resource, but nothing distinguishes it from siblings like public_get_order, public_get_order_deliverables, or list_assets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of the public-facing nature of the endpoint, and no pointer to alternatives such as public_get_order for a single order. Nothing is said about when this tool is appropriate versus the other order endpoints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

public_register_mediaC

Register Media Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
order_idYesThe ID of the order to which this media will be attached.
media_urlNoA URL to fetch the media file from. Mutually exclusive with media.
media_pathNoThe media file to upload. Mutually exclusive with media_url. Local path.
media_base64NoBase64 contents for "media". Use this when the server cannot read your local disk.
media_filenameNoFilename to send for "media". Some endpoints infer the audio format from it.
declared_languageYesThe language code of the media content (e.g. 'en', 'es-ES'). Must be a supported source language for some order item kind.
media_url_filenameNoThe filename for URL-sourced media (e.g. 'example.mp4'). Required when using media_url.
media_url_content_typeNoThe MIME type for URL-sourced media (e.g. 'video/mp4'). Required when using media_url.

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is covered. The description adds one genuinely useful trait not present in annotations – that the call spends ElevenLabs credits – but says nothing about the non-idempotent retry hazard or what registration actually mutates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence, front-loaded with the action and the cost consequence, so there is no waste. However, the brevity crosses into under-specification for a non-idempotent 8-parameter mutation rather than being efficiently concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating tool with 8 parameters, two media-source modes, conditional requirements, and no output schema, the description should at least state that media is attached to an order and what a successful registration yields. It omits all of that, leaving only the credit-cost note.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all 8 parameters, including the three mutually exclusive media source options and the conditional media_url_filename/media_url_content_type requirements, so the baseline of 3 applies. The description contributes no additional parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb+resource pair 'Register Media' names the action but leaves the resource vague – register media to what, and for which container? The schema reveals it attaches media to an order_id, but the description never says so, and it doesn't distinguish itself from siblings like public_upload_order_item or public_get_media_info.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, prerequisites, or named alternatives. The only contextual hint is that it consumes credits, which is a cost note rather than usage routing. The agent has to infer everything from the schema and sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

public_remove_order_itemC
DestructiveIdempotent

Remove Order Item

ParametersJSON Schema
NameRequiredDescriptionDefault
item_idYesThe ID of the order item.
order_idYesThe ID of the order.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false, and openWorldHint=true, so the safety profile is covered by structured data. The description adds nothing beyond that – it does not say what gets destroyed, whether removal is reversible, or what happens to the parent order, so it earns only minimal credit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words is front-loaded and free of filler, but this is under-specification rather than genuine conciseness. The brevity leaves no room for the context an agent needs on a destructive mutation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world mutation with no output schema, the description should at minimum clarify what is deleted and any ordering/consistency implications for the parent order. Instead it provides nothing beyond a title-like phrase, leaving the definition materially incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both required parameters (order_id, item_id) are documented in the schema itself. The description adds no syntax, format, or relationship detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Remove Order Item' is essentially the tool name restated, which is a tautology per the rubric. It does name a verb and resource, but offers no scope, no differentiating detail, and nothing to distinguish it from siblings like public_upsert_order_item or public_submit_order.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. An agent gets no signal about when removal is appropriate versus updating the order via public_update_order or upserting items via public_upsert_order_item.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

public_submit_orderB

Submit Order Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
order_idYesThe ID of the order.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is covered. The description adds the important credit-spend side effect, but does not describe return values, prerequisites, or recovery behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence with the action and cost warning front-loaded. It is appropriately sized, though the phrasing is slightly awkward and lacks punctuation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given one required parameter, no output schema, and annotations that cover the operation's safety profile, the description is nearly complete. The credit-spend warning is the key missing contextual detail, but routing guidance relative to sibling order tools is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single required parameter order_id, so the schema already documents the parameter fully. The description adds no parameter-level meaning beyond that baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Submit Order') and adds a side effect ('Spends ElevenLabs credits'). It does not, however, distinguish this tool from sibling order tools such as public_create_order or public_update_order.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use, when-not-to-use, or alternative guidance. The only context is the implied action and the credit-spend warning, which is too thin to route an agent between public_submit_order and public_create_order/public_update_order.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

public_update_orderC

Update Order Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes
order_idYesThe ID of the order.

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds one piece of genuinely new context – that the call consumes ElevenLabs credits – but says nothing about permissions, reversibility of the update, or which fields can change.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, which is good, but the single sentence is grammatically garbled and conflates the action with a cost side-effect rather than front-loading the true purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, non-idempotent tool with a nested request object, no output schema, and only 50% param coverage, the description is too thin. An agent still cannot tell what gets updated, what the response looks like, or what authorization is required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%, and the description adds no parameter meaning at all. It never explains that order_id identifies the target or that request carries the new name, so the nested request object is left entirely to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb (Update) and resource (Order), so it is not a tautology, but the run-on phrasing 'Update Order Spends ElevenLabs credits' blurs the actual action into a cost note. It does nothing to distinguish this from siblings like public_upsert_order_item or public_submit_order.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives, even though close siblings (public_upsert_order_item, public_submit_order, public_remove_order_item) make routing ambiguous. The only guidance given is a credit-cost warning, which is not a usage condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

public_upsert_order_itemC

Upsert Order Item Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes
order_idYesThe ID of the order.

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a real behavioral side effect beyond the annotations – it spends ElevenLabs credits – which the safety annotations (readOnlyHint=false, openWorldHint=true) do not convey. However, it says nothing about idempotency intent despite idempotentHint=false, nor about permissions or which order states allow the operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loaded, but the single sentence is grammatically garbled ('Spends ElevenLabs credits'), which harms readability without adding information efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with a nested object parameter and no output schema, the description is far too thin. It omits what the upsert returns, how the free-form item payload is shaped, and any cost magnitude for the mentioned credit spend.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 50% schema description coverage, the description must compensate but adds no parameter meaning. The nested 'request.item' object (free-form additionalProperties) and 'item_id' are entirely undocumented in both schema and description, leaving the agent guessing at required shape.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb+resource ('upsert order item') is identifiable, though the sentence is malformed ('Spends ElevenLabs credits'). It does not differentiate itself from siblings such as public_update_order, public_create_order, or public_remove_order_item.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no reference to alternative order tools. An agent must infer from the name alone when this upsert applies versus public_update_order or public_remove_order_item.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_agent_knowledge_base_rag_routeC

Query Agent Knowledge Base Rag

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesQuery to run against the agent's knowledge base RAG index.
agent_idYesThe id of an agent. This is returned on agent creation.
branch_idNoThe ID of the branch to use
use_agent_defaultsNoWhen true (the default), retrieval uses the agent's own RAG settings, reproducing exactly what the agent would retrieve. Set to false to retrieve with neutral default RAG settings instead (the agent's embedding model is always kept, since it determines which vector index exists). Useful for auditing
max_documents_lengthNo
max_retrieved_rag_chunks_countNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, but the description supplies none of the context this raises: it never says whether the call has side effects, whether retrieval is deterministic, or whether the agent's config is mutated. Notably the schema's own use_agent_defaults text describes audit behavior that the description omits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loaded, but that brevity is under-specification rather than conciseness. As a single verbatim-restatement sentence it wastes no words but conveys almost nothing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter RAG retrieval tool with an unusual use_agent_defaults audit mode and no output schema, the description should explain return shape and the agent-defaults behavior. It explains nothing, leaving the agent reliant entirely on partial schema text.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% and the two parameters that lack descriptions, max_documents_length and max_retrieved_rag_chunks_count, are undocumented here as well. The description adds zero parameter meaning, but the baseline for moderate schema coverage is 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Query Agent Knowledge Base Rag' is essentially a restatement of the tool name with no verb-object framing or scope. It does not say what querying returns (retrieved chunks? an answer?), nor how it differs from siblings like search_knowledge_base_content_route or get_knowledge_base_content. An agent cannot distinguish its purpose from the other knowledge-base read tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all. The schema's use_agent_defaults parameter hints at an auditing use case, but the description never states when to prefer this tool over search_knowledge_base_content_route or the document/chunk retrieval tools. Nothing steers the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rag_index_statusD

Compute Rag Index.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
documentation_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=false and openWorldHint=true, so the agent knows this is a non-idempotent mutating operation, but the description adds nothing on top of that. It never says whether the call kicks off an indexing job, whether it is synchronous or async, whether it requires the knowledge base to already exist, or what side effects it produces.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words and a period. This is under-specification rather than conciseness — there is no front-loaded scope, no object of the computation, and nothing for the agent to act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A tool with two required parameters, a mutating/non-idempotent annotation profile, and no output schema needs far more than three words. The description leaves the operation's trigger semantics, prerequisites, and return behavior entirely unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: documentation_id is documented, but the model parameter (an enum of two embedding models) has no description anywhere. The description supplies no parameter meaning at all, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Compute Rag Index.' is a near-tautology that restates the tool name in different words and never resolves the ambiguity between the name ('status', implying a read) and the verb ('compute', implying a write/trigger). It gives no basis for choosing it over the many siblings in the same domain (get_rag_indexes, get_rag_index_overview, get_or_create_rag_indexes, delete_rag_index).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, prerequisite, or alternative-tool guidance of any kind. This is especially costly here because at least three siblings cover overlapping RAG-index territory, and an agent has nothing to route on.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rebase_branch_onto_mainC

Rebase A Branch Onto Main

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.
branch_idYesUnique identifier for the source branch to merge from.

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, openWorldHint=true, and idempotentHint=false. The description adds nothing beyond these annotations—it does not explain what a rebase does, what state it requires, or what side effects it may have.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, but it is entirely redundant with the tool name and title, so it does not earn its place in the definition. Conciseness alone does not compensate for lack of substantive content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent, open-world mutation operation like rebasing a branch, the description omits any explanation of prerequisites, consequences, or expected behavior. With no output schema and annotations only covering safety hints, the description is critically incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (agent_id and branch_id) are fully documented in the schema. The description adds no parameter-level information, but with complete schema coverage the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name and title, adding no differentiation from siblings such as rebase_preview_route or merge_branch_into_target. While the name itself is descriptive, the description field provides no additional clarity about scope or behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, nor any mention of prerequisites or when not to use it. The description gives no context for selection among the many sibling branch operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rebase_preview_routeC
Read-onlyIdempotent

Preview Rebased Configuration

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.
branch_idYesUnique identifier for the source branch to merge from.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered structurally. The description adds nothing beyond that – it does not explain what the preview returns, whether it has side effects on branch state, or how it relates to the real rebase. With the lower bar set by annotations, a 2 reflects near-zero added value rather than a contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loaded, but it is under-specified rather than concise: the single phrase conveys no information an agent could act on, so the brevity is a defect, not a virtue.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining what a preview yields, and it does not. Combined with a two-parameter mutation-adjacent branch operation and no usage context, the definition leaves the agent without enough to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both agent_id and branch_id, so the baseline is 3. The description contributes no additional meaning about parameter roles or formats, and in particular does not clarify what 'source branch to merge from' implies in a rebase context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Preview Rebased Configuration" is essentially a restatement of the tool name rebase_preview_route; it names no resource actor beyond 'configuration' and does not distinguish itself from sibling operations such as rebase_branch_onto_main or merge_preview_route. An agent cannot tell the preview's scope, target, or relation to the actual rebase from this text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus rebase_branch_onto_main, merge_preview_route, or create_branch_route. No prerequisites, no indication that this is a dry-run preceding an actual rebase, nothing to route the agent correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

redirect_to_mintlifyD
Read-onlyIdempotent

Redirect To Mintlify

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

D1.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already specify readOnlyHint, openWorldHint, idempotentHint, and destructiveHint, but the description adds no behavioral context beyond restating the title. It does not disclose what a 'redirect' entails (e.g., URL redirect, session redirect) or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short but not concise in the useful sense: it is under-specified and fails to convey essential information. It duplicates the name without adding structure or clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that this is a redirect tool with no parameters and no output schema, the description should at least explain what a 'redirect to Mintlify' operation does. Instead, it provides no context, leaving the agent unable to understand the tool's behavior or purpose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool accepts zero parameters, so parameter semantics are not applicable. The schema has 100% description coverage for an empty object, and the description need not add parameter detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Redirect To Mintlify' is a verbatim restatement of the tool name, offering no distinct verb+resource or scope. It does not clarify what is being redirected, where, or for what purpose, making it a tautology that fails to distinguish the tool from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. The description provides no context, prerequisites, or exclusions, leaving the agent to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refresh_url_document_routeC

Refresh Url Document Content

ParametersJSON Schema
NameRequiredDescriptionDefault
documentation_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, covering the safety profile. The description adds nothing beyond that – it doesn't explain that refreshing re-fetches remote URL content, whether prior content is replaced, or what side effects occur.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a four-word fragment with no wasted words, but it is under-specified rather than genuinely concise – there is no sentence structure or front-loaded guidance for the agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent, non-read-only operation on a knowledge base document, with no output schema, the description omits what refresh means, what it affects, and any result expectations. The annotation set cannot compensate for this missing behavioral context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single documentation_id parameter, and the description adds no additional semantics (e.g., what happens if the ID is stale or invalid). Baseline 3 applies when the schema already documents the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Refresh Url Document Content' essentially restates the tool name and title without stating a distinct verb-object relationship or scope. It doesn't distinguish this from siblings like update_document_route, update_file_document_route, or create_url_document_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus alternatives such as update_document_route or create_url_document_route, nor any precondition (e.g., that the source URL changed). No exclusions or context are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

register_twilio_callC

Register A Twilio Call And Return Twiml

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes
directionNo
to_numberYes
from_numberYes
conversation_initiation_client_dataNo

TDQS

C2.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false and openWorldHint=true, so the mutation/non-idempotent profile is covered. The description adds one genuine piece of context by noting it returns TwiML (useful since there is no output schema), but omits registration side effects, auth requirements, and error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, so it is appropriately sized. It is under-specified rather than verbose, but the one sentence does carry the only two facts offered (registration and TwiML return).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with 0% schema description coverage, a nested client-data object, no output schema, and many competing call-handling siblings, this one line is far too thin to let an agent invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, including a complex nested conversation_initiation_client_data object, and the description supplies no parameter meaning at all. Nothing explains agent_id, direction, from_number, or to_number beyond their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Register) and resource (Twilio Call) and adds the notable side effect of returning TwiML. However, it does nothing to distinguish this from close siblings such as handle_twilio_outbound_call or handle_exotel_outbound_call, so an agent cannot tell which to pick from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no indication of when to use this tool versus the many sibling call-handling tools (handle_twilio_outbound_call, handle_sip_trunk_outbound_call, whatsapp_outbound_call). There are no prerequisites, triggers, or exclusions stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_mcp_server_tool_approval_routeC
DestructiveIdempotent

Delete Mcp Server Tool Approval

ParametersJSON Schema
NameRequiredDescriptionDefault
tool_nameYesName of the MCP tool to remove approval for.
mcp_server_idYesID of the MCP Server.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered structurally. However, the description adds nothing beyond that — no statement about what is destroyed (the approval route only, not the server or tool), required permissions, or effect on in-flight approvals.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is technically brief, but there is no real content to be concise about; the single phrase duplicates the name/title rather than front-loading useful information. This is under-specification rather than effective conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive two-parameter mutation with no output schema, the description should at minimum identify the target resource and the scope of deletion. Rich annotations carry the safety signal, but the text itself leaves behavioral and routing context entirely unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both mcp_server_id and tool_name clearly documented in the schema itself. Per the rubric, a high-coverage schema sets the baseline at 3 even though the description contributes no parameter detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Delete Mcp Server Tool Approval' is essentially a restatement of the tool name and annotation title, providing no additional specification of what deletion means in this context (removing an approval route for a specific tool on a specific MCP server). It is not misleading, but it is tautological rather than clarifying.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus siblings such as add_mcp_server_tool_approval_route, update_mcp_server_approval_policy_route, or remove_mcp_tool_config_override_route. The agent must infer the invocation context from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_mcp_tool_config_override_routeC
DestructiveIdempotent

Delete Mcp Tool Configuration Override

ParametersJSON Schema
NameRequiredDescriptionDefault
tool_nameYesName of the MCP tool to remove config overrides for.
mcp_server_idYesID of the MCP Server.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered. However, the description adds nothing beyond the name — it does not explain what gets deleted, whether the operation can be undone, or what happens after removal. Given the annotations carry significant weight, the description's lack of additional behavioral context is a clear gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no wasted words, but it is so minimal that it fails to front-load any useful context. Conciseness is achieved at the expense of informativeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, idempotent mutation tool with no output schema, the description is essentially just a title. It does not explain the effect of the deletion, required permissions, or side effects. An agent would need to infer behavior entirely from the name and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents both parameters (tool_name and mcp_server_id). The description adds no parameter details, which is the baseline expectation when the schema already does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name as a human-readable phrase ('Delete Mcp Tool Configuration Override'), which identifies the verb and resource but adds no specificity beyond what the name already conveys. It does not distinguish this tool from sibling tools like update_mcp_tool_config_override_route or get_mcp_tool_config_override_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as update_mcp_tool_config_override_route or get_mcp_tool_config_override_route. No context about prerequisites or scenarios is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_memberC

Delete Member From User Group

ParametersJSON Schema
NameRequiredDescriptionDefault
emailYesThe email of the target workspace member.
group_idYesThe ID of the target group.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is partly covered. The description adds nothing beyond the name: no permission requirements, no note on whether the member loses workspace access or only group membership, and no statement about repeat-call behavior despite idempotentHint=false. There is a mild tension between 'Delete' and destructiveHint=false, though removing a group membership while preserving the user account is a defensible reading.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase with zero redundancy, which is efficient, but it is under-specified rather than genuinely concise, and capitalized like a title rather than a usable instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation, the schema covers inputs and the annotations cover the safety flags, so the core information an agent needs is present. Gaps remain around permissions, failure modes (member not in group), and side effects, which the description does not address.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both group_id and email are documented in the schema as the target group ID and target member email. The description adds no further semantics (e.g., that email must match an existing workspace member), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Delete') and resource ('Member From User Group'), so an agent can tell it removes a group membership rather than a workspace account. It does not, however, distinguish itself from close siblings such as add_member or update_workspace_member beyond the obvious verb flip.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus alternatives (e.g., get_workspace_members to inspect membership, update_workspace_member to change roles, or a workspace-level removal). Usage is only implied by the verb, with no prerequisites or exclusions stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_procedure_routeC
DestructiveIdempotent

Remove Procedure

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesAgent ID to get the procedure draft from
branch_idYesBranch ID to get the procedure draft from
procedure_idYesThe procedure ID

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true, idempotentHint=true, openWorldHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds no behavioral context beyond that, such as whether removal is permanent, what happens to associated data, or any required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two words, which is under-specified rather than truly concise. It lacks the structure needed to convey scope, prerequisites, or routing despite having ample room to do so.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, multi-parameter mutation tool, the description is severely incomplete. Although the annotations and schema cover safety hints and parameters, the description itself provides no operational context, such as what a procedure route is, how it relates to agent and branch IDs, or any side effects of removal.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the three parameters (agent_id, branch_id, procedure_id) are documented in the input schema. The description adds no parameter meaning beyond what the schema already provides, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb and resource ('Remove Procedure'), so the basic action is identifiable. However, it is essentially a restatement of the title and does not clarify whether this removes a procedure definition, a procedure route, or something else, nor does it distinguish it from the many delete/remove siblings in the API.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as delete_procedure_draft_route or update_procedure_draft_route. No prerequisites, exclusions, or context for selection are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_rulesC

Remove Rules From The Pronunciation Dictionary

ParametersJSON Schema
NameRequiredDescriptionDefault
rule_stringsYesList of strings to remove from the pronunciation dictionary.
pronunciation_dictionary_idYesThe id of the pronunciation dictionary

TDQS

C2.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds no behavioral context beyond the name. It also contradicts the annotations: the tool is called remove_rules and described as removing rules, while destructiveHint is false, indicating the tool performs only additive updates. This is an annotation contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded phrase with no wasted words. It is appropriately sized for a simple two-parameter removal tool, even if additional detail would be useful elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description does not explain effects, prerequisites, error behavior, or alternatives. The schema covers parameters, but the definition is incomplete regarding operational context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the schema. The description adds no additional semantic detail about rule_strings or pronunciation_dictionary_id, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: remove rules from a pronunciation dictionary. It is specific enough to distinguish this from unrelated siblings, but it does not explicitly differentiate itself from closely related tools such as add_rules or set_rules.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like add_rules, set_rules, or patch_pronunciation_dictionary. The implied usage is obvious from the name, but no conditions, prerequisites, or exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

renderA

Render Audio Or Video For The Given Language Spends ElevenLabs credits. Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
languageYesThe target language code to render, eg. 'es'. To render the source track use 'original'.
dubbing_idYesID of the dubbing project.
render_typeYes
normalize_volumeNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description adds that it spends credits and is deprecated, which is useful behavioral context not in the annotations. However, it doesn't clarify permission requirements or what happens to existing renders.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no waste, front-loading the core action and noting the credit cost and deprecation. It is appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter mutation tool with no output schema, the description is minimal. It conveys credit spending and deprecation but doesn't cover return values, side effects, or parameter details, leaving gaps for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, with the 'language' parameter well-documented in the schema and 'dubbing_id' partially documented, while 'render_type' and 'normalize_volume' lack descriptions. The description adds no parameter-level detail beyond what is already in the schema. Baseline 3 is appropriate when schema covers half the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Render') and resource ('Audio Or Video') for a given language. It doesn't differentiate itself from dubbing siblings like 'dub' or 'generate', but the core action is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes it 'Spends ElevenLabs credits' and is 'Deprecated upstream', which gives some context for caution, but it does not specify when to use this tool versus alternatives like 'dub' or 'generate'. Usage context is implied but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

replicate_voice_to_isolated_environmentC

Replicate Voice To Isolated Environment

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
preserve_voice_idNoWhen true (default) the replicated voice keeps the same voice ID in the target residency; set to false to assign a new voice ID.
target_workspace_idYesID of the workspace to replicate the voice into. It must belong to the same consolidated billing group as the calling workspace; the target's data residency is derived from that link.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, which covers the basic safety profile. However the description adds zero behavioral context beyond the name — nothing about cross-workspace side effects, whether the source voice is affected, or permission requirements. With annotations present the bar is lower, but this still contributes nothing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is not concise in a useful sense — it is under-specified. It repeats the tool title with no front-loaded information an agent could act on, so the brevity does not earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a mutation tool that crosses workspace/data-residency boundaries with no output schema and no description guidance. The critical prerequisite (same consolidated billing group) lives only in the schema, and the description never explains the operation's effect, so completeness is inadequate for what an agent needs before invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The schema itself fully documents voice_id, preserve_voice_id, and target_workspace_id (including the consolidated-billing-group constraint); the description adds no parameter meaning on top of it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is verbatim the tool title/name, 'Replicate Voice To Isolated Environment.' While it names a verb and resource, it is a tautological restatement of the identifier and adds no distinguishing detail versus siblings like add_voice, create_voice, or edit_voice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance whatsoever on when to use this tool versus the many other voice-management siblings (add_voice, create_voice, edit_voice, add_sharing_voice). No prerequisites, no exclusions, nothing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_pvc_manual_verificationC

Request Manual Verification

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
extra_textNoExtra text to be used in the manual verification process.
files_pathsNoVerification documents Local paths. Required for this call.
files_filenamesNoFilenames to send for "files". Some endpoints infer the audio format from them.
files_base64_listNoBase64 contents for "files", one entry per file. Use this when the server cannot read your local disk.

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false, covering basic safety traits. However, the description adds zero context about what 'manual verification' entails, whether it blocks, what documents are needed, or any side effects. It fails to add value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only three words, which is under-specification rather than conciseness. It is not front-loaded with useful information and does not earn its place as a meaningful tool description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation tool with no output schema, the description is completely inadequate. It omits what the tool does, required inputs, expected behavior, and any context needed for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters (voice_id, extra_text, files_paths, files_filenames, files_base64_list) are documented in the schema. The description adds no additional meaning, so the baseline of 3 is appropriate when the schema does all the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Request Manual Verification' merely restates the tool name and title without specifying the resource (PVC voice), the action's effect, or differentiating it from sibling tools like verify_pvc_voice_captcha or create_pvc_voice. It is a tautology, not a meaningful statement of purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites, and no mention of the manual verification workflow or required conditions. The description offers no context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

requests_listD

List Api Requests

ParametersJSON Schema
NameRequiredDescriptionDefault
sortNo
limitNo
searchNo
filtersNo
end_timeNo
start_timeNo

TDQS

D1.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description says 'List', which strongly implies a read-only operation, but the annotations declare readOnlyHint=false. That is a contradiction between the description and the structured safety metadata. The description also adds no other behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a three-word fragment with no structure or front-loaded detail. Brevity here reflects under-specification rather than useful conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with six undocumented parameters, no output schema, and contradictory annotations, the description is almost entirely incomplete. An agent lacks the information needed to invoke it correctly or understand its return behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All six parameters (sort, limit, search, filters, end_time, start_time) have 0% schema description coverage, and the description provides no parameter information at all. It therefore gives no meaning for sorting, filtering, pagination, or time-range behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb ('List') and a resource ('Api Requests'), so the basic action is identifiable. However, 'Api Requests' is ambiguous — it does not clarify whether these are logs, proxied API calls, audit records, or something else. It also does not distinguish this tool from the many other list/get siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives. The many sibling list_* and get_* tools make routing ambiguous, and the description provides no context, exclusions, or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_conversation_reference_routeD
Read-onlyIdempotent

Resolve Conversation Reference

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesAgent id (agent_…) or speech engine external id (seng_), resolved to the same underlying resource.
referenceYesA Slack message URL or a Zendesk ticket URL.

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, destructiveHint=false, which already carries the safety profile. The description adds nothing beyond that — notably it does not disclose that resolution depends on an external Slack/Zendesk lookup (openWorld), nor what a resolution yields. With annotations the bar is lower, but this still contributes zero behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but under-specification rather than genuine conciseness — a single title fragment carrying no information. There is nothing to waste, but also nothing earning its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-required-param resolver with no output schema, the description should at least say what resolution produces (conversation id? route?). It leaves the agent to infer everything from the name and schema, which is inadequate even though the schema itself is well documented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: the schema already explains that agent_id accepts agent_… or seng_ ids and that reference is a Slack message URL or Zendesk ticket URL. Per the rubric, full coverage sets baseline 3; the description adds no extra parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is the bare title 'Resolve Conversation Reference' — a restatement of the name that adds no verb specificity, no resource explanation, and no differentiation from siblings like get_conversation_* or the many *_route tools. An agent cannot tell from this text what 'resolving a reference' actually does or returns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no alternatives named. The agent has no cue for choosing this over the ~300 sibling conversation/route tools, or when a Slack/Zendesk reference should be resolved at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resubmit_tests_routeD

Resubmit Tests

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesAgent ID to resubmit tests for
branch_idNo
test_run_idsYesList of test run IDs to resubmit
test_invocation_idYesThe id of a test invocation. This is returned when tests are run.
agent_config_overrideNo

TDQS

D1.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the agent knows this is a non-idempotent write that is not destructive. The description adds nothing about what resubmission entails, whether it creates new test invocations, what happens to prior results, or any side effects. With annotations providing basic safety profile, the near-total absence of behavioral context beyond that is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two words, which is under-specified rather than concise. It lacks any structure or front-loaded information needed for correct invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation tool with no output schema and partial schema coverage, the description is completely inadequate. It omits essential context such as what resubmission does, required IDs, and expected behavior, making it impossible for an agent to use confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 60%, leaving gaps such as branch_id and agent_config_override undocumented in the schema. The description provides no parameter information at all, so it fails to compensate for the uncovered parameters. The baseline for 60% coverage without description help is below 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Resubmit Tests' merely restates the tool name (resubmit_tests_route) with no added specificity, scope, or distinction from siblings like run_agent_test_suite_route or create_agent_response_test_route. An agent cannot tell from the description what 'tests' means here (test runs, test invocations, test suites) or what resubmission implies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as run_agent_test_suite_route or create_agent_response_test_route. There is no mention of prerequisites, conditions, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retry_batch_callC

Retry A Batch Call.

ParametersJSON Schema
NameRequiredDescriptionDefault
batch_idYes

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose the important behavior profile: readOnlyHint=false, idempotentHint=false, openWorldHint=true, destructiveHint=false. The description contributes no additional context, such as what specifically gets re-executed, whether a new batch id is returned, or any error-state prerequisites. For a non-idempotent mutation, the description carries none of its own weight.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The one-line description is not padded, but that brevity reflects under-specification rather than efficiency. There is no structure, no front-loaded detail, and nothing an agent can act on beyond the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a single undocumented parameter, the description should explain prerequisites, retry scope, and return behavior. It supplies none of this, leaving the agent unable to determine when retry is appropriate versus when to cancel or recreate the batch.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter batch_id is undocumented in both schema and description. The description says nothing about where batch_id comes from or what form it takes, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb (retry) and resource (batch call), so the basic purpose is legible, but it adds nothing beyond the tool name and gives no detail on what 'retry' means (re-run all calls, only failed ones, resume from failure point). It also does not distinguish itself from the many sibling batch tools such as create_batch_call, cancel_batch_call, or get_batch_call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is provided. It does not say what state a batch must be in for retry to be valid (e.g. failed/partial) or point to alternatives like create_batch_call to start fresh. The only context an agent gets is the tool name itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_agent_test_suite_routeC

Run Tests On The Agent

ParametersJSON Schema
NameRequiredDescriptionDefault
testsYesList of tests to run on the agent
agent_idYesThe id of an agent. This is returned on agent creation.
branch_idNo
repeat_countNoNumber of times to run each test. When greater than 1, results are grouped and summarized.
agent_config_overrideNo

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide safety profile (readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false), but the description adds no behavioral context such as auth, rate limits, what gets created, or what happens on repeat_count. It merely restates the title.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase, but it is under-specified rather than concise; like the "Process" calibration, brevity without information does not earn a higher score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation-style route with nested agent_config_override, no output schema, and many sibling test routes, the description is far too thin. Annotations and partial schema descriptions help, but key routing and parameter details remain missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 60%; agent_id, tests, and repeat_count have schema descriptions, while branch_id and agent_config_override are undocumented. The description adds no parameter semantics at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Run Tests On The Agent" restates the tool name/title rather than specifying what kind of tests, on which agent artifacts, or how this differs from siblings like resubmit_tests_route or list_chat_response_tests_route. It gives a verb and resource but no scope or distinguishing detail, fitting the tautology/restatement level.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or alternatives are provided. The agent cannot infer whether this is for response tests, test suites, or a specific route relative to sibling test tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_conversation_analysisC

Run Conversation Analysis

ParametersJSON Schema
NameRequiredDescriptionDefault
conversation_idYesID of the conversation

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=false, and openWorldHint=true, which already signal that this is a mutating, non-idempotent operation with external effects. The description adds nothing to this - it doesn't explain what the analysis does, whether it incurs costs, how long it takes, or what happens after running. With annotations covering the safety profile, the description still fails to provide any operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but it's under-specified rather than concise. There's no structure or front-loading of useful information - it's just a restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that likely triggers a complex analysis process, the description is completely inadequate. It provides no information about what the analysis entails, what output to expect, or any side effects. The presence of an output schema would be helpful, but there is none.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the single parameter 'conversation_id' clearly documented as 'ID of the conversation'. The description adds no further parameter semantics. When schema coverage is this high, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Run Conversation Analysis' is essentially a restatement of the tool name. It doesn't specify what kind of analysis is performed, what output to expect, or how it differs from sibling tools like run_conversation_evaluations. It's a tautology that adds no clarity beyond the name itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when or why to use this tool versus alternatives. The sibling tools include run_conversation_evaluations and run_conversation_simulation_route, but the description provides no indication of which one to choose. No when-to-use, no prerequisites, no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_conversation_evaluationsC

Run Conversation Evaluation

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNo
evaluation_idYesID of the single evaluation criterion to rerun.
conversation_idYesID of the conversation

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the description should add context about what this write actually does — does it trigger a rerun of scoring, does it modify stored results, does it require an existing evaluation run. Instead the description adds nothing beyond the annotations, leaving the agent with no understanding of the mutation's effects or output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase with no waste, but it is under-specified rather than concise — it front-loads nothing beyond the title. It fails to earn its place as a standalone description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent, open-world write tool with an undocumented enum parameter and no output schema, the description should explain triggers, effects, and expected results. It does none of this, leaving the definition materially incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%, with conversation_id and evaluation_id documented in the schema while 'scope' (enum: conversation, agent) has no description. The tool description supplies no parameter-level detail at all. The schema carries most of the load, so the baseline 3 applies, but the description does not compensate for the undocumented scope parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name ('Run Conversation Evaluation') without adding specificity about what an evaluation is, how it is triggered, or what result it produces. It states a verb and resource, so it clears the vague bar, but it offers nothing an agent could not infer from the name itself. Among many sibling 'run_*' and 'conversation_*' tools, it does not distinguish itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites, and no reference to alternative tools such as run_conversation_analysis, run_conversation_simulation_route, or resubmit_tests_route. The agent must guess when this is the right tool versus those siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_conversation_simulation_routeC

Simulates A Conversation Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.
new_turns_limitNoMaximum number of new turns to generate in the conversation simulation
simulation_specificationYesA specification that will be used to simulate a conversation between an agent and an AI user.
extra_evaluation_criteriaNo

TDQS

C2.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is covered by structured data. The description adds only the deprecation status, which is real value not present in annotations, but discloses nothing about side effects, auth, or what a simulation actually creates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At two fragments it is not bloated, but this is under-specification rather than conciseness. Nothing meaningful is front-loaded because almost nothing is said.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A four-parameter tool with nested objects, no output schema, and only partial schema coverage needs substantive description. 'Simulates A Conversation Deprecated upstream.' is completely inadequate for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, below the >80% threshold, and the description supplies zero parameter meaning. With a nested simulation_specification object and an undocumented extra_evaluation_criteria, the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Simulates A Conversation' essentially restates the tool name (run_conversation_simulation_route) rather than explaining what a simulation run produces or how it differs from the sibling run_conversation_simulation_route_stream. There is a verb+resource, but no scope, no output, and no sibling differentiation, so it lands just above tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Deprecated upstream' is a genuine usage signal that the agent should prefer something else, which is better than nothing. However, no alternative is named even though run_conversation_simulation_route_stream sits in the sibling list, and there is no when/when-not guidance, leaving the agent to guess the replacement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_conversation_simulation_route_streamC

Simulates A Conversation (Stream) Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe id of an agent. This is returned on agent creation.
new_turns_limitNoMaximum number of new turns to generate in the conversation simulation
simulation_specificationYesA specification that will be used to simulate a conversation between an agent and an AI user.
extra_evaluation_criteriaNo

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is covered. The description adds the behavioral disclosure that the tool is deprecated upstream, which is useful but limited; it does not describe side effects, output, or streaming behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is a single very short sentence. It is front-loaded but severely under-specified for a complex simulation tool, so it is too brief rather than appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a complex simulation operation, nested parameters, and no output schema, the description omits return behavior, simulation semantics, and any migration guidance for the deprecation. It is not complete enough for an agent to call the tool confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, with nested objects and one undocumented top-level parameter (extra_evaluation_criteria). The description adds no parameter meaning at all, so it fails to compensate for the gaps in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Simulates') and resource ('A Conversation (Stream)'), and adds the critical qualifier 'Deprecated upstream'. It distinguishes this stream variant from the non-stream sibling run_conversation_simulation_route, though it offers no further scope detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says 'Deprecated upstream,' which is a usage warning, but it does not say when to use this tool instead of the non-stream simulation or any alternative. No explicit when/when-not guidance or migration path is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_pvc_voice_trainingD

Run Pvc Training

ParametersJSON Schema
NameRequiredDescriptionDefault
model_idNo
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations state readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the agent knows this is a non-read-only, non-idempotent operation. However, the description adds no behavioral context beyond the title, and does not explain what training entails, what state changes occur, or any side effects or prerequisites.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but this is under-specification rather than concise structure. It lacks any front-loaded useful detail, so brevity here does not earn a higher score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation-like training operation with no output schema and only partial annotation coverage, the description is completely inadequate. It gives an agent no basis to understand what the tool returns, what permissions are needed, or how to handle the training process.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: voice_id has a description pointing to a list endpoint, but model_id has no description. The tool description provides no additional parameter meaning, leaving model_id entirely undocumented and failing to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Run Pvc Training' is essentially a restatement of the tool name and title. It vaguely indicates an action on a 'Pvc Training' resource, but gives no detail on what training means, what it operates on, or how it differs from sibling tools like create_pvc_voice or add_pvc_voice_samples.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, when not to, or what alternatives exist. With many related PVC voice tools in the sibling list, the absence of any routing information is a critical gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_groupsC
Read-onlyIdempotent

Search User Groups

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesName of the target group.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds nothing about search behavior, matching semantics, pagination, or result limits, so it contributes no behavioral context beyond the title.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three words, front-loaded, and free of filler. It is highly concise, though its extreme brevity leaves the tool under-specified for an agent that must differentiate it from similar siblings.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter search tool with full schema coverage and rich annotations, the description covers the core purpose. However, it does not help the agent select this tool over sibling group-listing or member-listing tools, and it omits any search-specific context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single parameter 'name' is documented in the schema as 'Name of the target group.' The description adds no additional meaning beyond what the schema already provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Search) and resource (User Groups), so the basic action is clear. However, it does not distinguish this tool from sibling tools like get_groups_endpoint or get_workspace_members, leaving selection ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no when-to-use guidance, no conditions, and no alternatives. The agent is left to infer that this tool searches groups by name, but no explicit context or exclusions are offered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_knowledge_base_content_routeC
Read-onlyIdempotent

Search Knowledge Base Content

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesThe search query text
typesNoIf present, the endpoint will return only documents of the given types.
cursorNoUsed for fetching next page. Cursor is returned in the response.
page_sizeNoHow many documents to return at maximum. Can not exceed 100, defaults to 30.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is covered by structured data. The description adds nothing beyond that — no mention of pagination via cursor, the 100-item page cap, or what corpus is searched.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single short line has no wasted words and is front-loaded, but it is under-specified rather than genuinely concise — it conveys almost no usable information for a four-parameter search tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no description of return shape, ranking, or result scope, the definition leaves key behavior unexplained. The schema covers inputs, but the description does nothing to compensate for the absent output contract on a search endpoint.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with query, types, cursor and page_size all documented in the schema, so the baseline is 3. The description contributes no additional parameter meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb (Search) and a resource (Knowledge Base Content), so the basic operation is inferable, but it is essentially a restatement of the tool name/title with no differentiation from close siblings such as get_knowledge_base_content or query_agent_knowledge_base_rag_route. An agent cannot tell from this text what distinguishes this search endpoint from those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all — no indication of when this search should be preferred over get_knowledge_base_content, get_knowledge_base_list_route, or the RAG query route. The agent must infer selection purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

separate_song_stemsC

Stem Separation Spends ElevenLabs credits. Returns application/zip bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathNoThe audio file to separate into stems. Local path. Required for this call.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
sign_with_c2paNoWhether to sign the generated song with C2PA. Applicable only for mp3 files.
stem_variation_idNoThe id of the stem variation to use.

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description adds two genuinely useful pieces beyond that: the operation consumes ElevenLabs credits, and the response is zip bytes. It says nothing about permissions, processing time, or how many stems are produced, so it adds moderate rather than rich context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with no filler, and the credit-cost warning is front-loaded. The first clause is awkwardly phrased, which slightly hurts readability, but there is no wasted content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a credit-consuming, non-idempotent generation tool with no output schema, the description covers cost and return type but leaves notable gaps: which stems the two_stems_v1/six_stems_v1 variations yield, expected latency, and input format constraints. Annotations plus the rich schema carry much of the load, so 3 is appropriate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all seven parameters are documented in the schema itself, and the description only restates output_path behavior ('pass output_path to save them'). It adds no meaning beyond the schema, which sets the baseline at 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The name and the phrase 'Stem Separation' identify the resource and action, and 'Returns application/zip bytes' clarifies the output type. However, the sentence is grammatically garbled ('Stem Separation Spends ElevenLabs credits') and does not distinguish this tool from closely related siblings like start_speaker_separation or audio_isolation, so an agent cannot reliably route between them from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives such as start_speaker_separation or audio_isolation, nor any prerequisites (e.g. supported input formats). The only actionable guidance is a hint about output_path, which is invocation detail rather than usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_rulesC

Set Rules On The Pronunciation Dictionary

ParametersJSON Schema
NameRequiredDescriptionDefault
rulesYesList of pronunciation rules. Rule can be either: an alias rule: {'string_to_replace': 'a', 'type': 'alias', 'alias': 'b', } or a phoneme rule: {'string_to_replace': 'a', 'type': 'phoneme', 'phoneme': 'b', 'alphabet': 'ipa' }
pronunciation_dictionary_idYesThe id of the pronunciation dictionary

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond the word 'Set'; it does not explain replacement semantics, how existing rules are affected, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short, front-loaded phrase with no wasted words, which is structurally clean. However, it is under-specified for a mutation tool and reads more like a title than a complete instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool that sets rules on a pronunciation dictionary, the description omits crucial context: whether existing rules are replaced or merged, how duplicates are handled, and what happens to rules not included in the request. With no output schema and sibling add/remove tools, the agent lacks enough detail to call this confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both required parameters, including the detailed rule object structure. The description adds no additional meaning or syntax for pronunciation_dictionary_id or rules, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Set') and resource ('Rules On The Pronunciation Dictionary'), so the agent knows this acts on pronunciation dictionary rules. However, it does not distinguish the operation from sibling tools like add_rules or remove_rules, leaving the exact scope ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no when-to-use guidance, prerequisites, or alternatives. It does not mention when to use set_rules versus add_rules or remove_rules, leaving the agent to infer selection from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_third_party_disabling_policyD

Set Workspace Third-Party Disabling Policy

ParametersJSON Schema
NameRequiredDescriptionDefault
third_party_disable_allowedNo

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false, and the description adds nothing beyond them—no statement of scope, reversibility, or what disabling/enabling implies. For a workspace-wide security policy mutation, this is a serious gap with no compensating context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loaded, but it is under-specified rather than concise—it contributes no information beyond the name. Length is fine; content is absent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no annotations covering semantics (only safety hints), and a critical untyped boolean parameter leave the definition inadequate for an agent to invoke this workspace security-policy tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter, `third_party_disable_allowed`, with 0% schema description coverage and no explanation in the description. An agent cannot tell what true vs false does (allow disabling? disable third-party apps?) or whether the parameter is optional-with-default, since required count is 0.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description simply restates the tool name/title verbatim ('Set Workspace Third-Party Disabling Policy') with no added detail about what the policy governs or what setting it means. It is effectively a tautology of the identifier, with no differentiation from the many other `set_*`/`update_*` siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, what prerequisites exist, or what alternatives (e.g. update_settings_route, update_mcp_server_approval_policy_route) might apply. An agent has no basis to choose this tool over its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

share_resource_endpointD

Share Workspace Resource

ParametersJSON Schema
NameRequiredDescriptionDefault
roleYesRole to grant to the target: one of 'admin', 'editor', 'commenter', or 'viewer'.
group_idNo
user_emailNo
resource_idYesThe ID of the target resource.
resource_typeYesResource types that can be shared in the workspace. The name always need to match the collection names
workspace_api_key_idNo

TDQS

D1.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, openWorldHint=true and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that: it does not say whether existing shares are replaced, what permissions are required, or whether an email/group must be pre-existing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words with no filler, but this is under-specification rather than conciseness. Nothing is front-loaded because there is effectively no content to structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent, open-world mutation with six parameters (three undescribed) and no output schema, the description is completely inadequate. An agent cannot call this correctly from the description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50% and the description provides zero parameter meaning. It does not clarify the critical target ambiguity (user_email vs group_id vs workspace_api_key_id) or how role and resource_type interact, so the missing half of the schema is left uncompensated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Share Workspace Resource' essentially restates the tool name and title rather than telling the agent what sharing means here (granting a role to a user or group on a specific resource type). It gives no differentiation from close siblings like unshare_resource_endpoint or add_sharing_voice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, prerequisite, or alternative guidance at all. The agent gets no signal about choosing this over unshare_resource_endpoint or other sharing-related siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

smart_search_conversation_messages_routeC
Read-onlyIdempotent

Smart Search Conversation Messages

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
agent_idNoAgent id (agent_…) or speech engine external id (seng_), resolved to the same underlying resource.
page_sizeNoNumber of results per page. Max 50.
text_queryYesThe search query text for semantic similarity matching

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond those annotations, such as pagination behavior, result ordering, or authentication requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but its brevity reflects under-specification rather than efficiency. It provides only a title-like phrase and does not front-load any useful scope, constraints, or routing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of a close sibling search tool and four parameters, the description is not complete enough to guide correct invocation. It omits what makes this search 'smart', how it differs from text search, and any usage constraints beyond what the schema and annotations already provide.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for all four parameters, so the schema already documents cursor, agent_id, page_size, and text_query. The description adds no additional parameter meaning, which is the baseline when schema coverage is high.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description simply restates the tool name/title as 'Smart Search Conversation Messages' rather than specifying a verb and resource with distinguishing detail. It implies searching conversation messages, but does not distinguish this from sibling text_search_conversation_messages_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. In particular, the sibling text_search_conversation_messages_route suggests a competing search mode, but the description gives no criteria for choosing between them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sound_generationC

Sound Generation Spends ElevenLabs credits. Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
loopNoWhether to create a sound effect that loops smoothly. Only available for the 'eleven_text_to_sound_v2 model'.
textYesThe text that will get converted into a sound effect.
model_idNo
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
duration_secondsNo
prompt_influenceNo

TDQS

C2.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the operation profile (readOnly=false, idempotent=false, openWorld=true, destructive=false), and the description usefully adds context the annotations do not: that the call consumes ElevenLabs credits (a cost/side-effect disclosure) and that the return payload is audio/mpeg bytes. It stops short of permission/tier requirements or failure behavior, but it clearly pulls its weight beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with no filler. The cost warning leads, followed by the return-handling note, which is a sensible ordering for an agent. It is terse almost to a fault, but nothing in it is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, non-idempotent, credit-spending generation tool with no output schema, the description covers the biggest agent concern (output delivery and cost) but omits the actual generative purpose and leaves duration_seconds and prompt_influence entirely undocumented anywhere. It is adequate but leaves real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 57%, and the description only echoes the output_path behavior ("pass output_path to save them") that the schema's own output_path description already states in more detail (relative path resolution against ELEVENLABS_OUTPUT_DIR, inline base64 fallback). It adds no meaning for duration_seconds, prompt_influence, or model_id, which carry no schema descriptions at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Sound Generation" essentially restates the tool name/title rather than stating a verb+resource with scope; the actual function (converting a text prompt into a sound effect) is only discoverable from the schema's `text` description. It also fails to distinguish itself from near-neighbors like text_to_speech_full, create_video_generation, or video_to_music, which occupy adjacent audio-generation space.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative-tool guidance. The only usage-relevant content is the implicit cost warning ("Spends ElevenLabs credits") and the output_path hint, neither of which helps an agent decide between this and the other audio-generation siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

speech_to_speech_fullB

Speech To Speech Spends ElevenLabs credits. Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNoIf specified, our system will make a best effort to sample deterministically, such that repeated requests with the same seed and parameters should return the same result. Determinism is not guaranteed. Must be integer between 0 and 4294967295.
model_idNoIdentifier of the model that will be used, you can query them using GET /v1/models. The model needs to have support for speech to speech, you can check this using the can_do_voice_conversion property.
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
audio_pathNoThe audio file which holds the content and emotion that will control the generated speech. Local path. Required for this call.
file_formatNoThe format of input audio. Options are 'pcm_s16le_16' or 'other' For `pcm_s16le_16`, the input audio must be 16-bit PCM at a 16kHz sample rate, single channel (mono), and little-endian byte order. Latency will be lower than with passing an encoded waveform.
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
audio_base64NoBase64 contents for "audio". Use this when the server cannot read your local disk.
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM and WAV formats with 44.1kHz sample rate requires you to be
audio_filenameNoFilename to send for "audio". Some endpoints infer the audio format from it.
enable_loggingNoWhen enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.
voice_settingsNoVoice settings overriding stored settings for the given voice. They are applied only on the given request. Needs to be send as a JSON encoded string.
remove_background_noiseNoIf set, will remove the background noise from your audio input using our audio isolation model. Only applies to Voice Changer.
optimize_streaming_latencyNoYou can turn on latency optimizations at some cost of quality. The best possible final latency varies by model. Possible values: 0 - default mode (no latency optimizations) 1 - normal latency optimizations (about 50% of possible latency improvement of option 3) 2 - strong latency optimizations (abou

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare a non-read-only, non-idempotent, open-world operation, and the description adds genuinely useful behavioral context beyond that: it consumes ElevenLabs credits and returns audio/mpeg bytes rather than structured data. It stops short of disclosing latency, long-running behavior, or enterprise/tier constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler, and the cost warning is front-loaded ahead of the output-format detail. The first sentence reads awkwardly ('Speech To Speech Spends ElevenLabs credits') but the content is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter generation tool with no output schema, the description covers the return type and cost, which is the right emphasis. However, it omits input requirements (voice_id required, audio_path or audio_base64 supply path) and any note on runtime or tier restrictions, leaving meaningful gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 13 parameters are already documented in the schema. The description only restates one of them (output_path for saving bytes), which is already described there, and adds nothing about voice_id, audio_path, or model_id selection.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific operation, 'Speech To Speech', on an audio resource, and implies it produces converted audio. It is distinguishable from the sibling speech_to_speech_stream only by the implicit 'full' vs 'stream' contrast, which the description does not spell out, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It notes that credits are spent, which hints at a cost consideration, but gives no explicit when-to-use guidance, no conditions, and no mention of when to prefer speech_to_speech_stream or other conversion tools in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

speech_to_speech_streamB

Speech To Speech Streaming Spends ElevenLabs credits. Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNoIf specified, our system will make a best effort to sample deterministically, such that repeated requests with the same seed and parameters should return the same result. Determinism is not guaranteed. Must be integer between 0 and 4294967295.
model_idNoIdentifier of the model that will be used, you can query them using GET /v1/models. The model needs to have support for speech to speech, you can check this using the can_do_voice_conversion property.
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
audio_pathNoThe audio file which holds the content and emotion that will control the generated speech. Local path. Required for this call.
file_formatNoThe format of input audio. Options are 'pcm_s16le_16' or 'other' For `pcm_s16le_16`, the input audio must be 16-bit PCM at a 16kHz sample rate, single channel (mono), and little-endian byte order. Latency will be lower than with passing an encoded waveform.
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
audio_base64NoBase64 contents for "audio". Use this when the server cannot read your local disk.
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
audio_filenameNoFilename to send for "audio". Some endpoints infer the audio format from it.
enable_loggingNoWhen enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.
voice_settingsNoVoice settings overriding stored settings for the given voice. They are applied only on the given request. Needs to be send as a JSON encoded string.
remove_background_noiseNoIf set, will remove the background noise from your audio input using our audio isolation model. Only applies to Voice Changer.
optimize_streaming_latencyNoYou can turn on latency optimizations at some cost of quality. The best possible final latency varies by model. Possible values: 0 - default mode (no latency optimizations) 1 - normal latency optimizations (about 50% of possible latency improvement of option 3) 2 - strong latency optimizations (abou

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds non-obvious behavioral context: it spends ElevenLabs credits, returns audio/mpeg bytes, and the output_path parameter controls whether bytes are saved or returned inline. Annotations cover safety (readOnlyHint=false, destructiveHint=false, etc.) but this credit-spending and output-format detail goes beyond what annotations declare.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the cost warning and return format. No wasted words, though it could be slightly more structured (e.g., separating cost from output behavior).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 13 parameters, no output schema, and rich sibling context, the description is minimal. It covers cost and output handling but omits required input parameters (voice_id is required, audio_path is 'Required for this call') and doesn't mention that this is a streaming endpoint. More would be expected for a tool with this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 13 parameters thoroughly. The description only adds meaning for output_path (saving bytes) and implies the return type, but doesn't explain other important params like voice_id, audio_path, or audio_base64. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the resource (speech-to-speech streaming) and a key behavior (returns audio/mpeg bytes). However, it doesn't explicitly differentiate this streaming variant from the sibling 'speech_to_speech_full' or clarify that it performs voice conversion rather than text-to-speech.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like speech_to_speech_full or text_to_speech_stream. The only conditional hint is 'pass output_path to save them', which is about output handling, not tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

speech_to_textC

Speech To Text Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNoIf specified, our system will make a best effort to sample deterministically, such that repeated requests with the same seed and parameters should return the same result. Determinism is not guaranteed. Must be an integer between 0 and 2147483647.
tokenNoA single-use authentication token created via POST /v1/single-use-token/batch_scribe. This token can only be used once and expires after 15 minutes. Alternative to API key or bearer token authentication for frontend clients.
diarizeNoWhether to annotate which speaker is currently talking in the uploaded file.
webhookNoWhether to send the transcription result to configured speech-to-text webhooks. If set the request will return early without the transcription, which will be delivered later via webhook.
keytermsNoA list of keyterms to bias the transcription towards. The keyterms are words or phrases you want the model to recognise more accurately. The number of keyterms cannot exceed 1000. The length of each keyterm must be less than 50 characters. Keyterms can contain at most 5 words (after normalisation).
model_idYesThe ID of the model to use for transcription.
file_pathNoThe file to transcribe (100ms minimum audio length). All major audio and video formats are supported. Exactly one of the file or cloud_storage_url parameters must be provided. The file size must be less than 5.0GB. Local path.
source_urlNoThe URL of an audio or video file to transcribe. Supports hosted video or audio files, YouTube video URLs, TikTok video URLs, and other video hosting services.
webhook_idNoOptional specific webhook ID to send the transcription result to. Only valid when webhook is set to true. If not provided, transcription will be sent to all configured speech-to-text webhooks.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
file_formatNoThe format of input audio. Options are 'pcm_s16le_16' or 'other' For `pcm_s16le_16`, the input audio must be 16-bit PCM at a 16kHz sample rate, single channel (mono), and little-endian byte order. Latency will be lower than with passing an encoded waveform.
no_verbatimNoIf true, the transcription will not have any filler words, false starts and non-speech sounds. Only supported with scribe_v2 model.
temperatureNoControls the randomness of the transcription output. Accepts values between 0.0 and 2.0, where higher values result in more diverse and less deterministic results. If omitted, we will use a temperature based on the model you selected which is usually 0.
num_speakersNoThe maximum amount of speakers talking in the uploaded file. Can help with predicting who speaks when. The maximum amount of speakers that can be predicted is 32. Defaults to null, in this case the amount of speakers is set to the maximum value the model supports.
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.
language_codeNoAn ISO-639-1 or ISO-639-3 language_code corresponding to the language of the audio file. Can sometimes improve transcription performance if known beforehand. Defaults to null, in this case the language is predicted automatically.
enable_loggingNoWhen enable_logging is set to false zero retention mode will be used for the request. This will mean log and transcript storage features are unavailable for this request. Zero retention mode may only be used by enterprise customers.
entity_detectionNoDetect entities in the transcript. Can be 'all' to detect all entities, a single entity type or category string, or a list of entity types/categories. Categories include 'pii', 'phi', 'pci', 'other', 'offensive_language'. When enabled, detected entities will be returned in the 'entities' field with
entity_redactionNoRedact entities from the transcript text. Accepts the same format as entity_detection: 'all', a category ('pii', 'phi'), or specific entity types. Must be a subset of entity_detection. When redaction is enabled, the entities field will not be returned. Usage of this parameter will incur an additiona
tag_audio_eventsNoWhether to tag audio events like (laughter), (footsteps), etc. in the transcription.
webhook_metadataNoOptional metadata to be included in the webhook response. This should be a JSON string representing an object with a maximum depth of 2 levels and maximum size of 16KB. Useful for tracking internal IDs, job references, or other contextual information.
cloud_storage_urlNo[Deprecated] This parameter is deprecated and will be removed in the future. Use 'source_url' instead.The HTTPS URL of the file to transcribe. Exactly one of the file or cloud_storage_url parameters must be provided. The file must be accessible via HTTPS and the file size must be less than 2GB. Any
use_multi_channelNoWhether the audio file contains multiple channels where each channel contains a single speaker. When enabled, each channel is transcribed independently. By default a separate transcript is returned per channel; set multichannel_output_style='combined' to instead receive a single transcript with all
additional_formatsNo
use_speaker_libraryNoWhether to use the speaker library for identifying known speakers during diarization. When enabled and diarize is true, detected speakers will be matched against registered speakers in the workspace's speaker library.
detect_speaker_rolesNoWhether to detect speaker roles (agent vs customer). Requires diarize=true. Cannot be used with use_multi_channel=true. When enabled, speaker_id values will be 'agent' and 'customer' instead of 'speaker_0', 'speaker_1', etc. Usage incurs an additional 10% surcharge on base transcription cost.
diarization_thresholdNoDiarization threshold to apply during speaker diarization. A higher value means there will be a lower chance of one speaker being diarized as two different speakers but also a higher chance of two different speakers being diarized as one speaker (less total speakers predicted). A low value means the
entity_redaction_modeNoHow to format redacted entities. 'redacted' replaces with {REDACTED}, 'entity_type' replaces with {ENTITY_TYPE}, 'enumerated_entity_type' replaces with {ENTITY_TYPE_N} where N enumerates each occurrence. Only used when entity_redaction is set.
timestamps_granularityNoThe granularity of the timestamps in the transcription. 'word' provides word-level timestamps and 'character' provides character-level timestamps per word.
multichannel_output_styleNoControls the response shape when use_multi_channel is enabled. 'separate' (default) returns one transcript per channel under 'transcripts'. 'combined' merges all channels into a single transcript whose words are sorted by start time, each carrying a 'channel_index' - matching the single-channel resp

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=false, and openWorldHint=true, so the agent already knows this is a non-repeatable write-ish operation. The description's only added trait is that it consumes ElevenLabs credits, which is genuinely useful but far short of what a 30-parameter transcription job needs (async webhook delivery, retention/zero-retention mode, auth token one-time use).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but shortness here is under-specification rather than conciseness. A single sentence that mostly repeats the title and mentions billing leaves the essential behavior unstated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 30-parameter, open-world, non-idempotent transcription tool with no output schema, one sentence of billing trivia is completely inadequate. Nothing about file/source selection, model choice, webhook mode, or result delivery is conveyed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 97%, so the schema already documents all 30 parameters in detail. The description adds no parameter meaning whatsoever, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description only restates the tool name ('Speech To Text') and appends a billing note. It never states the actual operation (transcribing audio/video to text), the accepted input forms, or how it differs from nearest siblings like 'transcribe' or 'forced_alignment'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of alternatives (transcribe, create_batch_call), and no note about the async webhook path. Only the credit-cost warning hints at any decision input.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_speaker_separationC

Start Speaker Separation

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
sample_idYesSample ID to be used

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a non-read-only, open-world, non-idempotent operation, but the description adds no behavioral context such as whether it starts an async job, what permissions are needed, or what the result contains. It does not contradict the annotations, but it also adds no value beyond them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase, which is concise but severely under-specified. It is not wasteful, but it lacks any structure or front-loaded useful information beyond the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two required parameters, no output schema, and only partial annotation coverage of behavior, the description is too thin. It omits what speaker separation produces, whether it is asynchronous, and how results are retrieved, leaving the agent under-informed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters are fully described in the schema. The description adds no parameter meaning, so the baseline of 3 applies because the schema already does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description simply repeats the tool name and title, 'Start Speaker Separation,' adding no specific verb-resource detail beyond the name itself. It does not distinguish this tool from siblings such as separate_song_stems, audio_isolation, or create_speaker.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The title implies a general purpose, but there are no prerequisites, exclusions, or sibling routing hints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stream_chapter_snapshot_audioA

Stream Chapter Audio Spends ElevenLabs credits. Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
chapter_idYesThe ID of the chapter.
project_idYesThe ID of the Studio project.
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
convert_to_mpegNoWhether to convert the audio to mpeg format.
chapter_snapshot_idYesThe ID of the chapter snapshot.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds a genuine non-obvious cost warning ('Spends ElevenLabs credits') that annotations do not convey, and pairs it with the idempotentHint=false implication that repeated calls keep consuming credits. It also discloses the return format, which the absence of an output schema leaves otherwise unknown.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no filler, with the cost warning and return type front-loaded. Slightly marred by the clipped opening clause ('Stream Chapter Audio Spends ElevenLabs credits'), which reads as a truncated title fragment.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-format burden and does it correctly (audio/mpeg bytes, saved or returned inline), plus the credit cost. It stops short of covering errors such as credit exhaustion or missing snapshot IDs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters, including the ELEVENLABS_OUTPUT_DIR resolution and base64 fallback for output_path, are already documented. The description's output_path mention restates the schema rather than adding syntax or constraints, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Stream) and resource (Chapter Audio) plus the return type (audio/mpeg bytes). An agent can distinguish it from stream_project_snapshot_audio_endpoint, though the description never names that sibling to make the distinction explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives operational guidance for one use case ('pass output_path to save them'), which implies streaming vs inline base64 handling, but never says when to choose this tool over get_chapter_snapshot_endpoint or the project-snapshot audio variant.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stream_composeB

Stream Composed Music Spends ElevenLabs credits. Returns audio/* bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
promptNo
model_idNo
finetune_idNo
lyrics_textNo
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
music_promptNoComposition plan for the `music_v1` model. Using this field with any other model will result in an error.
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. Use "auto" (the default) to let the API pick the best format for the selected model: mp3_44100_128 for v1 models and mp3_48000_192 for v2 models.
generation_modeNo
music_length_msNo
composition_planNo
finetune_strengthNoHow strongly the finetune influences the generation. Defaults to 1.0 (full strength). Lower values soften the influence of the finetune, leaving more room for prompt-level steering. Only meaningful when `finetune_id` is also provided.
force_instrumentalNoIf true, guarantees that the generated song will be instrumental. If false, the song may or may not be instrumental depending on the `prompt`. Can only be used with `prompt`.
use_phonetic_namesNoIf true, proper names in the prompt will be phonetically spelled in the lyrics for better pronunciation by the music model. The original names will be restored in word timestamps.
store_for_inpaintingNoWhether to store the generated song for inpainting.

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations state readOnlyHint=false, idempotentHint=false, openWorldHint=true, destructiveHint=false. The description adds two useful behavioral facts: it consumes ElevenLabs credits and returns audio/* bytes, with output_path controlling inline vs. saved. That covers cost and return channel, but not auth, errors, streaming semantics, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded, no filler. It is slightly terse given a 15-parameter schema, but every sentence earns its place and there is no structural waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter, zero-required, streaming generation tool with no output schema and 47% schema coverage, the description is far too thin. It omits prompt/lyrics/model relationships, generation_mode meaning, music_length_ms, seed, finetune interactions, and the music_v1-only constraint on music_prompt/composition_plan already noted in schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 47%, so the schema explains some parameters (output_path, output_format, finetune_strength, force_instrumental, use_phonetic_names, store_for_inpainting, music_prompt) while several remain bare (seed, prompt, model_id, finetune_id, lyrics_text, generation_mode, music_length_ms, composition_plan). The description adds output_path semantics but does not compensate for the undocumented half.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it composes music by streaming. 'Stream Composed Music' names the operation and the artifact clearly, distinguishing it from text-to-speech or sound_generation siblings. It falls short of a 5 only because it doesn't discuss ties to sibling compose tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus siblings like compose_detailed, compose_detailed_stream, compose_plan, or sound_generation. The only routing hint is the output_path behavior, which is operational rather than selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stream_project_snapshot_archive_endpointA

Stream Archive With Studio Project Audio Spends ElevenLabs credits. Returns application/x-zip bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesThe ID of the Studio project.
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
project_snapshot_idYesThe ID of the Studio project snapshot.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover safety hints (readOnly=false, destructive=false, idempotent=false, openWorld=true); they do not mention cost. The description adds a genuinely useful behavior not in structured data: it consumes ElevenLabs credits. It also discloses the return media type (application/x-zip bytes), which matters absent an output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler; the return type and the output_path hint are front-loaded enough to be useful. The awkward 'Stream Archive With Studio Project Audio Spends ElevenLabs credits' run-on slightly hurts first-pass parsing but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly states the return type and that bytes can be captured via output_path. Cost disclosure and the open-world annotation together give an agent enough to call this safely, though the missing when-to-use routing against sibling stream/get endpoints leaves a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description only reinforces output_path behavior ('pass output_path to save them') that the schema already documents in more detail, adding no new semantics for project_id or project_snapshot_id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific operation — streaming an archive (zip) of a Studio project snapshot — which an agent can distinguish from the sibling stream_project_snapshot_audio_endpoint (audio only) and get_project_snapshot_endpoint (metadata). The phrasing 'Stream Archive With Studio Project Audio' is garbled, but the resource and action are recoverable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage instruction is the output_path mechanics ('pass output_path to save them'), which is a parameter hint rather than when-to-use guidance. There is no statement of when this archive stream should be chosen over the sibling audio stream or the snapshot metadata endpoint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stream_project_snapshot_audio_endpointC

Stream Studio Project Audio Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesThe ID of the Studio project.
convert_to_mpegNoWhether to convert the audio to mpeg format.
project_snapshot_idYesThe ID of the Studio project snapshot.

TDQS

C2.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true). The description adds one genuinely useful fact beyond them: the operation consumes ElevenLabs credits, which is critical cost context. It does not say whether credits are spent per call, on failure, or whether output is audio bytes, so the disclosure is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, but the sentence itself is mangled and ungrammatical ('Stream Studio Project Audio Spends ElevenLabs credits'), which harms readability rather than achieving conciseness. The core verb-resource meaning is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a streaming/mutating endpoint with no output schema, three parameters, and no annotations telling the agent what the return shape is. The description provides cost context but omits the pointer to the sibling chapter variant, doesn't explain the options, and doesn't clarify that 'snapshot' is required. Incomplete for the complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each of the three parameters is already documented in the schema. The description adds no parameter detail, so the baseline 3 applies. The credits note does not tell the agent anything about parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

Attempts to state a verb and resource ('Stream Studio Project Audio'), but the phrasing is clumsy and omits the snapshot dimension entirely, despite 'snapshot' being in the tool name and a required parameter. An agent cannot distinguish this from the sibling stream_chapter_snapshot_audio. The name is clearer than the description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, no alternative named despite the obvious sibling stream_chapter_snapshot_audio and stream_project_snapshot_archive_endpoint. Nothing tells the agent which of the similar streaming endpoints to pick.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_search_conversation_messages_routeD
Read-onlyIdempotent

Text Search Conversation Messages

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoUsed for fetching next page. Cursor is returned in the response.
sort_byNoSort order for search results. 'search_score' sorts by search score, 'created_at' sorts by conversation start time.
user_idNoFilter conversations by the user ID who initiated them.
agent_idNoAgent id (agent_…) or speech engine external id (seng_), resolved to the same underlying resource.
branch_idNoFilter conversations by branch ID.
page_sizeNoNumber of results per page. Max 50.
text_onlyNo
topic_idsNoFilter conversations by topic IDs assigned during topic discovery.
rating_maxNoMaximum overall rating (1-5).
rating_minNoMinimum overall rating (1-5).
text_queryYesThe search query text for full-text and fuzzy matching
tool_namesNoFilter conversations by tool names used during the call.
version_idNoFilter conversations by version ID.
summary_modeNoWhether to include transcript summaries in the response.
main_languagesNoFilter conversations by detected main language (language code).
call_successfulNoThe result of the success evaluation
exclude_statusesNoExclude conversations with the given statuses. Useful for hiding in-progress / processing conversations from list views.
evaluation_paramsNoEvaluation filters. Repeat param. Format: criteria_id:result. Example: eval=value_framing:success
visited_agent_idsNoFilter conversations where any of these agents participated. Can not exceed 50 values.
tool_names_erroredNoFilter conversations by tool names that had errored calls.
termination_reasonsNoFilter conversations by their stored termination_reason (metadata.termination_reason). Repeat param to match any of several.
has_feedback_commentNoFilter conversations with user feedback comments.
call_start_after_unixNoUnix timestamp (in seconds) to filter conversations after to this start date.
tool_names_successfulNoFilter conversations by tool names that had successful calls.
call_duration_max_secsNoMaximum call duration in seconds.
call_duration_min_secsNoMinimum call duration in seconds.
call_start_before_unixNoUnix timestamp (in seconds) to filter conversations up to this start date.
data_collection_paramsNoData collection filters. Repeat param. Format: id:op:value where op is one of eq|gt|gte|lt|lte|missing.
dynamic_variable_paramsNoDynamic variable filters. Repeat param. Format: name:op:value where op is one of eq|gt|gte|lt|lte. Comparison operators require a numeric value. Names containing ':' cannot be expressed.
triggered_procedure_idsNoFilter conversations where any of these procedures were triggered. Can not exceed 50 values.
visited_agent_branch_idsNoFilter conversations where any of these agent branches participated. Can not exceed 50 values.
conversation_product_typeNoRestrict results to a single conversation product surface.
include_invalid_tool_callsNoAlso match tool calls that never ran.
conversation_initiation_sourceNo

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false. The description adds no additional behavioral context such as pagination behavior, search semantics, required permissions, or result characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single phrase, but it is under-specified rather than meaningfully concise. It repeats the title and provides no front-loaded operational information beyond the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 34-parameter conversation-message search tool with a closely related sibling and no output schema, this description is inadequate. It does not explain search scope, filtering behavior, pagination, or how to choose this tool over alternatives.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 94%, so the schema already documents nearly all 34 parameters thoroughly. The description adds no extra parameter meaning, so the baseline of 3 is appropriate when the schema carries the burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is only "Text Search Conversation Messages", which essentially restates the tool name/title. It names the resource and search action but adds no explanatory detail and does not distinguish it from the sibling smart_search_conversation_messages_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, when not to use it, or how it differs from alternatives such as smart_search_conversation_messages_route. The description gives no context for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_dialogueB

Text To Dialogue (Multi-Voice) Spends ElevenLabs credits. Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
inputsYesA list of dialogue inputs, each containing text and a voice ID which will be converted into speech. The maximum number of unique voice IDs is 10. For reliable generation, keep the total character count across all `inputs[].text` values at or below 2,000 characters per request. Longer requests can te
model_idNoIdentifier of the model that will be used, you can query them using GET /v1/models. The model needs to have support for text to speech, you can check this using the can_do_text_to_speech property.
settingsNo
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
language_codeNo
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM and WAV formats with 44.1kHz sample rate requires you to be
enable_loggingNoWhen enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.
apply_text_normalizationNoThis parameter controls text normalization with three modes: 'auto', 'on', and 'off'. When set to 'auto', the system will automatically decide whether to apply text normalization (e.g., spelling out numbers). With 'on', text normalization will always be applied, while with 'off', it will be skipped.
pronunciation_dictionary_locatorsNo

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (which only mark it non-readonly/openWorld/non-idempotent/non-destructive), the description adds genuinely useful behavior: it consumes ElevenLabs credits and returns audio/mpeg bytes, with the output-saving mechanism surfaced. These are not derivable from the annotations and directly guide correct invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences with no filler, and the credit-cost warning is front-loaded. It is efficient though perhaps too terse given the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter tool with no output schema, the description valuably names the return type, but it omits sibling differentiation, stream-vs-full guidance, and any hint at the input constraints (max 10 voices, character limits) that live only in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60% and the description only touches output_path semantics, which the schema already documents. It adds no meaning for the voice-ID inputs, model, settings, or output_format parameters, so the baseline of 3 is appropriate with the schema doing most of the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific transformation (text to multi-voice dialogue) that names the resource and scope. It does not, however, distinguish this base variant from the several sibling variants (text_to_dialogue_stream, text_to_dialogue_full_with_timestamps, text_to_dialogue_stream_with_timestamps), which an agent must choose between.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage context is an implicit 'multi-voice' cue plus a cost warning that it spends credits. There is no statement of when to use this versus the streaming or timestamped siblings, and no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_dialogue_full_with_timestampsC

Text To Dialogue With Timestamps Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
inputsYesA list of dialogue inputs, each containing text and a voice ID which will be converted into speech. The maximum number of unique voice IDs is 10. For reliable generation, keep the total character count across all `inputs[].text` values at or below 2,000 characters per request. Longer requests can te
model_idNoIdentifier of the model that will be used, you can query them using GET /v1/models. The model needs to have support for text to speech, you can check this using the can_do_text_to_speech property.
settingsNo
language_codeNo
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM and WAV formats with 44.1kHz sample rate requires you to be
enable_loggingNoWhen enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.
apply_text_normalizationNoThis parameter controls text normalization with three modes: 'auto', 'on', and 'off'. When set to 'auto', the system will automatically decide whether to apply text normalization (e.g., spelling out numbers). With 'on', text normalization will always be applied, while with 'off', it will be skipped.
pronunciation_dictionary_locatorsNo

TDQS

C2.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the mutation/safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false). The description adds one genuinely new behavioral fact — that the call consumes ElevenLabs credits — which is useful cost/billing context. However it omits other important traits such as synchronous blocking, latency, and the 2,000-character request limits that matter for this long-running generation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence, but the brevity is under-specification rather than concision — half the sentence merely repeats the title. It is not front-loaded with actionable purpose information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, credit-consuming audio generation tool with no output schema, the description is grossly incomplete: no explanation of inputs, no note about the timestamps in the response, no limits, and no differentiation from streaming or non-timestamp variants.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 56%, so the description is expected to compensate for undocumented parameters (seed, settings.stability, language_code, pronunciation_dictionary_locators). It contributes nothing about any parameter, leaving gaps in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description essentially restates the tool name ('Text To Dialogue With Timestamps') and adds only a credit-cost note. It never states what the tool actually produces — dialogue audio with timestamp alignment — in a way that distinguishes it from siblings like text_to_dialogue, text_to_dialogue_stream, or text_to_speech_full_with_timestamps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives, despite the crowded sibling space (text_to_dialogue, text_to_dialogue_stream, text_to_dialogue_stream_with_timestamps, text_to_speech_full_with_timestamps). The agent is left to infer selection entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_dialogue_streamB

Text To Dialogue (Multi-Voice) Streaming Spends ElevenLabs credits. Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
inputsYesA list of dialogue inputs, each containing text and a voice ID which will be converted into speech. The maximum number of unique voice IDs is 10. For reliable generation, keep the total character count across all `inputs[].text` values at or below 2,000 characters per request. Longer requests can te
model_idNoIdentifier of the model that will be used, you can query them using GET /v1/models. The model needs to have support for text to speech, you can check this using the can_do_text_to_speech property.
settingsNo
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
language_codeNo
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
enable_loggingNoWhen enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.
apply_text_normalizationNoThis parameter controls text normalization with three modes: 'auto', 'on', and 'off'. When set to 'auto', the system will automatically decide whether to apply text normalization (e.g., spelling out numbers). With 'on', text normalization will always be applied, while with 'off', it will be skipped.
pronunciation_dictionary_locatorsNo

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it is a non-read-only, open-world, non-idempotent operation, and the description usefully adds that it consumes credits and returns audio/mpeg bytes. These two facts (billing impact and return media type) are not derivable from the annotations and materially affect invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler, and the credit-cost warning is placed early. The phrasing 'Text To Dialogue (Multi-Voice) Streaming Spends ElevenLabs credits' is slightly run-on, but the content is tightly written.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter generation tool with no output schema, the description covers cost, return type, and persistence, but omits the dialogue input format constraints, the meaning of streaming, and voice/model selection context. Adequate as a minimum but leaves gaps an agent would hit when constructing inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 10 parameters at 60% schema coverage, roughly four parameters rely on the description for meaning, yet it only restates the output_path behavior that the schema already documents. Nothing is added about inputs[], seed, model_id, language_code, settings, or pronunciation_dictionary_locators.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (stream multi-voice text-to-dialogue audio) and clarifies it is the streaming variant, which separates it from text_to_dialogue and text_to_dialogue_full_with_timestamps. However, it does not explicitly contrast itself with those siblings, so an agent still has to infer the difference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance is a cost warning ('Spends ElevenLabs credits'). There is no statement of when to pick this over text_to_dialogue, text_to_dialogue_stream_with_timestamps, or text_to_speech_stream, and no prerequisites or exclusion conditions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_dialogue_stream_with_timestampsD

Text To Dialogue Streaming With Timestamps Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
inputsYesA list of dialogue inputs, each containing text and a voice ID which will be converted into speech. The maximum number of unique voice IDs is 10. For reliable generation, keep the total character count across all `inputs[].text` values at or below 2,000 characters per request. Longer requests can te
model_idNoIdentifier of the model that will be used, you can query them using GET /v1/models. The model needs to have support for text to speech, you can check this using the can_do_text_to_speech property.
settingsNo
language_codeNo
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
enable_loggingNoWhen enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.
apply_text_normalizationNoThis parameter controls text normalization with three modes: 'auto', 'on', and 'off'. When set to 'auto', the system will automatically decide whether to apply text normalization (e.g., spelling out numbers). With 'on', text normalization will always be applied, while with 'off', it will be skipped.
pronunciation_dictionary_locatorsNo

TDQS

D1.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag non-readonly, non-idempotent, open-world behavior; the description adds that the call consumes ElevenLabs credits, which is a useful cost/side-effect note beyond annotations. It does not describe streaming output or other operational traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is extremely short, but it is a sentence fragment that mostly restates the tool name; the only substantive content is the credit warning. It is not front-loaded with actionable guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter streaming generation tool with no output schema and a sparse schema description, the definition is woefully incomplete. It does not explain usage, parameter behavior, or differentiation from closely named siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 56% and several parameters (seed, settings, language_code, pronunciation_dictionary_locators) are undocumented. The description provides no parameter semantics at all, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is essentially the tool name plus a credit warning; it does not state a distinct verb or scope relative to text_to_dialogue_stream or text_to_dialogue_full_with_timestamps. An agent cannot tell from the text why streaming-with-timestamps is needed over siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no alternatives named, no prerequisites. The many sibling text-to-dialogue tools are not mentioned, so selection is left entirely to the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_speech_fullB

Text To Speech Spends ElevenLabs credits. Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
textYesThe text that will get converted into speech.
model_idNoIdentifier of the model that will be used, you can query them using GET /v1/models. The model needs to have support for text to speech, you can check this using the can_do_text_to_speech property.
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
next_textNo
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
language_codeNo
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM and WAV formats with 44.1kHz sample rate requires you to be
previous_textNo
enable_loggingNoWhen enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.
use_pvc_as_ivcNoIf true, we won't use PVC version of the voice for the generation but the IVC version. This is a temporary workaround for higher latency in PVC versions.
voice_settingsNo
next_request_idsNo
previous_request_idsNo
apply_text_normalizationNoThis parameter controls text normalization with three modes: 'auto', 'on', and 'off'. When set to 'auto', the system will automatically decide whether to apply text normalization (e.g., spelling out numbers). With 'on', text normalization will always be applied, while with 'off', it will be skipped.
optimize_streaming_latencyNoYou can turn on latency optimizations at some cost of quality. The best possible final latency varies by model. Possible values: 0 - default mode (no latency optimizations) 1 - normal latency optimizations (about 50% of possible latency improvement of option 3) 2 - strong latency optimizations (abou
apply_language_text_normalizationNoThis parameter controls language text normalization. This helps with proper pronunciation of text in some supported languages. WARNING: This parameter can heavily increase the latency of the request. Currently only supported for Japanese.
pronunciation_dictionary_locatorsNo

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false). The description adds genuinely non-structured context: the call spends ElevenLabs credits (a cost side effect) and returns audio/mpeg bytes. It stops short of noting that omitting output_path falls back to inline base64 with size limits, which the schema covers but the description does not reinforce.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler and the cost warning front-loaded. The first sentence reads slightly awkwardly ('Text To Speech Spends ElevenLabs credits') but wastes nothing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 18-parameter generation tool with no output schema, the description covers the return type and the save-to-disk path, which is the minimum an agent needs. It omits the credit-cost magnitude, model/voice prerequisites, and the inline-base64 fallback behavior, leaving notable gaps for a tool this complex.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 56% across 18 parameters, and the description adds meaning for just one of them (output_path), which the schema already documents in more detail. The many undocumented parameters such as next_text, previous_text, seed, and pronunciation_dictionary_locators get no help from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (text to speech) and adds the key differentiator that this returns audio/mpeg bytes rather than a stream. However, it never distinguishes itself from close siblings such as text_to_speech_stream, text_to_speech_full_with_timestamps, or text_to_voice, so an agent still has to guess why 'full' is the right choice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes that passing output_path saves bytes, which is a usage hint, but gives no when-to-use or when-not-to-use guidance. It never says whether to prefer this over the streaming or timestamped variants, which are the obvious alternatives in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_speech_full_with_timestampsD

Text To Speech With Timestamps Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
textYesThe text that will get converted into speech.
model_idNoIdentifier of the model that will be used, you can query them using GET /v1/models. The model needs to have support for text to speech, you can check this using the can_do_text_to_speech property.
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
next_textNo
language_codeNo
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM and WAV formats with 44.1kHz sample rate requires you to be
previous_textNo
enable_loggingNoWhen enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.
use_pvc_as_ivcNoIf true, we won't use PVC version of the voice for the generation but the IVC version. This is a temporary workaround for higher latency in PVC versions.
voice_settingsNo
next_request_idsNoA list of request_id of the samples that come after this generation. next_request_ids is especially useful for maintaining the speech's continuity when regenerating a sample that has had some audio quality issues. For example, if you have generated 3 speech clips, and you want to improve clip 2, pas
previous_request_idsNoA list of request_id of the samples that were generated before this generation. Can be used to improve the speech's continuity when splitting up a large task into multiple requests. The results will be best when the same model is used across the generations. In case both previous_text and previous_r
apply_text_normalizationNoThis parameter controls text normalization with three modes: 'auto', 'on', and 'off'. When set to 'auto', the system will automatically decide whether to apply text normalization (e.g., spelling out numbers). With 'on', text normalization will always be applied, while with 'off', it will be skipped.
optimize_streaming_latencyNoYou can turn on latency optimizations at some cost of quality. The best possible final latency varies by model. Possible values: 0 - default mode (no latency optimizations) 1 - normal latency optimizations (about 50% of possible latency improvement of option 3) 2 - strong latency optimizations (abou
apply_language_text_normalizationNoThis parameter controls language text normalization. This helps with proper pronunciation of text in some supported languages. WARNING: This parameter can heavily increase the latency of the request. Currently only supported for Japanese.
pronunciation_dictionary_locatorsNoA list of pronunciation dictionary locators (id, version_id) to be applied to the text. They will be applied in order. You may have up to 3 locators per request

TDQS

D1.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose readOnlyHint=false, idempotentHint=false and openWorldHint=true. The description earns credit for adding that the call consumes ElevenLabs credits, a cost/rate-limit fact annotations do not convey, but it says nothing about the generation being non-reversible or the timestamp payload.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Only a single ungrammatical sentence fragment, and it is under-specified rather than concise. It is not so much wasteful as absent, which is a different structural failure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter, non-idempotent, credit-consuming generation tool with no output schema, one fragment is wholly inadequate. Nothing tells the agent how timestamps are returned or what a successful call produces.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 17 parameters and 71% schema description coverage, the schema does most of the work, but the description contributes zero parameter meaning. Nothing explains required voice_id/text, the timestamp-relevant options, or the continuity parameters (previous/next_request_ids).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool's own name ('Text To Speech With Timestamps') without a distinct verb or scoping statement. It adds only a credit-cost note, so an agent learns nothing about what differentiates it from text_to_speech_full or text_to_speech_stream_with_timestamps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative-tool guidance is given, despite several closely named siblings (text_to_speech_full, text_to_speech_stream, text_to_speech_stream_with_timestamps). The credit remark hints at cost but does not route the agent to any alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_speech_streamB

Text To Speech Streaming Spends ElevenLabs credits. Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
textYesThe text that will get converted into speech.
model_idNoIdentifier of the model that will be used, you can query them using GET /v1/models. The model needs to have support for text to speech, you can check this using the can_do_text_to_speech property.
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
next_textNo
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
language_codeNo
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
previous_textNo
enable_loggingNoWhen enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.
use_pvc_as_ivcNoIf true, we won't use PVC version of the voice for the generation but the IVC version. This is a temporary workaround for higher latency in PVC versions.
voice_settingsNo
next_request_idsNo
previous_request_idsNo
apply_text_normalizationNoThis parameter controls text normalization with three modes: 'auto', 'on', and 'off'. When set to 'auto', the system will automatically decide whether to apply text normalization (e.g., spelling out numbers). With 'on', text normalization will always be applied, while with 'off', it will be skipped.
optimize_streaming_latencyNoYou can turn on latency optimizations at some cost of quality. The best possible final latency varies by model. Possible values: 0 - default mode (no latency optimizations) 1 - normal latency optimizations (about 50% of possible latency improvement of option 3) 2 - strong latency optimizations (abou
apply_language_text_normalizationNoThis parameter controls language text normalization. This helps with proper pronunciation of text in some supported languages. WARNING: This parameter can heavily increase the latency of the request. Currently only supported for Japanese.
pronunciation_dictionary_locatorsNo

TDQS

B3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Given readOnlyHint=false and no output schema, the description usefully discloses two behavioral facts the annotations don't: it consumes ElevenLabs credits (a cost/side-effect) and it returns audio/mpeg bytes that can be persisted via output_path. It does not explain streaming semantics (chunking, latency, partial delivery), which is the one notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the cost warning and then the return-handling instruction. Slightly wasteful in that the opening phrase just echoes the tool title, but nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 18-parameter generation tool with no output schema, the description covers cost and return handling but leaves streaming behavior and most parameters unexplained. It is adequate for a basic call but incomplete given the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 18 parameters at only 56% schema coverage, roughly half the parameters (seed, next_text, language_code, previous_text, next/previous_request_ids, pronunciation_dictionary_locators) are undocumented in both schema and description. The description only touches output_path, which the schema already explains in more detail, so it does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description largely restates the tool name ("Text To Speech Streaming") and only adds the credit-spend fact. It does not differentiate from close siblings like text_to_speech_full or text_to_speech_stream_with_timestamps, so an agent cannot tell from the text alone why it would pick this over the full or timestamped variants.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus text_to_speech_full, text_to_speech_stream_with_timestamps, or text_to_dialogue_stream. The only routing-ish hint is that output_path can be passed, which is a parameter detail rather than a usage rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_speech_stream_with_timestampsC

Text To Speech Streaming With Timestamps Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
textYesThe text that will get converted into speech.
model_idNoIdentifier of the model that will be used, you can query them using GET /v1/models. The model needs to have support for text to speech, you can check this using the can_do_text_to_speech property.
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
next_textNo
language_codeNo
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
previous_textNo
enable_loggingNoWhen enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.
use_pvc_as_ivcNoIf true, we won't use PVC version of the voice for the generation but the IVC version. This is a temporary workaround for higher latency in PVC versions.
voice_settingsNo
next_request_idsNo
previous_request_idsNo
apply_text_normalizationNoThis parameter controls text normalization with three modes: 'auto', 'on', and 'off'. When set to 'auto', the system will automatically decide whether to apply text normalization (e.g., spelling out numbers). With 'on', text normalization will always be applied, while with 'off', it will be skipped.
optimize_streaming_latencyNoYou can turn on latency optimizations at some cost of quality. The best possible final latency varies by model. Possible values: 0 - default mode (no latency optimizations) 1 - normal latency optimizations (about 50% of possible latency improvement of option 3) 2 - strong latency optimizations (abou
apply_language_text_normalizationNoThis parameter controls language text normalization. This helps with proper pronunciation of text in some supported languages. WARNING: This parameter can heavily increase the latency of the request. Currently only supported for Japanese.
pronunciation_dictionary_locatorsNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is known. The only behavioral fact added is 'Spends ElevenLabs credits', which is useful cost context, but the description says nothing about streaming semantics, timestamp delivery, partial output, or retention behavior that would matter for this specific variant.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence with no filler and the credit-cost note is front-loaded alongside the name. However, it is arguably too terse rather than concise-structuring useful content, and the sentence is essentially a name sandwich rather than a structured description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter, streaming, credit-consuming operation with no output schema, the description is insufficient. It omits streaming behavior, timestamp availability, tier/plan constraints mentioned in parameters (e.g., Creator tier for high-bitrate MP3), and return shape, while annotations cover only safety hints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 53%, so several parameters (seed, next_text, previous_text, language_code, previous/next_request_ids, pronunciation_dictionary_locators) are undocumented in both places. The description supplies no parameter guidance at all, so it neither compensates for the gap nor exceeds the baseline for the documented half.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description only restates the tool name ('Text To Speech Streaming With Timestamps') and appends a billing note. It conveys the core action (TTS streaming with timestamps) but adds no verb+resource detail beyond the title, and does not distinguish this streaming variant from text_to_speech_stream or text_to_speech_full_with_timestamps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance. The description does not tell the agent when to choose streaming-with-timestamps over streaming-without or the full-with-timestamps sibling, despite the tool being one of four near-identical TTS variants in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_voiceC

[Deprecated] Generate A Voice Preview From Description Spends ElevenLabs credits. Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
textNo
qualityNoHigher quality results in better voice output but less variety.
loudnessNoControls the volume level of the generated voice. -1 is quietest, 1 is loudest, 0 corresponds to roughly -24 LUFS.
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
guidance_scaleNoControls how closely the AI follows the prompt. Lower numbers give the AI more freedom to be creative, while higher numbers force it to stick more to the prompt. High numbers can cause voice to sound artificial or robotic. We recommend to use longer, more detailed prompts at lower Guidance Scale.
should_enhanceNoWhether to enhance the voice description using AI to add more detail and improve voice generation quality. When enabled, the system will automatically expand simple prompts into more detailed voice descriptions. Defaults to False
voice_descriptionYesDescription to use for the created voice.
auto_generate_textNoWhether to automatically generate a text suitable for the voice description.

TDQS

C2.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose readOnlyHint=false, idempotentHint=false and openWorldHint=true, so the safety profile is covered. The description adds a genuinely useful trait not in the annotations: the call spends ElevenLabs credits, and it is deprecated upstream, both of which materially affect tool selection. It still omits whether a persistent voice is created versus a throwaway preview.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loads the deprecation notice, which is the right priority. But the sentence runs together into an ambiguous fragment ('Generate A Voice Preview From Description Spends ElevenLabs credits'), making the cost statement read as part of the purpose rather than a separate fact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, credit-consuming generation tool with no output schema and no return-value description, the definition is thin. It never clarifies that the result is an audio preview rather than a saved voice, nor does it quantify the credit cost or point at a replacement, leaving key decision inputs missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 78%, close to the high-coverage baseline, so the schema already carries most parameter meaning (quality, loudness, guidance_scale, output_format, should_enhance are documented inline). The description adds no parameter detail whatsoever, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete verb and resource ('Generate A Voice Preview From Description'), which is more specific than a tautology. However, it does not distinguish this tool from close siblings such as text_to_voice_design, text_to_voice_remix, or text_to_voice_preview_stream, so an agent cannot tell which voice-generation variant to pick from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The '[Deprecated]' and 'Deprecated upstream' markers signal that this tool should be avoided, which is weak negative guidance, but no alternative is named and no condition for legitimate use is given. An agent is left to infer that a sibling deprecated-adjacent tool should be used instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_voice_designC

Design A Voice. Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
textNo
qualityNo
loudnessNoControls the volume level of the generated voice. -1 is quietest, 1 is loudest, 0 corresponds to roughly -24 LUFS.
model_idNoModel to use for the voice generation. Possible values: eleven_multilingual_ttv_v2, eleven_ttv_v3.
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
guidance_scaleNoControls how closely the AI follows the prompt. Lower numbers give the AI more freedom to be creative, while higher numbers force it to stick more to the prompt. High numbers can cause voice to sound artificial or robotic. We recommend to use longer, more detailed prompts at lower Guidance Scale.
should_enhanceNoWhether to enhance the voice description using AI to add more detail and improve voice generation quality. When enabled, the system will automatically expand simple prompts into more detailed voice descriptions. Defaults to False
prompt_strengthNo
stream_previewsNoDetermines whether the Text to Voice previews should be included in the response. If true, only the generated IDs will be returned which can then be streamed via the /v1/text-to-voice/:generated_voice_id/stream endpoint.
voice_descriptionYesDescription to use for the created voice.
auto_generate_textNoWhether to automatically generate a text suitable for the voice description.
remixing_session_idNo
reference_audio_base64No
remixing_session_iteration_idNo

TDQS

C2.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-read-only, non-idempotent, open-world behavior, so mutation and repeat-call semantics are covered. The description does add one genuine behavioral fact beyond annotations — that the call consumes ElevenLabs credits — which is real value for a billable operation. However, it says nothing about whether design is reversible, whether permission tiers matter, or what is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded, with no filler — the credit-cost warning is the one sentence that genuinely earns its place. But the extreme brevity is under-specification rather than conciseness for a 15-parameter, credit-consuming generation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter, non-idempotent, credit-spending voice-design tool with no output schema and only 53% schema coverage, this description is far too thin. It omits required inputs, model/format choices, preview-streaming behavior, and cost magnitude — none of which the structured fields fully carry.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema documents only about 53% of the 15 parameters, so the description would need to compensate for the remaining gap — yet it mentions no parameter at all, not even the required voice_description. Parameters like guidance_scale, prompt_strength, and stream_previews are left to be inferred from partial schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Design A Voice" is essentially a restatement of the tool name text_to_voice_design plus the title "Text To Voice Design" — it adds no specific verb+resource detail beyond what the identifier already conveys. It also fails to distinguish this tool from close siblings such as text_to_voice, text_to_voice_remix, text_to_voice_preview_stream, or create_voice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance and no alternatives: an agent cannot tell from this text why it would call text_to_voice_design rather than text_to_voice or text_to_voice_remix. The only implicit cue is cost awareness, which is not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_voice_preview_streamC
Read-onlyIdempotent

Text To Voice Preview Streaming Returns audio/mpeg bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
generated_voice_idYesThe generated_voice_id to stream.

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description usefully discloses the concrete return payload (audio/mpeg bytes) and that output_path persists them, but says nothing about streaming behavior, size limits, or auth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no padding; the return type is stated before the optional parameter hint. It loses a point only because the leading phrase is a near-verbatim restatement of the title rather than new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly covers the return type (audio/mpeg bytes) and the save-vs-inline tradeoff. However, for a streaming preview tool it never explains what the preview represents or when an agent should call it rather than the non-streaming voice-preview siblings, leaving the definition merely adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already explains output_path resolution (ELEVENLABS_OUTPUT_DIR, inline base64 fallback) and generated_voice_id. The description only echoes the output_path behavior, adding no meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name ('Text To Voice Preview Streaming') and adds only that it returns audio/mpeg bytes. A reader can infer it streams preview audio, but the description never says what is being previewed or how this differs from the many sibling text_to_voice*/text_to_speech* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all — no condition, no alternative, no prerequisite such as needing a previously created generated_voice_id. With ~15 similar audio-synthesis siblings, the absence of routing guidance is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_voice_remixC

Remix A Voice. Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
textNo
loudnessNoControls the volume level of the generated voice. -1 is quietest, 1 is loudest, 0 corresponds to roughly -24 LUFS.
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
guidance_scaleNoControls how closely the AI follows the prompt. Lower numbers give the AI more freedom to be creative, while higher numbers force it to stick more to the prompt. High numbers can cause voice to sound artificial or robotic. We recommend to use longer, more detailed prompts at lower Guidance Scale.
prompt_strengthNo
stream_previewsNoDetermines whether the Text to Voice previews should be included in the response. If true, only the generated IDs will be returned which can then be streamed via the /v1/text-to-voice/:generated_voice_id/stream endpoint.
voice_descriptionYesDescription of the changes to make to the voice.
auto_generate_textNoWhether to automatically generate a text suitable for the voice description.
remixing_session_idNo
remixing_session_iteration_idNo

TDQS

C2.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true. The description adds the useful behavioral fact that the operation spends ElevenLabs credits, but says nothing about required permissions, rate limits, or what happens to the original voice.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The two short sentences are concise and front-load a rough purpose and a cost warning, but the first sentence adds little beyond the name, and the overall structure is too sparse to guide correct invocation of a 12-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutation tool with 12 parameters, no output schema, and only partial schema descriptions, the description is severely incomplete. It omits usage context, parameter semantics, and nearly all behavioral details beyond a generic cost note.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description provides no parameter-level guidance despite 12 parameters and only 58% schema description coverage. Several nullable parameters (seed, text, prompt_strength, remixing_session_id, remixing_session_iteration_id) are undocumented in both the schema and the description, so the description fails to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Remix A Voice' restates the tool name almost verbatim, and adds only a billing note ('Spends ElevenLabs credits'). It does not explain what remixing a voice actually does or how it differs from siblings like text_to_voice, text_to_voice_design, or text_to_voice_preview_stream.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool instead of alternatives such as text_to_voice or text_to_voice_design. The only usage-relevant information is the cost warning, which is a constraint but not selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transcribeC

Transcribes Segments Spends ElevenLabs credits. Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
segmentsYesTranscribe this specific list of segments.
dubbing_idYesID of the dubbing project.

TDQS

C2.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare the safety profile (readOnly=false, destructive=false, idempotent=false, openWorld=true), so the bar is lower, and the description still adds two things structured fields cannot: it consumes ElevenLabs credits and it is deprecated upstream. Both are material to a calling agent weighing cost and lifecycle. It stops short of saying what the transcription returns or what happens to the segment list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The payload is short and the cost/deprecation facts are front-loaded, which is good. However the two sentences are malformed and jammed together ("Segments Spends"), which slows parsing rather than helping it. Brevity is not the problem; clarity is.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the return-value burden, and it says nothing about what transcription output the caller gets back. It does at least flag credit consumption and deprecation, which are the two most decision-relevant facts for this mutating tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — both dubbing_id and segments are documented in the schema, including that segments is the specific list to transcribe. The description adds nothing beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description reads as two fragments run together ("Transcribes Segments Spends ElevenLabs credits. Deprecated upstream."), so the stated purpose is little more than the tool name plus a parameter name. It never distinguishes this from the many sibling transcription tools (speech_to_text, dubbing_target_transcript_get, dubbing_target_transcript_regenerate), leaving the agent to guess which transcription path applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Deprecated upstream" is a useful steer away from the tool, but no positive when-to-use guidance or alternative is named. With dozens of transcription/dubbing transcript siblings, the agent gets no help deciding whether to call this at all versus a non-deprecated equivalent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

translateB

Translates All Or Some Segments And Languages Spends ElevenLabs credits. Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
segmentsYesTranslate only this list of segments.
languagesYes
dubbing_idYesID of the dubbing project.

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety hints, and the description adds two important behaviors beyond them: it spends ElevenLabs credits and is deprecated upstream. This is useful cost/deprecation context, though mutation specifics remain unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short and front-loaded, with both sentences carrying useful information. The first sentence is grammatically awkward and lacks punctuation, but it is still concise and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a deprecated, cost-incurring translation tool with no output schema, the description includes critical cost and deprecation warnings. However, it lacks usage alternatives and parameter detail, leaving gaps for an agent choosing among many sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is moderate at 67%, and the description does not clarify the parameters beyond loosely naming segments and languages. It does not explain the languages array format or how segments and languages interact, so it adds little beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb and resource: translating segments and languages. It does not distinguish itself from sibling dubbing/transcript tools, but the core action is identifiable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage signal is 'Deprecated upstream,' which warns against use but does not name a replacement or explain when this tool should still be used. No when-to-use or alternative guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unassign_conversation_tag_routeD
DestructiveIdempotent

Unassign Conversation Tag

ParametersJSON Schema
NameRequiredDescriptionDefault
tag_idYes
conversation_idYes

TDQS

D1.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is partly covered. However, the description adds nothing: it does not say whether the unassignment is reversible, whether it applies per-conversation or globally, or what happens if the tag was never assigned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is technically brief with no filler, but this is under-specification rather than conciseness. A single noun phrase cannot front-load anything useful for a two-parameter mutation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A destructive, non-idempotent-obvious mutation with two undocumented required parameters and no output schema leaves the agent with almost nothing to reason about. The description should at minimum state effect, scope, and required identifiers.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both required parameters (conversation_id, tag_id), so the description must compensate and does not. The parameter names are inferable from the tool name, but no format, ID conventions, or ambiguity resolution is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Unassign Conversation Tag' effectively restates the tool name and adds no verb scope, target semantics, or differentiation from siblings like assign_conversation_tags_route or delete_conversation_tag_route. It is a tautology rather than an explanation of what unassigning accomplishes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus assign_conversation_tags_route, update_conversation_tag_route, or delete_conversation_tag_route. No preconditions, no alternatives, no exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unshare_resource_endpointC

Unshare Workspace Resource

ParametersJSON Schema
NameRequiredDescriptionDefault
group_idNo
user_emailNo
resource_idYesThe ID of the target resource.
resource_typeYesResource types that can be shared in the workspace. The name always need to match the collection names
workspace_api_key_idNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, openWorldHint=true, and destructiveHint=false, so the agent knows this is a non-idempotent write. The description adds nothing beyond that: it does not say whether unsharing is reversible, whether it requires ownership/admin permission, or whether the target must be a user or a group. For an access-revocation mutation, that context is valuable and absent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words is concise but under-specified rather than efficient; the single fragment carries almost no information and there is no structure beyond a restated title. It reads as a placeholder rather than a front-loaded, purposeful first sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter mutation tool with no output schema and no usage guidance, the description is far too thin. An agent cannot tell which parameters are required to scope the unshare, what resource types are valid, or what the result of the call is.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 40%: resource_id and resource_type are documented in the schema, but group_id, user_email, and workspace_api_key_id are not. The description supplies no parameter meaning at all, so it fails to compensate for the three undocumented parameters and does not clarify the user_email-vs-group_id targeting choice.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb and resource ('Unshare Workspace Resource'), so an agent knows it reverses sharing. However, it is essentially a restatement of the tool name/title and gives no differentiation from the sibling share_resource_endpoint, nor any clue about what 'unsharing' actually does (revoking a user's or group's access to a specific resource_type).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the obvious alternative share_resource_endpoint or any related sibling. The agent gets no signal about when this tool should be chosen over other access-management routes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_agent_conversation_ticket_routeD

Update Agent Conversation Ticket

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNo
assignee_user_idNo
agentqa_ticket_idYes

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true. The description adds no further behavioral context such as which fields are mutable, permission requirements, or expected effects, so it does not enrich the annotation-provided safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but this is under-specification rather than useful conciseness. It does not front-load actionable information because there is none to front-load.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with three parameters, 0% schema description coverage, and no output schema, the description is completely inadequate. It omits what can be updated, which fields are optional, and any prerequisites or side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description mentions none of the three parameters (agentqa_ticket_id, status, assignee_user_id). With low coverage, the description is responsible for compensating, but it adds no parameter meaning at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update Agent Conversation Ticket' simply humanizes the tool name and does not specify what aspects of the ticket are updated or how it differs from siblings such as update_agent_response_test_route or create_agent_conversation_ticket_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. The description gives no context, prerequisites, or exclusions, matching the 'no guidance' level.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_agent_response_test_routeD
Idempotent

Update Agent Response Test

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
typeNo
test_idYesThe id of a chat response test. This is returned on test creation.
environmentNo
chat_historyNo
evaluation_modelNo
failure_examplesNoNon-empty list of example responses that should be considered failures
parent_folder_idNo
success_examplesNoNon-empty list of example responses that should be considered successful
tool_mock_configNoSimulation/preview-side config: tools are identified by IDs, resolved to names at runtime.
dynamic_variablesNoDynamic variables to replace in the agent config during testing
success_conditionNo
success_conditionsNoList of prompts that evaluate whether the simulation was successful. If provided, all criteria are evaluated and merged into a final result. Capped at the maximum number of evaluation criteria.
simulation_scenarioNoDescription of the simulation scenario and user persona for simulation tests.
tool_mock_overridesNoTest-specific response mocks, keyed by tool ID. Applied ahead of the tool's shared mocks and only within this test. Only take effect for tools that are mocked (see tool_mock_config).
simulated_user_modelNo
simulation_max_turnsNoMaximum number of conversation turns for simulation tests.
tool_call_parametersNo
check_any_tool_matchesNo
simulation_environmentNo
from_conversation_metadataNo
conversation_initiation_sourceNoEnum representing the possible sources for conversation initiation.

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare a non-read-only, idempotent, non-destructive, open-world mutation, and the description adds nothing on top — no note on which fields are replaced vs preserved, no auth or permission context, no side effects. The description carries zero behavioral value beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but this is under-specification rather than conciseness — a single noun phrase that omits everything the agent needs before invoking a 22-parameter mutation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutation with 22 params, nested objects, and no output schema, the definition is completely inadequate. An agent cannot determine what the tool changes, what the required test_id refers to, or what happens to unspecified fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 22 parameters and only 45% schema description coverage, the description was the place to explain fields like success_conditions, tool_mock_overrides, or from_conversation_metadata, but it names none of them. Nothing compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is essentially a restatement of the tool name: 'Update Agent Response Test'. It implies a verb and resource but adds no scope, no indication of what fields can be updated, and no differentiation from siblings like create_agent_response_test_route, get_agent_response_test_route, or delete_chat_response_test_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites, and no reference to alternative tools. An agent has nothing to distinguish this update route from the create/get/delete/run test routes in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_agent_test_folder_routeC

Update Agent Test Folder

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe new name for the folder
folder_idYesThe folder ID.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that — it doesn't state required permissions, failure modes, or what happens to the folder on rename.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short phrase with no waste, but it is under-specified rather than genuinely concise — there is no front-loaded detail because there is no detail at all.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a mutation tool with no output schema, and the description omits any behavioral context: what the update entails, whether renaming is reversible, or what the response contains. For a write operation it should do considerably more.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (name, folder_id) are documented in the schema. The description contributes no additional parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description merely restates the tool name ('Update Agent Test Folder') — verb and resource are present but there is no further specification of what is being updated or how it differs from siblings like create_agent_test_folder_route or delete_agent_test_folder_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives. An agent must infer from the name alone that this renames an existing test folder rather than creating or deleting one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_auth_connectionD

Update Workspace Auth Connection

ParametersJSON Schema
NameRequiredDescriptionDefault
tokenNo
issuerNo
key_idNo
scopesNo
subjectNo
audienceNo
passwordNo
providerNo
usernameNo
algorithmNo
auth_typeNo
client_idNo
secret_keyNo
extra_paramsNo
client_secretNo
custom_headersNo
auth_connection_idYes
expiration_secondsNo
basic_auth_in_headerNo
token_response_fieldNo

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already state readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is partially known. However the description adds no behavioral context beyond that—it does not clarify whether this is a partial or full update, what happens to unspecified fields, whether secrets are re-encrypted, or what permissions are required. For a 20-param mutation touching tokens and client secrets, this is a significant omission.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase that is front-loaded only in the sense of being brief. It is under-specified rather than appropriately concise: for a complex 20-parameter mutation, the reader gets no actionable detail. The brevity does not earn its place because the content is essentially a title repeated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a 20-parameter mutation with no output schema and only partial annotation coverage, the description is grossly incomplete. It does not explain the resource, the update semantics, required permissions, side effects, or any parameter meaning. An agent would struggle to invoke this tool correctly based on the definition alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema contains 20 parameters with 0% description coverage, and the description supplies no parameter information whatsoever. The agent receives no explanation of what can be updated, which fields are optional versus required beyond the schema's required list, or how sensitive fields like token, secret_key, client_secret, and password are handled. The description completely fails to compensate for the schema's lack of documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update Workspace Auth Connection' is a verbatim expansion of the tool name and annotation title. It restates what the name already says without adding any differentiating scope, fields, or context relative to siblings like create_auth_connection or delete_auth_connection. This is a tautology per the rubric.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no alternatives mentioned. An agent cannot tell from this description alone whether to use this tool versus update_secret_route, update_environment_variable, or other update tools. The definition provides zero routing help for a mutation with 20 parameters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_branch_routeD

Update Agent Branch

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
agent_idYesThe id of an agent. This is returned on agent creation.
branch_idYesUnique identifier for the branch.
is_archivedNo
protection_statusNo

TDQS

D1.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=false, openWorldHint=true, destructiveHint=false, but the description adds nothing. It does not disclose which updates are irreversible, how protection_status affects later operations, or that renewal of protection perms may be required. For a mutation tool this is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words that merely echo the tool name. It is not overlong, but it is under-specified rather than concise, so no sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutation tool with 5 parameters (40% described), no output schema, and open-world/non-idempotent annotations demands far more than a restated title. Nothing in the description helps an agent call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 40%, so the description is expected to compensate, and it does not. The enum-valued protection_status (writer_perms_required vs admin_perms_required) and the is_archived/name semantics are undocumented in the schema and unexplained in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Update Agent Branch" restates the name and title rather than describing the operation. It does not say what fields can be changed (name, is_archived, protection_status) or what an agent branch is in this system, and it has no differentiation against siblings like create_branch_route, get_branch_route, or rebase_branch_onto_main.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance whatsoever. Nothing tells the agent when to update a branch vs. rebase it, merge it, or read it, even though several sibling tools touch the same branch resource.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_conversation_tag_routeD

Update Conversation Tag

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNo
tag_idYes
descriptionNo

TDQS

D1.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose that this is a non-read-only, non-idempotent, non-destructive, open-world operation. The description adds nothing beyond "update" and does not state permissions needed, what happens to unspecified fields, whether changes are reversible, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence, so it is not verbose, but it is under-specified rather than usefully concise. The description does not front-load any actionable detail beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, 0% parameter description coverage, and only basic annotations, the description is wholly inadequate. An agent cannot determine required inputs, expected behavior, or how to distinguish this operation from sibling tag tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention any of the three parameters. It does not clarify that tag_id is required, or what title and description represent in an update context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Update Conversation Tag" states a verb and resource, but it essentially restates the tool name and title. It provides no scope or differentiation from siblings such as create_conversation_tag_route, delete_conversation_tag_route, or get_conversation_tag_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no indication of when to use this tool versus create, delete, get, list, assign, or unassign conversation tag tools. There are no conditions, prerequisites, or alternative-selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_dashboard_settings_routeC

Update Convai Dashboard Settings

ParametersJSON Schema
NameRequiredDescriptionDefault
chartsNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already state readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description only restates that an update occurs and adds nothing about permissions, side effects, or what happens to existing dashboard settings.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. However, it is under-specified for a mutation tool with an undocumented payload parameter, so its brevity is not fully appropriate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a mutation tool with one nested-array parameter, zero schema descriptions, and no output schema, the definition is incomplete. An agent cannot determine what 'charts' should contain or how the update behaves beyond the high-level operation name.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter 'charts' is undocumented in both the schema and the description. The description never mentions charts, their format, or what updating them means, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: update Convai Dashboard Settings. It distinguishes the operation from the sibling read tool get_dashboard_settings_route, but does not differentiate it from the broader update_settings_route or clarify what dashboard settings are included.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives. It does not mention prerequisites, exclusions, or the sibling update_settings_route, leaving the agent to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_document_routeC

Update Document

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
contentNo
documentation_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a non-read-only, non-idempotent, open-world, non-destructive mutation. The description adds no behavioral context beyond those hints, such as what fields are affected, authorization requirements, or side effects. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is technically concise, but it is under-specified rather than appropriately sized. For a mutation tool with three parameters, two words do not provide enough front-loaded structure to be useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a mutation tool with incomplete parameter documentation, no output schema, and many sibling tools, this description omits essential context about purpose, usage, affected fields, and expected behavior. It is not complete enough for an agent to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, with just documentation_id documented. The description 'Update Document' adds no meaning for the name, content, or documentation_id parameters, so it fails to compensate for the sparse schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb and resource ('Update Document'), but it does not clarify the scope: whether this is a knowledge-base document, a file document, or another document type. It also does not distinguish this tool from nearby siblings such as update_file_document_route or refresh_url_document_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites, and no mention of related update/list/delete document tools. The description gives no routing or selection context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_environment_variableD

Update Environment Variable

ParametersJSON Schema
NameRequiredDescriptionDefault
valuesYesValues to replace. Set to null to remove an environment (except 'production').
env_var_idYes

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond the title: no permissions, no effect on unspecified values, no reversibility, no rate-limit or auth notes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single sentence, but the issue is under-specification rather than concision. Nothing is front-loaded because nothing useful is stated beyond the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with a nested replacement object, no output schema, and only half of the schema parameters described, the definition omits nearly all operational context an agent would need to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: the values object has a schema description, but env_var_id has none. The description provides no parameter meaning at all, so it does not compensate for the undocumented required parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is only 'Update Environment Variable', identical to the tool name/title. It restates rather than specifies purpose; there is no scope, no indication of what is updated, and no differentiation from siblings such as create_environment_variable or get_environment_variable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or alternatives are provided. The description does not tell the agent when to choose update_environment_variable over create_environment_variable or list_environment_variables.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_file_document_routeC

Update File Document

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathNoDocumentation that the agent will have access to in order to interact with users. Local path. Required for this call.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.
documentation_idYesThe id of a document from the knowledge base. This is returned on document addition.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, covering the basic safety profile. The description adds nothing beyond that: it does not say what is replaced, whether the prior file is destroyed, what permissions are needed, or what the response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is extremely short and front-loaded, but the brevity reflects under-specification rather than efficient communication. There is no wasted sentence because there is effectively no content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent, open-world mutation tool with no output schema, the description omits what the update does, what inputs are required versus optional in practice, and what the caller gets back. Only the 100% schema coverage keeps this from being entirely inadequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with 4 documented parameters, so the schema carries parameter meaning on its own. The description adds no additional parameter semantics, which is the expected baseline of 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is essentially a restatement of the tool name ('Update File Document'), with no verb+resource detail beyond the title. It does not distinguish this from sibling tools such as update_document_route or create_file_document_route, so an agent cannot tell what resource is being updated or how it differs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives among the many sibling update/create document tools. The agent receives no signal about when this tool is the correct choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_finetuneC

Update Music Finetune Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
tagsNo
visibilityNo
finetune_idYes
primary_genreNo

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is largely covered. The description adds a meaningful cost trait by noting that the operation spends ElevenLabs credits. However, it does not explain what mutations are allowed, permission requirements, or other side effects beyond that credit consumption.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is very short and front-loads the core action, so it is not verbose. But the sentence is malformed and lacks clear separation between the action and the credit-spending note, which weakens structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an update tool with five parameters, no schema descriptions, and no output schema, the description is too thin. It gives the purpose and a credit-cost hint but leaves parameter meanings, usage conditions, and mutation details completely unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has five parameters with 0% description coverage, and the description provides no information about any of them. It does not clarify what finetune_id identifies, which fields can be updated, or what values like visibility, tags, or primary_genre mean.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Update Music Finetune.' That is enough to distinguish it from create_finetune, get_finetune, and delete_finetune, though the sentence itself does not explicitly say so. The trailing 'Spends ElevenLabs credits' is an odd addition but does not obscure the core action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives such as create_finetune or get_finetune. It implies updating an existing finetune, but never states prerequisites or when the operation is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_mcp_server_approval_policy_routeC

Update Mcp Server Approval Policy Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
mcp_server_idYesID of the MCP Server.
approval_policyYesDefines the MCP server-level approval policy for tool execution.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds only the cryptic 'Deprecated upstream' note, which raises a behavioral question (is this tool still supported?) without answering it, so the added value is minimal and ambiguous.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loads the action, which is good. However, the dangling sentence fragment 'Deprecated upstream' is not a well-formed statement and costs clarity rather than earning its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter mutation with complete schema coverage, the description still omits essential context: prerequisites, what the policy change affects, and, most importantly, what 'deprecated upstream' means for the caller. The deprecation note is raised but never explained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both mcp_server_id and the enumerated approval_policy are fully documented in the schema, including the three policy values. The description adds nothing beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Update Mcp Server Approval Policy' names a verb and resource, so an agent can guess the operation. But the trailing fragment 'Deprecated upstream' muddies whether this tool should be called at all, and there is no differentiation from siblings like update_mcp_server_config_route or update_mcp_tool_config_override_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given. The only contextual signal, 'Deprecated upstream,' does not tell the agent whether to call this tool, avoid it, or prefer a replacement, and no alternative is named despite several similar update_* MCP siblings existing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_mcp_server_config_routeC

Update Mcp Server Configuration

ParametersJSON Schema
NameRequiredDescriptionDefault
request_metaNo
secret_tokenNoUsed to reference a secret from the agent's secret store.
mcp_server_idYesID of the MCP Server.
execution_modeNo
approval_policyNoDefines the MCP server-level approval policy for tool execution.
auth_connectionNoOptional auth connection to use for authentication with this MCP server
pre_tool_speechNo
request_headersNo
tool_call_soundNoPredefined tool call sounds; ``None`` means no sound.
interruption_modeNo
disable_compressionNo
disable_interruptionsNo
force_pre_tool_speechNo
response_timeout_secsNo
tool_call_sound_behaviorNoDetermines how the tool call sound should be played.

TDQS

C2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint=false, idempotentHint=false, and destructiveHint=false, which tells the agent this is a non-idempotent write that is not destructive. However, the description adds nothing beyond this: it doesn't clarify what happens when fields are omitted (partial update vs full replace), whether the server must exist, or any side effects of reconfiguration. With annotations covering the safety profile, a 3 is the appropriate baseline for the description adding no extra context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, short phrase that is front-loaded but largely uninformative. It is concise but borders on the tautological, providing no additional structure or actionable detail. The brevity is not a virtue here since it sacrifices substance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 15 parameters, no output schema, no annotations beyond the basic safety hints, and only 40% schema coverage, the description is completely inadequate. It offers no information about the tool's behavior, parameter meanings, or usage context, leaving an agent unable to use the tool correctly without resorting to external knowledge.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 40%, meaning over half of the 15 parameters have no description in the schema. The tool description provides zero parameter information to compensate. Critical parameters like execution_mode, interruption_mode, disable_compression, and response_timeout_secs are entirely undocumented in both the description and the schema. The description does not help an agent understand what to pass.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update Mcp Server Configuration' is essentially a restatement of the tool name (update_mcp_server_config_route) with no added specificity about what configuration fields are affected. It does distinguish the resource (MCP server config) but gives no verb-level detail beyond the name itself. Given the many sibling tools around MCP servers (update_mcp_server_approval_policy_route, update_mcp_tool_config_override_route), this description fails to differentiate scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance whatsoever about when to use this tool versus siblings like update_mcp_server_approval_policy_route or update_mcp_tool_config_override_route. There is no mention of prerequisites, alternatives, or conditions. An agent would have to guess which MCP update tool to invoke.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_mcp_tool_config_override_routeD

Update Mcp Tool Configuration Override

ParametersJSON Schema
NameRequiredDescriptionDefault
tool_nameYesName of the MCP tool to update config overrides for.
assignmentsNo
environmentNoEnvironment whose values are used when the MCP server URL, headers, or auth connection reference environment variables. Mirrors the environment a conversation would run in; defaults to production.
mcp_server_idYesID of the MCP Server.
execution_modeNo
response_mocksNo
input_overridesNo
pre_tool_speechNo
tool_call_soundNoOverrides the server's tool_call_sound setting for this tool. A sound name plays that sound; 'off' overrides to no sound (silence); null means do not override (inherit the server default).
interruption_modeNo
disable_interruptionsNo
force_pre_tool_speechNo
response_timeout_secsNo
tool_call_sound_behaviorNoDetermines how the tool call sound should be played.

TDQS

D1.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-read-only, non-idempotent, open-world mutation, so the safety profile is covered. The description adds nothing on top: with 14 parameters exposing partial-update semantics, it doesn't say what a partial payload does, whether omitted fields are cleared or left alone, or whether the override replaces or merges with existing config.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but only because it is under-specified rather than efficient. A single restated title earns no informational value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter mutation tool with no output schema and low schema coverage, a one-line tautology is wholly inadequate. Nothing about required fields, merge semantics, or environment scoping is conveyed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 36%, so most of the 14 parameters are undocumented in both schema and description. The description contributes zero parameter meaning, leaving the agent guessing about fields like input_overrides, response_mocks, assignments, and interruption_mode.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name/title ('Update Mcp Tool Configuration Override') with no verb+resource elaboration. It tells the agent nothing beyond what the name already encodes, and offers no differentiation against siblings like add_mcp_tool_config_override_route or remove_mcp_tool_config_override_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, no mention of alternatives (add/remove/get variants exist as siblings). The agent must infer everything from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_phone_number_routeD

Update Phone Number

ParametersJSON Schema
NameRequiredDescriptionDefault
labelNo
agent_idNo
branch_idNo
environmentNo
livekit_stackNo
phone_number_idYesThe phone number ID. This is returned when a phone number is imported.
store_sip_messagesNo
inbound_trunk_configNo
outbound_trunk_configNo

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the agent knows this is a non-idempotent, open-world mutation. The description adds nothing beyond that — it does not describe what fields are mutable, reversibility, required permissions, or side effects. With annotations covering the safety profile, the bar is lower, but the description still contributes no behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At three words, the description is not verbose, but it is under-specified rather than concise. The single phrase is too sparse to earn its place as useful guidance, matching the calibration pattern where under-specification yields a 2.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a high-complexity mutation tool with nine parameters (including nested trunk configuration objects), no output schema, and only minimal annotations. The description provides none of the context an agent needs — no parameter details, no mutation side effects, no usage context — making it completely inadequate for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 11%: nine parameters exist, yet only phone_number_id has a description (in the schema), and the description text adds nothing about any parameter. Parameters like label, agent_id, branch_id, environment, livekit_stack, store_sip_messages, inbound_trunk_config, and outbound_trunk_config are undocumented in both the schema and the description, so the description fails to compensate for the very low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update Phone Number' is essentially a tautological restatement of the tool title 'Update Phone Number Route'. It names a verb and resource but provides no distinguishing detail from siblings like create_phone_number_route, delete_phone_number_route, or get_phone_number_route, so an agent cannot tell what makes this update unique.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives such as delete_phone_number_route or get_phone_number_route, and no preconditions or context are stated. The description offers zero when-to-use or when-not-to-use information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_procedure_draft_routeC

Update Procedure Draft

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesProcedure name
typeYes
contentYesProcedure content
triggerNo
agent_idYesAgent ID to get the procedure draft from
branch_idYesBranch ID to get the procedure draft from
procedure_idYesThe procedure ID

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the operation is a mutation (readOnlyHint false), non-destructive, non-idempotent, and open-world. The description adds no behavioral context beyond those annotations, such as required permissions or what specifically gets updated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short but under-specified rather than concise. It consists of a single four-word phrase that does not earn its place by providing any useful information beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with six required parameters and no output schema, the description gives no context about what is being updated or how to invoke it correctly. Annotations and schema provide some structured data, but the definition as a whole leaves significant gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 71%, covering 5 of 7 parameters, but the description contributes no parameter meaning at all. Undescribed fields like the 'type' enum and optional 'trigger' receive no clarification, so the description fails to compensate for the remaining gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update Procedure Draft' restates the tool name and title almost verbatim, making it a tautology. It names an action and resource at a high level but gives no scope, distinguishing details, or sibling differentiation to help an agent select it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus siblings like create_procedure_route, get_procedure_draft_route, or delete_procedure_draft_route. No prerequisites, context, or exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_pronunciation_dictionariesD

Create Pronunciation Dictionaries Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesThe ID of the Studio project.
invalidate_affected_textNoThis will automatically mark text in this project for reconversion when the new dictionary applies or the old one no longer does.
pronunciation_dictionary_locatorsYesA list of pronunciation dictionary locators (pronunciation_dictionary_id, version_id) encoded as a list of JSON strings for pronunciation dictionaries to be applied to the text. A list of json encoded strings is required as adding projects may occur through formData as opposed to jsonBody. To specif

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation set already declares this is a non-read-only, open-world, non-idempotent write operation. The description adds only a credit-cost warning, but fails to explain what is updated, what the dictionary locators do, or what happens to affected project text.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short, but it is malformed and front-loads the incorrect verb "Create." The credit-cost phrase is useful, but the overall structure does not coherently communicate the tool's operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool that applies pronunciation dictionaries to a Studio project, the description is critically incomplete. Even with annotations and full schema coverage, an agent cannot confidently tell what the tool updates or how it differs from related pronunciation-dictionary tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents project_id, invalidate_affected_text, and pronunciation_dictionary_locators. The description adds no parameter-level meaning beyond that baseline, which is appropriate for full schema coverage but not helpful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says "Create Pronunciation Dictionaries," but the tool name and annotation title describe an update operation for pronunciation dictionaries. It gives a resource but uses the wrong verb, which is misleading rather than clarifying, and it does not distinguish this tool from siblings like patch_pronunciation_dictionary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites, and no exclusions. The only added context is that credits are spent, which is not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_secret_routeC

Update Convai Workspace Secret

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
typeYes
valueYes
secret_idYes

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true), so the agent knows this is a non-idempotent write with external effects. The description adds nothing beyond that: no mention of overwrite behavior, required permissions, or what happens to unmentioned fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At four words it is not verbose, but it is severely under-specified rather than concise; there is no front-loaded detail beyond the bare operation name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with four required parameters, no output schema, and 0% parameter documentation, the description is far too thin. Annotations only cover the safety profile; the description should at minimum clarify what updating a secret does and how the required fields are used.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four required parameters (secret_id, type, name, value), and the description lists none of them. The description does nothing to compensate for the fully undocumented schema, leaving the agent with no meaning for any parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb and resource ('Update Convai Workspace Secret'), which is clearer than a bare tautology. However, it essentially rephrases the tool name and does not distinguish it from close siblings such as create_secret_route, delete_secret_route, or update_environment_variable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives like create_secret_route or update_environment_variable, nor any prerequisites. The agent must infer all usage context from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_segment_languageC

Modify A Single Segment Spends ElevenLabs credits. Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
textNo
end_timeNo
languageYesID of the language.
dubbing_idYesID of the dubbing project.
segment_idYesID of the segment
start_timeNo

TDQS

C2.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, and openWorldHint=true, so the mutation profile is covered structurally. The description adds two facts not in the annotations: that the call spends ElevenLabs credits (cost implication) and that the tool is deprecated upstream (lifecycle status). Those are genuine behavioral additions beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two terse clauses, front-loaded with the action and then the deprecation caveat; nothing is padded. The phrasing is slightly garbled ("Modify A Single Segment Spends ElevenLabs credits"), which costs a point but not efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent, credit-consuming mutation with 6 parameters and no output schema, the description omits parameter meaning, required-field obligations, and any indication of what the response contains. The credit and deprecation notes are helpful but leave the definition substantially incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50% across 6 parameters, and the description explains none of them. It gives no hint about what text, language, dubbing_id, segment_id, start_time, or end_time do, so it fails to compensate for the undocumented half of the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says it modifies a single segment, which is a verb+resource, but it never states the actual resource being changed (the segment's language) that the tool name implies. An agent must open the schema to learn that language, text, and timing fields are involved. It is not a tautology, but the purpose is only vaguely conveyed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Deprecated upstream" is a useful caution flag but not usage guidance. Nothing says when to use this versus siblings like dubbing_target_transcript_segment_update, dubbing_transcript_segment_update, or migrate_segments, all of which operate on similar dubbing transcript segment data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_settings_routeD

Update Convai Settings

ParametersJSON Schema
NameRequiredDescriptionDefault
webhooksNo
can_use_mcp_serversNoWhether the workspace can use MCP servers
default_livekit_stackNo
rag_retention_period_daysNo
conversation_embedding_retention_daysNo
conversation_initiation_client_data_webhookNo

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so safety basics are covered externally. The description adds nothing beyond that - it doesn't explain that all six parameters are optional (partial/patch semantics), what the effect of the update is, or whether it overwrites or merges settings.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single four-word fragment. While it wastes no words, it is under-specified rather than concise - it lacks the substance an agent needs to act, so brevity here is a deficit, not an efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with six parameters, nested objects, no required fields, a 17% schema coverage, no output schema, and no clarifying annotations, this description is grossly incomplete. An agent cannot determine scope, semantics, or side effects from this definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% across six parameters, including nested webhook objects and an enum. The description provides zero parameter information, so it does not compensate for the sparsely documented schema and leaves parameters like default_livekit_stack and rag_retention_period_days unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update Convai Settings' essentially restates the tool name with the product name appended. It gives no indication of which settings are modified (webhooks, retention, MCP servers, livekit stack) and does not distinguish itself from siblings like get_settings_route, update_dashboard_settings_route, or patch_agent_settings_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites, and no reference to alternatives. The agent gets no signal about when this workspace-level settings update is appropriate versus the many other update/get settings tools in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_speakerC

Update Metadata For A Speaker Spends ElevenLabs credits. Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idNo
languagesNo
dubbing_idYesID of the dubbing project.
speaker_idYesID of the speaker.
voice_styleNo
speaker_nameNo
voice_stabilityNo
voice_similarityNo

TDQS

C2.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the mutation/safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the bar is lower. The description contributes genuinely non-obvious operational context beyond the annotations: the call spends ElevenLabs credits and the endpoint is deprecated upstream, both of which affect whether an agent should call it at all.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Only two sentences, so size is fine and the purpose is front-loaded. However, the first sentence is grammatically malformed ('Update Metadata For A Speaker Spends ElevenLabs credits'), running the action and the cost note together, which slightly obscures both.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an eight-parameter, non-idempotent mutation tool with no output schema and 25% schema coverage, the description omits the field semantics, null-handling, and credit-cost magnitude. It would need to carry more weight here; the deprecation and cost notes are valuable but insufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% – just dubbing_id and speaker_id are documented, while voice_id, languages, voice_style, speaker_name, voice_stability, and voice_similarity are undocumented. The description adds no parameter meaning whatsoever (no field list, no null/omission semantics, no valid ranges for the numeric voice settings), so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb (update) and resource (speaker), but 'Metadata' is vague about which of the eight fields are actually mutated, and it does not distinguish itself from siblings like create_speaker, edit_pvc_voice, or get_similar_voices_for_speaker. An agent knows roughly what the tool does but not its precise scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes no when-to-use guidance and no alternatives, even though many speaker and voice tools exist. 'Deprecated upstream' is a genuine usage-relevant signal, but it is stated as a status rather than as a directive (e.g., 'prefer X instead'), so the agent is left to infer what to do.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_speech_engineD

Update Speech Engine

ParametersJSON Schema
NameRequiredDescriptionDefault
asrNo
ttsNo
vadNo
nameNo
tagsNo
turnNo
privacyNo
languageNo
overridesNo
call_limitsNo
conversationNo
speech_engineNo
speech_engine_idYesThe speech engine ID (accepts seng_ or agent_ prefix)

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false and openWorldHint=true, so the safety profile is covered structurally; the description adds zero behavioral context beyond them. It does not say what gets overwritten, whether partial updates are supported, or any auth/rate constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words carry no waste but also no information; this is under-specification rather than true conciseness. For a 13-parameter mutation tool, the brevity is a deficiency, not a virtue.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 13 parameters, 8% schema coverage, no output schema, and only baseline annotations, the definition is completely inadequate for calling the tool correctly. An agent cannot determine what fields to send, what changes, or what preconditions apply.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 8% across 13 parameters, many of which are nested objects (asr, tts, vad, turn, privacy, call_limits, conversation, speech_engine). With such low coverage the description must compensate, and it supplies no parameter meaning whatsoever, not even for the required speech_engine_id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Update Speech Engine" merely restates the tool name and title with no additional specificity. It does state a verb and resource, but nothing distinguishes it from the other speech-engine siblings (create/delete/get/list_speech_engine) or clarifies what updating entails.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus create_speech_engine, get_speech_engine, or delete_speech_engine, nor any prerequisites such as needing an existing speech_engine_id. The description provides no usage context at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_tool_routeD

Update Tool

ParametersJSON Schema
NameRequiredDescriptionDefault
tool_idYesID of the requested tool.
tool_configYesConfiguration for the tool
response_mocksNo

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already disclose that this is a non-read-only, non-idempotent, open-world mutation. The description adds no behavioral context beyond that, such as what gets updated, side effects, permission requirements, or whether partial config replacement occurs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The two-word description is not usefully concise; it is under-specified. It neither front-loads meaningful constraints nor earns its place by conveying anything beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with a nested configuration object, required parameters, and no output schema, the description is completely inadequate. It leaves an agent unable to understand what the tool updates or how to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67% and the schema documents tool_id and the nested tool_config object, but the description contributes no parameter meaning at all. With moderate coverage and a complex nested config, "Update Tool" does not compensate for the remaining undocumented fields such as response_mocks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Update Tool" restates the tool name and annotation title without adding specificity. It gives only a generic verb and resource, so an agent cannot distinguish this from sibling update routes beyond the obvious update-versus-get/add/delete distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool, what prerequisites are required, or when to prefer alternatives such as add_tool_route, delete_tool_route, or get_tool_route. It is a bare purpose statement with no usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_whatsapp_accountD

Update Whatsapp Account

ParametersJSON Schema
NameRequiredDescriptionDefault
phone_number_idYes
enable_messagingNo
assigned_agent_idNo
enable_typing_indicatorNo
enable_audio_message_responseNo

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, openWorldHint=true, and destructiveHint=false, so the mutation/safety profile is covered. The description adds nothing beyond that — no note on whether omitted nullable fields are left unchanged, no auth/permission requirements, no side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words with zero waste, but this is under-specification rather than conciseness — there is no substantive content to be concise about.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter mutation tool with 0% schema coverage, no output schema, and no annotation detail beyond generic hints, the description is completely inadequate. An agent has no basis to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, and the description supplies no meaning for any of them. An agent cannot tell from the definition alone what enable_messaging, assigned_agent_id, enable_typing_indicator, or enable_audio_message_response control or how null values are interpreted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name verbatim ('Update Whatsapp Account') with no added specificity about what fields or behavior are affected. It does not distinguish this tool from siblings like update_phone_number_route, update_settings_route, or get_whatsapp_account.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the read/list counterparts (get_whatsapp_account, list_whatsapp_accounts) that an agent would need to consult first to obtain phone_number_id.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_workspace_memberD

Update Member

ParametersJSON Schema
NameRequiredDescriptionDefault
emailYesEmail of the target user.
is_lockedNo
workspace_roleNoSeat types for workspace members.
workspace_seat_typeNoSeat types for workspace members.

TDQS

D1.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnly=false, idempotent=false, destructive=false, and openWorld=true, covering the safety profile. The description adds no behavioral context beyond that: it says nothing about what fields are mutable, permission requirements, or side effects. It does not contradict annotations, but contributes nothing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

'Update Member' is two words with no structure or front-loading. This is under-specification rather than effective conciseness, as it omits all essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter mutation tool with ambiguous role/seat parameters and no output schema, the description is essentially absent. An agent cannot determine what the tool updates, what prerequisites exist, or what effects it has.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%: email is described, but is_locked has no description and workspace_role/workspace_seat_type share an identical ambiguous description ('Seat types for workspace members.'). The description itself adds no parameter meaning, leaving the role-vs-seat ambiguity unresolved.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update Member' merely restates the tool name and title. It does not specify which member attributes can be updated or how it differs from siblings like add_member, remove_member, or invite_user.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no routing to alternatives. With roughly 200 sibling tools including add_member and remove_member, an agent receives zero help selecting this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_assetD

Upload Asset

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesDisplay name for the asset.
asset_pathNoThe file to upload. Local path. Required for this call.
asset_base64NoBase64 contents for "asset". Use this when the server cannot read your local disk.
asset_filenameNoFilename to send for "asset". Some endpoints infer the audio format from it.

TDQS

D1.4/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds zero behavioral context beyond that — nothing about auth requirements, overwrite semantics for an existing asset name, or rate limits, which matter for a non-idempotent write.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words are technically short, but this is under-specification rather than conciseness — the phrase carries no information an agent can act on. There is no front-loaded statement of purpose or scope to build on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter mutation tool with no output schema, the description supplies none of the context needed: no destination, no overwrite behavior, no relationship to upload_file_route or upload_song. The schema alone cannot disambiguate this tool from its siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (name, asset_path, asset_base64, asset_filename) are already documented in the schema, including the mutually exclusive path-vs-base64 guidance. Per the baseline rule for high coverage with no param info in the description, a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Upload Asset" merely restates the tool name with no verb+resource specificity, no scope, and no differentiation from siblings such as list_assets, get_asset, or delete_asset_endpoint. An agent learns nothing about what is being uploaded, where it goes, or what constitutes an asset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of any alternative. Notably, the sibling set contains upload_file_route and upload_song, which an agent has no basis to choose between.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_file_routeD

Upload File

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathNoImage or PDF file to upload Local path. Required for this call.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.
conversation_idYes

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond that, e.g. no note about required auth, size limits, or interaction with cancel_file_upload_route.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words are technically concise and front-loaded, but the description is under-specified rather than efficient. It spends no budget on the information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter, multi-modality (image/PDF) upload tool with no output schema, the description is wholly inadequate. It omits what is uploaded, where it lands, or what happens on success or failure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, close to but below the high-coverage baseline. The description names no parameters and does not clarify the file_path vs file_base64 file_path/file_base64 trade-off or why conversation_id is required, so it fails to compensate for the remaining gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Upload File" merely restates the tool name and adds no distinguishing detail. It does not clarify which file domain this covers (assets, songs, knowledge base documents) versus siblings like upload_asset, upload_song, or cancel_file_upload_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or mention of alternatives is provided. The agent is left to infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_songC

Upload Music Spends ElevenLabs credits.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathNoThe audio file to upload. Local path. Required for this call.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.
with_timestampsNoWhether to transcribe the uploaded song and return word-level timestamps. If True, the response will include words_timestamps but will increase the latency.
with_waveform_visualNoWhether to return the visual waveform of the uploaded song.
extract_composition_planNoWhether to generate and return the composition plan for the uploaded song. Pass a model id (`music_v1`, `music_v2` or `music_v2_5`) to control which composition plan format is returned. Passing `true`/`false` is deprecated; `true` defaults to the `music_v1` plan format. Enabling this will increase t

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, and idempotentHint=false, so the safety profile is mostly covered. The description adds an important behavioral trait, 'Spends ElevenLabs credits,' but omits upload semantics such as supported formats, destination behavior, and latency/return characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loads the action before the credit warning. However, the grammar is rough, and the first fragment largely restates the tool name rather than adding structured detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter upload tool with no output schema, the description does not explain when to use it, how file_path and file_base64 relate, or what the upload produces. Annotations cover safety, but the description leaves significant operational context missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters, including file_path, file_base64, file_filename, and the optional timestamp/waveform/composition-plan flags. The description adds no additional meaning or constraints for these parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb and resource ('Upload Music') but does not distinguish this tool from sibling upload tools such as upload_asset or upload_file_route, and the phrasing is close to restating the tool name upload_song. The added credit warning is useful but does not clarify the core purpose beyond a vague upload action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no indication of when to call this tool versus alternatives like upload_asset, add_from_file, or create_audio_native_project. The only guidance-like statement is an implicit cost warning, with no prerequisites or route-selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

usage_by_product_over_timeD

Get Workspace Usage

ParametersJSON Schema
NameRequiredDescriptionDefault
filtersNo
end_timeYesEnd of the time range as a Unix timestamp in milliseconds. Must be at least 2020-01-01.
group_byNo
time_zoneNoIANA time zone identifier (e.g. 'America/New_York', 'Europe/London', 'UTC') used to align bucket boundaries for eligible `interval_seconds` values. Whole-day multiples start at local midnight; whole-hour multiples up to 24 hours align to local hour boundaries from midnight. Sub-hour intervals and ot
start_timeYesStart of the time range as a Unix timestamp in milliseconds. Must be at least 2020-01-01.
interval_secondsNoBucket size in seconds. Each row in the response covers this many seconds of the selected time range. For example, pass 3600 for hourly buckets or 86400 for daily buckets. Whether `time_zone` shifts bucket boundaries depends on this value: whole-day multiples (e.g. 86400) align to local midnight; wh

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description says 'Get', implying a safe read, while annotations declare readOnlyHint=false and idempotentHint=false, which the description neither explains nor reconciles. No behavioral traits (aggregation semantics, bucket boundaries, cost of large ranges) are disclosed, and the text actively conflicts with the annotation profile. Annotation Contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but this is under-specification rather than conciseness: a three-word sentence cannot carry the semantics of a six-parameter time-series query.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description says nothing about the response shape, bucket granularity, or how grouping affects rows. For an analytics tool with 17 group_by options, this leaves the agent unable to predict results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% and the description adds nothing: the meaning of filters, group_by, interval_seconds, and time_zone must be inferred entirely from the schema. The undocumented filters and group_by parameters receive no compensating explanation in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get Workspace Usage' only partially restates the tool name and never mentions the two defining dimensions: 'by product' and 'over time'. An agent cannot tell from the text that this is a time-bucketed, groupable analytics query rather than a simple total.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus the sibling usage_characters, nor any mention of prerequisites such as the time range or grouping needs. The agent is left to guess.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

usage_charactersC
Read-onlyIdempotent

Get Characters Usage Metrics (Deprecated) Deprecated upstream.

ParametersJSON Schema
NameRequiredDescriptionDefault
metricNoWhich metric to aggregate.
end_unixYesUTC Unix timestamp for the end of the usage window, in milliseconds. To include the last day of the window, the timestamp should be at 23:59:59 of that day.
start_unixYesUTC Unix timestamp for the start of the usage window, in milliseconds. To include the first day of the window, the timestamp should be at 00:00:00 of that day.
breakdown_typeNoHow to break down the information. Cannot be "user" if include_workspace_metrics is False.
aggregation_intervalNoHow to aggregate usage data over time. Can be "hour", "day", "week", "month", or "cumulative".
aggregation_bucket_sizeNoAggregation bucket size in seconds. Overrides the aggregation interval.
include_workspace_metricsNoWhether or not to include the statistics of the entire workspace.

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds the deprecation status, which is genuinely useful behavioral context, but it does not say whether the endpoint still returns data, errors out, or will be removed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is very short, which is appropriate for a deprecated stub, but 'Deprecated' appears twice ('(Deprecated)' and 'Deprecated upstream'), which is pure redundancy rather than front-loaded signal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter analytics tool with no output schema, the description should at minimum state what a caller gets back or what deprecation implies operationally. Instead it provides only a name and a status flag, leaving an agent unable to judge whether the call is worth making.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with 7 well-documented parameters and 3 enums, so the schema carries the full burden and baseline 3 applies. The description adds no meaning about metrics, intervals, or breakdown types beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (Characters Usage Metrics) and an implied read verb, so the basic purpose is inferable. However, it does not distinguish this tool from the sibling usage_by_product_over_time, and the parenthetical repeats the deprecation twice without clarifying what is being measured or for whom.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance offered is that the tool is deprecated, which weakly signals 'avoid', but no alternative or replacement is named even though usage_by_product_over_time exists as an obvious sibling. There is no statement of when this should still be called, if ever.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_pvc_voice_captchaC

Verify Pvc Voice Captcha

ParametersJSON Schema
NameRequiredDescriptionDefault
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
recording_pathNoAudio recording of the user Local path. Required for this call.
recording_base64NoBase64 contents for "recording". Use this when the server cannot read your local disk.
recording_filenameNoFilename to send for "recording". Some endpoints infer the audio format from it.

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, and idempotentHint=false, but the description adds nothing about side effects, what state changes on success/failure, or whether the verification can be retried. It fails to carry any of the behavioral burden left open by the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but this is under-specification rather than conciseness - a single four-word title restating the name provides no front-loaded information an agent can act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-read-only, open-world verification operation with four parameters and no output schema, the description omits everything an agent needs to call it correctly: expected input pairing, success/failure semantics, and prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% - the three optional recording variants and voice_id are fully documented in the schema itself. The description adds no parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a verbatim restatement of the tool name with no added explanation of what the verification actually does, what a 'captcha' check entails, or how it differs from siblings like get_pvc_voice_captcha or request_pvc_manual_verification.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to invoke this tool, what precondition (a pending captcha challenge?) must hold, or which sibling to use instead. The agent must guess from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

video_to_musicC

Video To Music Spends ElevenLabs credits. Returns application/zip bytes; pass output_path to save them.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNoOptional list of style tags (e.g. ['upbeat', 'cinematic']). A maximum of 10 tags is allowed.
model_idNo
descriptionNoOptional text description of the music you want. A maximum of 1000 characters is allowed.
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
videos_pathsNoOne or more video files sent via FormData array (multipart/form-data). They will be combined into one codec in order. A maximum of 10 videos is allowed, where the total size of the combined video is limited to 200MB. In total, the video can be up to 600 seconds long. Note that combining multiple vid
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
sign_with_c2paNoWhether to sign the generated song with C2PA. Applicable only for mp3 files.
videos_filenamesNoFilenames to send for "videos". Some endpoints infer the audio format from them.
videos_base64_listNoBase64 contents for "videos", one entry per file. Use this when the server cannot read your local disk.

TDQS

C2.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite sparse annotations (readOnlyHint=false, destructiveHint=false, openWorldHint=true), the description adds genuinely useful behavioral context: that it spends ElevenLabs credits and that it returns application/zip bytes with output_path controlling persistence. This is real value beyond the annotation set, though it omits failure modes and size/time limits already noted in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the credit-cost caveat front-loaded, which is the most decision-relevant fact. No filler, though the extreme terseness leaves the core operation unexplained.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly covers the return type (zip bytes) and how to persist it, which is important for a 9-parameter tool. However, it omits any statement of the actual transformation performed and any credit-amount or size guidance, leaving the agent without enough to confidently select the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 89%, so the schema already documents parameters like output_path, tags, and output_format. The description only echoes output_path, adding no syntax or format detail beyond what the structured schema provides; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description never states the actual operation – that it generates music from video input – it merely restates the tool name ('Video To Music') and then talks about credits and output bytes. No verb+resource explanation of what the tool produces distinguishes it from siblings like sound_generation or create_video_generation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of when to prefer this over sound_generation, create_video_generation, or separate_song_stems, and no prerequisites described. The only implicit constraint is the credit cost, which is not framed as a selection criterion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

whatsapp_outbound_callC

Make An Outbound Call Via Whatsapp

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes
whatsapp_user_idYes
whatsapp_phone_number_idYes
conversation_initiation_client_dataNo
whatsapp_call_permission_request_template_nameYes
whatsapp_call_permission_request_template_language_codeYes

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is covered structurally. The description adds nothing beyond that: it does not say that a permission-request template must be sent first, that calls may fail if the user has not consented, or anything about the asynchronous nature of the result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no wasted words and the action front-loaded, which is structurally fine for a title-style summary. But it is under-specified rather than concise in the helpful sense, and adds no sentence that earns its place beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter, five-required mutation with no output schema and no parameter documentation, the description is far too thin. It omits the permission-request flow, the meaning of agent_id, and any distinction between this and the other outbound-call handlers, so an agent cannot confidently construct a call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Six parameters, five of them required, with 0% schema description coverage, and the description explains none of them. Critical fields like whatsapp_call_permission_request_template_name, its language code, agent_id, and the nested conversation_initiation_client_data are completely undocumented, leaving the agent to guess at formats and IDs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb and resource ('Make An Outbound Call Via Whatsapp'), which is enough to know it initiates a WhatsApp voice call rather than sending a message (cf. sibling whatsapp_outbound_message). However, it gives no detail about the permission-request/template mechanism that the required parameters imply, so it is barely more informative than the name itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose this over sibling outbound-call tools such as handle_twilio_outbound_call, handle_exotel_outbound_call, or handle_sip_trunk_outbound_call, nor any prerequisites (e.g., WhatsApp account must exist, user must have granted call permission). Usage context is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

whatsapp_outbound_messageC

Send An Outbound Message Via Whatsapp

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes
template_nameYes
template_paramsYes
whatsapp_user_idYes
template_language_codeYes
whatsapp_phone_number_idYes
conversation_initiation_client_dataNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond that: it does not mention that WhatsApp outbound messages require pre-approved templates, that template_params must match the template, or any rate/auth constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero fluff. It is efficient, though its brevity comes at the cost of substance rather than being a model of tight specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, 6-required mutation tool with 0% schema coverage and no output schema, the description is far too thin. It should at minimum explain the template-based nature of WhatsApp messaging and what each required identifier refers to.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 parameters (6 required), and the description names none of them. An agent has no clue from the text that template_name, template_language_code, template_params, agent_id, or the phone/user IDs are needed, nor what format template_params takes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Send An Outbound Message Via Whatsapp'), so an agent knows this initiates a WhatsApp message. It does not, however, distinguish itself from siblings like whatsapp_outbound_call or handle_twilio_outbound_call, leaving channel selection to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as whatsapp_outbound_call for voice. An agent gets no help deciding between the outbound-message and outbound-call tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 390 tool updatesv0.1.0
    • First observedadd_chapter
    • First observedadd_documentation_to_knowledge_base
    • First observedadd_from_file
    • First observedadd_from_rules
    • First observedadd_language
    • First observedadd_mcp_server_tool_approval_route
    • First observedadd_mcp_tool_config_override_route
    • First observedadd_member
    • First observedadd_project
    • First observedadd_pvc_voice_samples
    • First observedadd_rules
    • First observedadd_sharing_voice
    • First observedadd_ticket_comment_route
    • First observedadd_tool_route
    • First observedadd_turn_comment_route
    • First observedadd_voice
    • First observedagent_testing_bulk_move_route
    • First observedassign_conversation_tags_route
    • First observedaudio_isolation
    • First observedaudio_isolation_stream
    • First observedaudio_native_project_update_content_endpoint
    • First observedaudio_native_update_content_from_url
    • First observedcancel_batch_call
    • First observedcancel_crawl_job_route
    • First observedcancel_file_upload_route
    • First observedcompile_procedures_route
    • First observedcompose_detailed
    • First observedcompose_detailed_stream
    • First observedcompose_plan
    • First observedconvert_chapter_endpoint
    • First observedconvert_project_endpoint
    • First observedcreate_agent_conversation_ticket_route
    • First observedcreate_agent_deployment_route
    • First observedcreate_agent_draft_route
    • First observedcreate_agent_response_test_route
    • First observedcreate_agent_route
    • First observedcreate_agent_test_folder_route
    • First observedcreate_audio_native_project
    • First observedcreate_auth_connection
    • First observedcreate_batch_call
    • First observedcreate_branch_route
    • First observedcreate_clip
    • First observedcreate_conversation_tag_route
    • First observedcreate_crawl_job_route
    • First observedcreate_dubbing
    • First observedcreate_environment_variable
    • First observedcreate_file_document_route
    • First observedcreate_finetune
    • First observedcreate_folder_route
    • First observedcreate_image_generation
    • First observedcreate_manual_agent_ticket_route
    • First observedcreate_mcp_server_route
    • First observedcreate_phone_number_route
    • First observedcreate_podcast
    • First observedcreate_procedure_route
    • First observedcreate_pvc_voice
    • First observedcreate_secret_route
    • First observedcreate_service_account
    • First observedcreate_service_account_api_key
    • First observedcreate_speaker
    • First observedcreate_speech_engine
    • First observedcreate_text_document_route
    • First observedcreate_text_to_speech_generation
    • First observedcreate_url_document_route
    • First observedcreate_video_generation
    • First observedcreate_voice
    • First observedcreate_workspace_webhook_route
    • First observeddelete_agent_conversation_ticket_route
    • First observeddelete_agent_draft_route
    • First observeddelete_agent_hold_audio_route
    • First observeddelete_agent_route
    • First observeddelete_agent_test_folder_route
    • First observeddelete_asset_endpoint
    • First observeddelete_audio_isolation_history_item
    • First observeddelete_auth_connection
    • First observeddelete_batch_call
    • First observeddelete_chapter_endpoint
    • First observeddelete_chat_response_test_route
    • First observeddelete_conversation_route
    • First observeddelete_conversation_tag_route
    • First observeddelete_dubbing
    • First observeddelete_finetune
    • First observeddelete_invite
    • First observeddelete_knowledge_base_document
    • First observeddelete_mcp_server_route
    • First observeddelete_phone_number_route
    • First observeddelete_procedure_draft_route
    • First observeddelete_project
    • First observeddelete_pvc_voice_sample
    • First observeddelete_rag_index
    • First observeddelete_sample
    • First observeddelete_secret_route
    • First observeddelete_segment
    • First observeddelete_service_account_api_key
    • First observeddelete_speech_engine
    • First observeddelete_speech_history_item
    • First observeddelete_tool_route
    • First observeddelete_transcript_by_id
    • First observeddelete_voice
    • First observeddelete_whatsapp_account
    • First observeddelete_workspace_webhook_route
    • First observeddisable
    • First observeddownload_speech_history_items
    • First observeddub
    • First observeddubbing_language_create
    • First observeddubbing_language_delete
    • First observeddubbing_language_get
    • First observeddubbing_language_list
    • First observeddubbing_project_create
    • First observeddubbing_project_delete
    • First observeddubbing_project_get
    • First observeddubbing_project_list
    • First observeddubbing_target_transcript_get
    • First observeddubbing_target_transcript_regenerate
    • First observeddubbing_target_transcript_segment_update
    • First observeddubbing_target_transcript_segments_update
    • First observeddubbing_transcript_get
    • First observeddubbing_transcript_segment_add
    • First observeddubbing_transcript_segment_delete
    • First observeddubbing_transcript_segment_update
    • First observeddubbing_transcript_segments_update
    • First observedduplicate_agent_route
    • First observededit_chapter
    • First observededit_project
    • First observededit_project_content
    • First observededit_pvc_voice
    • First observededit_pvc_voice_sample
    • First observededit_service_account_api_key
    • First observededit_voice
    • First observededit_voice_settings
    • First observededit_workspace_webhook_route
    • First observedexport_batch_call
    • First observedforced_alignment
    • First observedgenerate
    • First observedget_agent_conversation_ticket_route
    • First observedget_agent_knowledge_base_size
    • First observedget_agent_knowledge_base_summaries_route
    • First observedget_agent_link_route
    • First observedget_agent_llm_expected_cost_calculation
    • First observedget_agent_response_test_route
    • First observedget_agent_response_tests_summaries_route
    • First observedget_agent_route
    • First observedget_agent_summaries_route
    • First observedget_agent_test_folder_route
    • First observedget_agent_topics_route
    • First observedget_agent_widget_route
    • First observedget_agents_route
    • First observedget_asset
    • First observedget_assignable_users_route
    • First observedget_audio_from_sample
    • First observedget_audio_full_from_speech_history_item
    • First observedget_audio_isolation_history
    • First observedget_audio_native_project_settings_endpoint
    • First observedget_batch_call
    • First observedget_branch_route
    • First observedget_branches_route
    • First observedget_chapter_by_id_endpoint
    • First observedget_chapter_snapshot_endpoint
    • First observedget_chapter_snapshots
    • First observedget_chapters
    • First observedget_conversation_audio_route
    • First observedget_conversation_histories_route
    • First observedget_conversation_history_route
    • First observedget_conversation_signed_link
    • First observedget_conversation_sip_messages
    • First observedget_conversation_summary_route
    • First observedget_conversation_tag_route
    • First observedget_conversation_users_route
    • First observedget_crawl_job_route
    • First observedget_dashboard_settings_route
    • First observedget_documentation_chunk_from_knowledge_base
    • First observedget_documentation_chunks_from_knowledge_base
    • First observedget_documentation_from_knowledge_base
    • First observedget_dubbed_file
    • First observedget_dubbed_metadata
    • First observedget_dubbed_transcript_file
    • First observedget_dubbing_resource
    • First observedget_dubbing_transcripts
    • First observedget_environment_variable
    • First observedget_finetune
    • First observedget_finetunes
    • First observedget_groups_endpoint
    • First observedget_image_generation
    • First observedget_knowledge_base_bulk_dependent_agents_route
    • First observedget_knowledge_base_content
    • First observedget_knowledge_base_dependent_agents
    • First observedget_knowledge_base_list_route
    • First observedget_knowledge_base_source_file_url
    • First observedget_library_voices
    • First observedget_live_count
    • First observedget_livekit_token
    • First observedget_mcp_route
    • First observedget_mcp_tool_config_override_route
    • First observedget_models
    • First observedget_or_create_rag_indexes
    • First observedget_phone_number_route
    • First observedget_procedure_draft_route
    • First observedget_procedure_route
    • First observedget_project_by_id
    • First observedget_project_muted_tracks_endpoint
    • First observedget_project_snapshot_endpoint
    • First observedget_project_snapshots
    • First observedget_projects
    • First observedget_pronunciation_dictionaries_metadata
    • First observedget_pronunciation_dictionary_metadata
    • First observedget_pronunciation_dictionary_version_pls
    • First observedget_public_llm_expected_cost_calculation
    • First observedget_pvc_sample_audio
    • First observedget_pvc_sample_speakers
    • First observedget_pvc_sample_visual_waveform
    • First observedget_pvc_voice_captcha
    • First observedget_rag_index_overview
    • First observedget_rag_indexes
    • First observedget_resource_metadata
    • First observedget_secret_dependencies_route
    • First observedget_secret_route
    • First observedget_secrets_route
    • First observedget_service_account_api_keys_route
    • First observedget_settings_route
    • First observedget_signed_url_deprecated
    • First observedget_similar_library_voices
    • First observedget_similar_voices_for_speaker
    • First observedget_single_use_token
    • First observedget_speaker_audio
    • First observedget_speech_engine
    • First observedget_speech_history
    • First observedget_speech_history_item_by_id
    • First observedget_test_invocation_route
    • First observedget_text_to_speech_generation
    • First observedget_tool_dependent_agents_route
    • First observedget_tool_executions_route
    • First observedget_tool_route
    • First observedget_tools_route
    • First observedget_transcript_by_id
    • First observedget_user_info
    • First observedget_user_subscription_info
    • First observedget_user_voices_v2
    • First observedget_version_metadata_route
    • First observedget_video_generation
    • First observedget_voice_accents
    • First observedget_voice_by_id
    • First observedget_voice_settings
    • First observedget_voice_settings_default
    • First observedget_voices
    • First observedget_whatsapp_account
    • First observedget_workspace_audit_logs
    • First observedget_workspace_batch_calls
    • First observedget_workspace_members
    • First observedget_workspace_service_accounts
    • First observedget_workspace_webhooks_route
    • First observedhandle_exotel_outbound_call
    • First observedhandle_sip_trunk_outbound_call
    • First observedhandle_twilio_outbound_call
    • First observedinvite_user
    • First observedinvite_users_bulk
    • First observedlist_agent_conversation_tickets_route
    • First observedlist_assets
    • First observedlist_auth_connections
    • First observedlist_available_llms
    • First observedlist_chat_response_tests_route
    • First observedlist_conversation_tags_route
    • First observedlist_crawl_jobs_route
    • First observedlist_dubs
    • First observedlist_environment_variables
    • First observedlist_image_generations
    • First observedlist_mcp_server_tools_route
    • First observedlist_mcp_servers_route
    • First observedlist_phone_numbers_route
    • First observedlist_procedures_route
    • First observedlist_sip_messages
    • First observedlist_speech_engines
    • First observedlist_test_invocations_route
    • First observedlist_text_to_speech_generations
    • First observedlist_video_generations
    • First observedlist_whatsapp_accounts
    • First observedlist_workspace_conversation_tickets_route
    • First observedmerge_branch_into_target
    • First observedmerge_preview_route
    • First observedmigrate_segments
    • First observedpatch_agent_settings_route
    • First observedpatch_pronunciation_dictionary
    • First observedpost_agent_avatar_route
    • First observedpost_agent_hold_audio_route
    • First observedpost_conversation_feedback_route
    • First observedpost_knowledge_base_bulk_delete_route
    • First observedpost_knowledge_base_bulk_move_route
    • First observedpost_knowledge_base_move_route
    • First observedpublic_create_order
    • First observedpublic_get_available_languages
    • First observedpublic_get_media_info
    • First observedpublic_get_order
    • First observedpublic_get_order_deliverables
    • First observedpublic_list_orders
    • First observedpublic_register_media
    • First observedpublic_remove_order_item
    • First observedpublic_submit_order
    • First observedpublic_update_order
    • First observedpublic_upsert_order_item
    • First observedquery_agent_knowledge_base_rag_route
    • First observedrag_index_status
    • First observedrebase_branch_onto_main
    • First observedrebase_preview_route
    • First observedredirect_to_mintlify
    • First observedrefresh_url_document_route
    • First observedregister_twilio_call
    • First observedremove_mcp_server_tool_approval_route
    • First observedremove_mcp_tool_config_override_route
    • First observedremove_member
    • First observedremove_procedure_route
    • First observedremove_rules
    • First observedrender
    • First observedreplicate_voice_to_isolated_environment
    • First observedrequest_pvc_manual_verification
    • First observedrequests_list
    • First observedresolve_conversation_reference_route
    • First observedresubmit_tests_route
    • First observedretry_batch_call
    • First observedrun_agent_test_suite_route
    • First observedrun_conversation_analysis
    • First observedrun_conversation_evaluations
    • First observedrun_conversation_simulation_route
    • First observedrun_conversation_simulation_route_stream
    • First observedrun_pvc_voice_training
    • First observedsearch_groups
    • First observedsearch_knowledge_base_content_route
    • First observedseparate_song_stems
    • First observedset_rules
    • First observedset_third_party_disabling_policy
    • First observedshare_resource_endpoint
    • First observedsmart_search_conversation_messages_route
    • First observedsound_generation
    • First observedspeech_to_speech_full
    • First observedspeech_to_speech_stream
    • First observedspeech_to_text
    • First observedstart_speaker_separation
    • First observedstream_chapter_snapshot_audio
    • First observedstream_compose
    • First observedstream_project_snapshot_archive_endpoint
    • First observedstream_project_snapshot_audio_endpoint
    • First observedtext_search_conversation_messages_route
    • First observedtext_to_dialogue
    • First observedtext_to_dialogue_full_with_timestamps
    • First observedtext_to_dialogue_stream
    • First observedtext_to_dialogue_stream_with_timestamps
    • First observedtext_to_speech_full
    • First observedtext_to_speech_full_with_timestamps
    • First observedtext_to_speech_stream
    • First observedtext_to_speech_stream_with_timestamps
    • First observedtext_to_voice
    • First observedtext_to_voice_design
    • First observedtext_to_voice_preview_stream
    • First observedtext_to_voice_remix
    • First observedtranscribe
    • First observedtranslate
    • First observedunassign_conversation_tag_route
    • First observedunshare_resource_endpoint
    • First observedupdate_agent_conversation_ticket_route
    • First observedupdate_agent_response_test_route
    • First observedupdate_agent_test_folder_route
    • First observedupdate_auth_connection
    • First observedupdate_branch_route
    • First observedupdate_conversation_tag_route
    • First observedupdate_dashboard_settings_route
    • First observedupdate_document_route
    • First observedupdate_environment_variable
    • First observedupdate_file_document_route
    • First observedupdate_finetune
    • First observedupdate_mcp_server_approval_policy_route
    • First observedupdate_mcp_server_config_route
    • First observedupdate_mcp_tool_config_override_route
    • First observedupdate_phone_number_route
    • First observedupdate_procedure_draft_route
    • First observedupdate_pronunciation_dictionaries
    • First observedupdate_secret_route
    • First observedupdate_segment_language
    • First observedupdate_settings_route
    • First observedupdate_speaker
    • First observedupdate_speech_engine
    • First observedupdate_tool_route
    • First observedupdate_whatsapp_account
    • First observedupdate_workspace_member
    • First observedupload_asset
    • First observedupload_file_route
    • First observedupload_song
    • First observedusage_by_product_over_time
    • First observedusage_characters
    • First observedverify_pvc_voice_captcha
    • First observedvideo_to_music
    • First observedwhatsapp_outbound_call
    • First observedwhatsapp_outbound_message

TDQS

C2/5.0

Scored across 390 tools

Disambiguation2/5

Many tools overlap heavily: text_to_speech_full/stream/with_timestamps variants, multiple delete/get tools for the same resource, and deprecated duplicates. Descriptions are present but the boundaries between similar tools are unclear, causing frequent misselection risk.

Naming Consistency2/5

Naming uses mixed patterns: verb_noun (get_agent_route), noun_verb (dubbing_language_create), inconsistent suffixes (_route, _endpoint), random prefixes (public_, deprecated markers), and several one-word tools (disable, dub, render). No predictable convention.

Tool Count1/5

390 tools is an extreme mismatch for a single API surface. The massive duplication and inclusion of deprecated endpoints indicate poor curation rather than a well-scoped set.

Completeness4/5

The tool set covers a wide range of ElevenLabs capabilities: TTS, STS, dubbing, agents, music, voices, workspace, billing, and more. Minor gaps exist (e.g., some deprecated tools lack modern replacements), but core workflows are largely present.

Related MCP Connectors

Related MCP Servers