download_caption
Download caption text from a YouTube video in SRT or VTT format by specifying the caption ID.
Instructions
Download a caption track's text. fmt: srt|vtt.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| fmt | No | srt | |
| caption_id | Yes |
Download caption text from a YouTube video in SRT or VTT format by specifying the caption ID.
Download a caption track's text. fmt: srt|vtt.
| Name | Required | Description | Default |
|---|---|---|---|
| fmt | No | srt | |
| caption_id | Yes |
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. 'Download' indicates a non-destructive read operation, but no details on output format (SRT/VTT file contents), error conditions, or rate limits are given. It is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise—one sentence plus format hint—with no wasted words. It delivers key information upfront.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple download tool with two parameters and no output schema, the description covers the core action and format. However, it omits context like the typical workflow (e.g., using list_captions first) and the exact nature of the output (e.g., text content in SRT/VTT format). Functional but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so description must compensate. It adds value by specifying format options ('fmt: srt|vtt'), which the schema lacks. However, it does not explain the required 'caption_id' parameter (e.g., where to obtain it). Partial compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool downloads a caption track's text, using a specific verb and resource. It distinguishes from siblings like list_captions (which lists metadata) and upload_caption (which uploads). The format hint further clarifies the output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when caption text is needed, but lacks explicit when/when-not guidance or alternatives. It does not mention prerequisites like obtaining a caption_id from list_captions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/aaronckj/yt-studio-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server