Capslane
Server Details
Retrieve YouTube transcripts with timestamps, native captions and asynchronous generation when captions are unavailable. Requires a Capslane API key.
- Status
- Healthy
- Uptime
- 98.7% over 21 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- Webba-Creative-Technologies/capslane-mcp
- GitHub Stars
- 0
TDQS
Scored across 3 tools
The three tools have clearly distinct purposes: retrieving a transcript, checking job status, and listing available languages. There is no overlap in functionality, and the descriptions explicitly differentiate behavior such as unit consumption and generation.
Tool names follow a predictable 'get_' and 'list_' prefix pattern, making the intent of each visible. Minor inconsistency exists between 'get_youtube_transcript' (resource-specific) and 'get_transcript_status' (status check), but the overall pattern remains coherent.
With only three tools, the server is tightly scoped to its core transcript workflow: fetch, check status, and list languages. Each tool is necessary and the count is well within the expected range for a focused service.
The tool surface covers the primary transcript lifecycle but lacks an explicit cancel or delete operation. However, the status tool and the mention of a generation job may make cancellation unnecessary; the current set handles retrieval and monitoring adequately.
Available Tools
3 toolsget_transcript_statusGet transcript job statusARead-onlyIdempotentInspect
Check the same accepted job without consuming another transcript unit. Content means success; stop on failed, cancelled or completed without content (stored result unavailable or expired). A successful tool call can still describe a pending or failed job. Poll with a delay and deadline, never by resubmitting the video.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | Job identifier returned by get_youtube_transcript |
Output Schema
| Name | Required | Description |
|---|---|---|
| lang | No | |
| error | No | |
| jobId | No | |
| cached | No | |
| source | No | |
| status | No | Job state, never an HTTP status. Completed without content means the stored result is unavailable or expired. |
| content | No | Transcript itself. Check this before jobId; it can coexist with a completed job ID. |
| progress | No | |
| requestId | No | |
| availableLangs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, the description discloses key behavioral traits: a successful API call can still represent a pending or failed job, content indicates success, and completed-without-content means the stored result is unavailable or expired. This is meaningful operational context the annotations do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, each earning its place: the core purpose and resource cost, success/failure semantics, the pending-job caveat, and concrete polling guidance. The most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter status-checking tool with readOnly and idempotent annotations plus an output schema, the description covers the essential operational edge cases and polling strategy. Nothing critical is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the jobId parameter is fully documented as the identifier returned by get_youtube_transcript, with a pattern. The description adds the notion of checking the 'same accepted job' but does not need to add more because the schema already defines the parameter precisely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks the status of an already-accepted transcription job and explicitly contrasts it with resubmitting the video. This makes it easy to distinguish from get_youtube_transcript and list_available_languages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: poll an existing job in place of resubmitting the video. It also defines stop conditions ('failed, cancelled or completed without content') and instructs to poll with a delay and deadline, leaving no ambiguity about the intended workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_youtube_transcriptGet YouTube transcriptAInspect
Retrieve captions or generate a transcript for a public YouTube video. Each submission consumes a transcript unit, including cache hits. Can create a generation job. For summaries, notes or timestamp citations, retrieve the transcript first; this tool does not summarize.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Public YouTube URL or 11-character video ID | |
| lang | No | Preferred ISO language code | |
| mode | No | Cache first in every mode. On a miss: native never generates; auto generates only after missing captions; generate requests audio transcription | auto |
| text | No | Plain text for an immediate response. Completed jobs return segments with offsets and durations in milliseconds | |
| waitForCompletion | No | Wait up to twenty minutes for generation. Set false for interactive clients, then poll the returned jobId with get_transcript_status |
Output Schema
| Name | Required | Description |
|---|---|---|
| lang | No | |
| error | No | |
| jobId | No | |
| cached | No | |
| source | No | |
| status | No | Job state, never an HTTP status. Completed without content means the stored result is unavailable or expired. |
| content | No | Transcript itself. Check this before jobId; it can coexist with a completed job ID. |
| progress | No | |
| requestId | No | |
| availableLangs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses non-obvious behaviors beyond the annotations: 'Each submission consumes a transcript unit, including cache hits' and 'Can create a generation job.' These additions clarify side effects (cost, background jobs) that the annotations (readOnlyHint=false, idempotentHint=false) only imply. This is valuable transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no fluff. It front-loads the core purpose, then discloses the side-effect cost, then gives usage guidance. Every sentence adds value and the structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential context: purpose, public video requirement, unit consumption, job creation, and non-summarization. Combined with the detailed schema and annotations, an agent has everything needed to decide when and how to invoke it. The output schema exists to describe return structure, so no additional explanation is required.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides 100% parameter coverage with detailed descriptions (e.g., mode explains caching and generation behavior, waitForCompletion explains polling). The description itself adds minimal parameter-specific information beyond the schema, so it does not significantly improve on the schema's already rich semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Retrieve captions or generate a transcript for a public YouTube video.' It uses a specific verb ('retrieve'/'generate') and resource ('transcript'), and differentiates itself from sibling tools by noting 'this tool does not summarize' and implying that it returns raw transcript for further processing. This distinguishes it from list_available_languages and get_transcript_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context on when to use the tool ('For summaries, notes or timestamp citations, retrieve the transcript first') and provides an exclusion ('this tool does not summarize'). However, it does not explicitly name alternative sibling tools or state when to prefer them, so the guidance is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_available_languagesList transcript languagesAInspect
Return languages observed by a native transcript request. Consumes one transcript unit and can populate the cache; it is not a free metadata lookup. Never starts generation. A cached result can have a generated source.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Public YouTube URL or 11-character video ID |
Output Schema
| Name | Required | Description |
|---|---|---|
| jobId | No | |
| status | No | |
| requestId | No | |
| selectedLang | No | |
| availableLangs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral context beyond the annotations: it consumes one transcript unit, can populate the cache, is not free, never starts generation, and cached results can have a generated source. This meaningfully enriches the raw annotation flags.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each adding important operational information with no redundancy. The primary action is front-loaded, followed by cost and side-effect caveats.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with a full output schema, the description covers the essential behavioral nuances: cost, cache effects, generation behavior, and the possibility of generated cached sources. Nothing critical appears missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents url as a public YouTube URL or 11-character video ID. The description adds no extra parameter-level detail, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Return languages observed by a native transcript request.' It also distinguishes the tool from a generic metadata lookup by stating it consumes a transcript unit, which helps differentiate it from the sibling transcript tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear contextual guidance: this is not a free metadata lookup, it consumes a transcript unit, and it never starts generation. However, it does not explicitly name alternatives or state when to prefer get_transcript_status or get_youtube_transcript instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- Changed
get_transcript_status1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "additionalProperties": true, + "properties": { + "availableLangs": { + "items": { + "type": "string" + }, + "type": "array" + }, + "cached": { + "type": "boolean" + }, + "content": { + "anyOf": [ + { + "type": "string" + }, + { + "items": { + "additionalProperties": false, + "properties": { + "duration": { + "description": "Duration in milliseconds", + "minimum": 0, + "type": "number" + }, + "lang": { + "type": "string" + }, + "offset": { + "description": "Start offset in milliseconds", + "minimum": 0, + "type": "number" + }, + "text": { + "type": "string" + } + }, + "required": [ + "text", + "offset", + "duration", + "lang" + ], + "type": "object" + }, + "type": "array" + } + ], + "description": "Transcript itself. Check this before jobId; it can coexist with a completed job ID." + }, + "error": { + "type": "string" + }, + "jobId": { + "type": "string" + }, + "lang": { + "type": "string" + }, + "progress": { + "type": "number" + }, + "requestId": { + "type": "string" + }, + "source": { + "enum": [ + "native", + "generated" + ], + "type": "string" + }, + "status": { + "description": "Job state, never an HTTP status. Completed without content means the stored result is unavailable or expired.", + "enum": [ + "queued", + "downloading", + "processing", + "persisting", + "completed", + "failed", + "cancelled" + ], + "type": "string" + } + }, + "type": "object" +}
- Changed
get_youtube_transcript1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "additionalProperties": true, + "properties": { + "availableLangs": { + "items": { + "type": "string" + }, + "type": "array" + }, + "cached": { + "type": "boolean" + }, + "content": { + "anyOf": [ + { + "type": "string" + }, + { + "items": { + "additionalProperties": false, + "properties": { + "duration": { + "description": "Duration in milliseconds", + "minimum": 0, + "type": "number" + }, + "lang": { + "type": "string" + }, + "offset": { + "description": "Start offset in milliseconds", + "minimum": 0, + "type": "number" + }, + "text": { + "type": "string" + } + }, + "required": [ + "text", + "offset", + "duration", + "lang" + ], + "type": "object" + }, + "type": "array" + } + ], + "description": "Transcript itself. Check this before jobId; it can coexist with a completed job ID." + }, + "error": { + "type": "string" + }, + "jobId": { + "type": "string" + }, + "lang": { + "type": "string" + }, + "progress": { + "type": "number" + }, + "requestId": { + "type": "string" + }, + "source": { + "enum": [ + "native", + "generated" + ], + "type": "string" + }, + "status": { + "description": "Job state, never an HTTP status. Completed without content means the stored result is unavailable or expired.", + "enum": [ + "queued", + "downloading", + "processing", + "persisting", + "completed", + "failed", + "cancelled" + ], + "type": "string" + } + }, + "type": "object" +}
- Changed
list_available_languages1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "additionalProperties": true, + "properties": { + "availableLangs": { + "items": { + "type": "string" + }, + "type": "array" + }, + "jobId": { + "type": "string" + }, + "requestId": { + "type": "string" + }, + "selectedLang": { + "type": "string" + }, + "status": { + "type": "string" + } + }, + "type": "object" +}
3 tool updates
- First observed
get_transcript_status - First observed
get_youtube_transcript - First observed
list_available_languages
Related MCP Connectors
Transcripts of YouTube videos, playlists and channels with timestamps; SRT or WebVTT too.
Fetch the full transcript of any YouTube video as clean text. No API key, no signup.
Free YouTube transcripts, no API key: videos, channel lists, latest uploads, bulk download links.
Transcripts of YouTube videos, playlists and channels in any language: text, SRT, VTT or JSON.
Related MCP Servers
- AlicenseAqualityDmaintenanceRetrieves transcripts from YouTube videos with support for multiple languages, timestamp control, and language detection. Enables video content analysis, summarization, and quote extraction without manually downloading or watching videos.272 npm15MIT
- AlicenseAqualityDmaintenanceEnables AI assistants to fetch YouTube video transcripts with precise timestamps, multi-language support, and time-range filtering.31MIT
- FlicenseAqualityCmaintenanceEnables fetching YouTube video transcripts with metadata, including timed captions in multiple formats (JSON, SRT, VTT, CSV, TXT) and preprocessing options.41-
- FlicenseNot gradedqualityCmaintenanceEnables retrieving plain-text transcripts from YouTube videos in any available language, including auto-generated captions, by providing a URL or video ID.-
Glama MCP Gateway
Add one secure layer between your agents and this server.