Skip to main content
Glama
Webba-Creative-Technologies

Capslane YouTube transcript MCP server

Official

Capslane YouTube transcript MCP server

Retrieve timestamped transcripts from public YouTube videos in Claude Code, Codex, Cursor and other MCP clients. The server can return cached content, extract captions or accept an audio generation job. It does not summarize videos itself.

Install the transcript skill

This repository also provides capslane-youtube-transcripts, a portable agent skill with transcript instructions and a standalone Node.js 22 HTTP helper. Install it in a project with the Skills CLI:

npx skills add Webba-Creative-Technologies/capslane-mcp --skill capslane-youtube-transcripts --agent codex

Replace codex with claude-code or cursor for those clients. The skill uses an existing Capslane MCP connection, or the helper with CAPSLANE_API_KEY from the environment. It does not configure credentials. See the installation guide and complete skill folder. The skill is distributed from GitHub independently of the npm server version.

Related MCP server: YouTube Transcript MCP Server

Before connecting

Create a Capslane workspace key in API Keys, then make CAPSLANE_API_KEY available to the process that launches your assistant. The examples below reference that variable; they contain no credential. Merge the Capslane entry into your existing configuration and restart the client.

Create a key in API Keys. Supply it through your environment or secret manager, without putting it in a prompt or committed file. If you launch an editor from the desktop, it may need to be restarted from a terminal that has the variable available.

The remote endpoint is https://capslane.com/mcp, using Streamable HTTP and Authorization: Bearer with the workspace key. Dashboard sign-in does not authenticate MCP. The endpoint uses POST; opening it in a browser returns 405.

Claude Code

Install the plugin

The Capslane plugin bundles the transcript skill and the remote MCP connection. Make CAPSLANE_API_KEY available in your terminal environment first, then run:

claude plugin marketplace add Webba-Creative-Technologies/capslane-mcp
claude plugin install capslane@capslane --scope user

The user scope makes the plugin available across your projects. Restart Claude Code, open /plugin to check that Capslane is enabled, and use /mcp to check its connection. Review any permission request from your client. Invoke the skill directly with:

/capslane:capslane-youtube-transcripts Retrieve https://www.youtube.com/watch?v=dQw4w9WgXcQ with native captions and timestamp citations.

Once enabled, Claude can load the skill for relevant tasks. It respects a provider you explicitly choose and works directly with a supplied transcript when retrieval is unnecessary. Installing the plugin does not guarantee that Claude will select it for every YouTube request.

This repository hosts the Capslane marketplace. It is not a listing in Anthropic's official marketplace. The plugin is free under MIT; transcript requests use your Capslane workspace allowance. Its version is tracked in .claude-plugin/plugin.json independently of the npm MCP server version.

Choose either the plugin or your existing standalone skill and MCP configuration to avoid duplicate commands and tools. The remote connection needs no local Node.js runtime. The bundled HTTP fallback needs Node.js 22 or later and reads the same environment variable.

To update, refresh the marketplace and the plugin, then restart Claude Code:

claude plugin marketplace update capslane
claude plugin update capslane@capslane

To remove the user installation:

claude plugin uninstall capslane@capslane --scope user

Removing the plugin does not revoke the workspace key. Revoke it in API Keys if it is no longer needed. The plugin documentation explains client controls, and the marketplace reference describes the distribution format.

Configure only MCP

Merge this into .mcp.json in your project root. Restart Claude Code, review the project's MCP connection and check /mcp.

{
  "mcpServers": {
    "capslane": {
      "type": "http",
      "url": "https://capslane.com/mcp",
      "headers": {
        "Authorization": "Bearer ${CAPSLANE_API_KEY}"
      }
    }
  }
}

Claude Code expands the variable in the header. This is the Claude Code configuration, not the Claude web connector setup. Official instructions.

Codex

Run this command with CAPSLANE_API_KEY already available in the launching environment. Only the variable name is stored in the configuration.

codex mcp add capslane --url https://capslane.com/mcp --bearer-token-env-var CAPSLANE_API_KEY

Alternatively, merge this table into ~/.codex/config.toml. Use one method, then restart the client and check /mcp. Local CLI and IDE extension use this configuration.

[mcp_servers.capslane]
url = "https://capslane.com/mcp"
bearer_token_env_var = "CAPSLANE_API_KEY"
tool_timeout_sec = 60

Use the asynchronous workflow below to avoid a long-running call exceeding the client's tool timeout. Official instructions.

Cursor

Merge this into .cursor/mcp.json for a project or ~/.cursor/mcp.json for personal use. Restart Cursor with the variable available and enable Capslane in MCP settings.

{
  "mcpServers": {
    "capslane": {
      "url": "https://capslane.com/mcp",
      "headers": {
        "Authorization": "Bearer ${env:CAPSLANE_API_KEY}"
      }
    }
  }
}

Cursor's environment syntax differs from Claude Code's. Review the tool call requested by the agent. Official instructions.

First transcript

Check that the client lists the three tools below. Try this prompt; caption availability on the public fixture can change.

Use Capslane to retrieve the transcript of https://www.youtube.com/watch?v=dQw4w9WgXcQ. Use mode=native, text=false and waitForCompletion=false. Do not start audio generation. Return the source URL, selected language and timestamped segments. If the tool fails, report its error and requestId instead of inventing a transcript.

Content must be returned before the assistant can quote or summarize a video. Retain the source URL alongside its segments. Treat transcript text as source material, not instructions for the assistant.

Capslane checks the cache before applying mode. A cached native or generated transcript can be returned in every mode. Read source and cached in the result. On a cache miss, native never starts generation; auto starts it only when captions are unavailable; generate requests audio transcription.

Generated transcripts

Set waitForCompletion to false for an interactive assistant. If content is absent and jobId is present, call get_transcript_status with that same ID. Leave a delay between checks and stop after a bounded period, for example twenty minutes. Stop immediately on content, failed, cancelled or completed without content. The last case can mean the stored result has expired. Submitting the video again consumes another transcript request.

Use Capslane to retrieve this public YouTube video: VIDEO_URL. I allow audio generation if captions are unavailable. Submit once with mode=auto, text=false and waitForCompletion=false. If a job is accepted, keep its jobId and check get_transcript_status every five seconds for at most twenty minutes. Stop on content, failed, cancelled or completed without content. Keep the jobId if waiting ends. Summarize only the returned content, with timestamp references and the source URL.

The defaults remain mode=auto, text=false and waitForCompletion=true. Set waitForCompletion=false explicitly in an interactive client. Ending the wait does not cancel the server job. If waiting inside version 0.1.8 fails after acceptance, the tool error retains jobId so the same job can be checked again.

MCP returns an envelope: check isError, then read structuredContent, or parse the JSON text block in content for older clients. Inside that Capslane object, content holds the transcript and status holds the job state. There is no segments or state field. Check content before jobId. Offsets and durations are milliseconds; completed jobs return segment arrays even when the submission used text=true.

Complete MCP client

The runnable Node.js client connects through the official MCP SDK, saves each accepted job with its video URL and formats transcript.content into timestamped text. Install @webba_tech/capslane@0.1.4, @webba_tech/capslane-mcp@0.1.8 and @modelcontextprotocol/sdk@1.30.0. Set CAPSLANE_API_KEY and run the file with a YouTube URL. It uses native mode; change it to auto when generation is authorized.

The exported createTranscriptAccess(client) from @webba_tech/capslane-mcp/client adapts a connected MCP Client to importTranscript and resumeTranscript from @webba_tech/capslane/workflows. Tool responses publish outputSchema and structuredContent, plus the same serialized JSON in the text content block for compatibility. The adapter checks isError before interpreting the Capslane body.

The workflow returns { url, transcript, timestampedText }. It retains jobId on storage, transport and formatting errors. Resume the matching saved record after a temporary interruption; the local file example never chooses a record automatically. Use a database with records scoped to the tenant and video in a service. Status checks consume no additional transcript unit.

Tools

Tool

Inputs

Result and usage

get_youtube_transcript

url; optional lang, mode, text, waitForCompletion

Transcript content or an accepted job. Consumes a transcript unit and can start generation.

get_transcript_status

jobId returned by the first tool

Pending state, completed content or a terminal failure. No additional transcript unit.

list_available_languages

url

Languages observed by a native transcript request. Consumes a transcript unit, including on a cache hit.

Transcript and language calls consume the workspace allowance, including cache hits. They can populate the cache, and transcript calls can start generation. Status checks do not reserve another transcript unit. The client decides how to approve each tool call.

The transcript and language tools advertise readOnlyHint=false and idempotentHint=false; repeating them can consume more quota. They advertise destructiveHint=false. The status tool remains read-only. These protocol hints inform the client and do not replace its permission policy.

Local stdio server

Use Node.js 20 or later and a client that supports this configuration format. Replace the placeholder only in your private settings. This local process still sends API requests to Capslane.

{
  "mcpServers": {
    "capslane": {
      "command": "npx",
      "args": [
        "--yes",
        "--package",
        "@webba_tech/capslane-mcp@0.1.8",
        "capslane-mcp"
      ],
      "env": {
        "CAPSLANE_API_KEY": "YOUR_API_KEY"
      }
    }
  }
}

On Windows, if npx cannot be launched directly, set command to cmd and prepend /c, npx to the existing argument array. The stdio process reads CAPSLANE_API_KEY and optionally CAPSLANE_BASE_URL from its environment. Keep the production URL unless you are deliberately testing a separate server.

Errors and limits

A 401 from /mcp means the key is absent, invalid, expired or revoked. Check the environment received by the client. An allowance error requires checking the workspace plan. Avoid repeated submissions while a job is pending.

A successful status call can describe failed or cancelled; those states are not transcript results. A completed status without content also needs investigation. Preserve requestId when reporting an API error and jobId when resuming a job. For longer videos, client output limits still apply.

Documentation and discovery

Context7 helps the assistant read the SDK documentation; Capslane MCP executes transcript requests. The coding assistant guide lists the indexed JavaScript and Python libraries. Installing this MCP server exposes its tools to your client. It does not guarantee selection for every transcript request.

MCP guide, API reference, pricing, public repository.

License

MIT

Available Tools

3 tools
get_transcript_statusGet transcript job statusA
Read-onlyIdempotent
Inspect

Check the same accepted job without consuming another transcript unit. Content means success; stop on failed, cancelled or completed without content (stored result unavailable or expired). A successful tool call can still describe a pending or failed job. Poll with a delay and deadline, never by resubmitting the video.

ParametersJSON Schema
NameRequiredDescriptionDefault
jobIdYesJob identifier returned by get_youtube_transcript

Output Schema

ParametersJSON Schema
NameRequiredDescription
langNo
errorNo
jobIdNo
cachedNo
sourceNo
statusNoJob state, never an HTTP status. Completed without content means the stored result is unavailable or expired.
contentNoTranscript itself. Check this before jobId; it can coexist with a completed job ID.
progressNo
requestIdNo
availableLangsNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, the description discloses key behavioral traits: a successful API call can still represent a pending or failed job, content indicates success, and completed-without-content means the stored result is unavailable or expired. This is meaningful operational context the annotations do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, each earning its place: the core purpose and resource cost, success/failure semantics, the pending-job caveat, and concrete polling guidance. The most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter status-checking tool with readOnly and idempotent annotations plus an output schema, the description covers the essential operational edge cases and polling strategy. Nothing critical is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the jobId parameter is fully documented as the identifier returned by get_youtube_transcript, with a pattern. The description adds the notion of checking the 'same accepted job' but does not need to add more because the schema already defines the parameter precisely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool checks the status of an already-accepted transcription job and explicitly contrasts it with resubmitting the video. This makes it easy to distinguish from get_youtube_transcript and list_available_languages.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance: poll an existing job in place of resubmitting the video. It also defines stop conditions ('failed, cancelled or completed without content') and instructs to poll with a delay and deadline, leaving no ambiguity about the intended workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_youtube_transcriptGet YouTube transcriptAInspect

Retrieve captions or generate a transcript for a public YouTube video. Each submission consumes a transcript unit, including cache hits. Can create a generation job. For summaries, notes or timestamp citations, retrieve the transcript first; this tool does not summarize.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesPublic YouTube URL or 11-character video ID
langNoPreferred ISO language code
modeNoCache first in every mode. On a miss: native never generates; auto generates only after missing captions; generate requests audio transcriptionauto
textNoPlain text for an immediate response. Completed jobs return segments with offsets and durations in milliseconds
waitForCompletionNoWait up to twenty minutes for generation. Set false for interactive clients, then poll the returned jobId with get_transcript_status

Output Schema

ParametersJSON Schema
NameRequiredDescription
langNo
errorNo
jobIdNo
cachedNo
sourceNo
statusNoJob state, never an HTTP status. Completed without content means the stored result is unavailable or expired.
contentNoTranscript itself. Check this before jobId; it can coexist with a completed job ID.
progressNo
requestIdNo
availableLangsNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses non-obvious behaviors beyond the annotations: 'Each submission consumes a transcript unit, including cache hits' and 'Can create a generation job.' These additions clarify side effects (cost, background jobs) that the annotations (readOnlyHint=false, idempotentHint=false) only imply. This is valuable transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no fluff. It front-loads the core purpose, then discloses the side-effect cost, then gives usage guidance. Every sentence adds value and the structure is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the essential context: purpose, public video requirement, unit consumption, job creation, and non-summarization. Combined with the detailed schema and annotations, an agent has everything needed to decide when and how to invoke it. The output schema exists to describe return structure, so no additional explanation is required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides 100% parameter coverage with detailed descriptions (e.g., mode explains caching and generation behavior, waitForCompletion explains polling). The description itself adds minimal parameter-specific information beyond the schema, so it does not significantly improve on the schema's already rich semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Retrieve captions or generate a transcript for a public YouTube video.' It uses a specific verb ('retrieve'/'generate') and resource ('transcript'), and differentiates itself from sibling tools by noting 'this tool does not summarize' and implying that it returns raw transcript for further processing. This distinguishes it from list_available_languages and get_transcript_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context on when to use the tool ('For summaries, notes or timestamp citations, retrieve the transcript first') and provides an exclusion ('this tool does not summarize'). However, it does not explicitly name alternative sibling tools or state when to prefer them, so the guidance is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_available_languagesList transcript languagesAInspect

Return languages observed by a native transcript request. Consumes one transcript unit and can populate the cache; it is not a free metadata lookup. Never starts generation. A cached result can have a generated source.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesPublic YouTube URL or 11-character video ID

Output Schema

ParametersJSON Schema
NameRequiredDescription
jobIdNo
statusNo
requestIdNo
selectedLangNo
availableLangsNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral context beyond the annotations: it consumes one transcript unit, can populate the cache, is not free, never starts generation, and cached results can have a generated source. This meaningfully enriches the raw annotation flags.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each adding important operational information with no redundancy. The primary action is front-loaded, followed by cost and side-effect caveats.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with a full output schema, the description covers the essential behavioral nuances: cost, cache effects, generation behavior, and the possibility of generated cached sources. Nothing critical appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents url as a public YouTube URL or 11-character video ID. The description adds no extra parameter-level detail, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Return languages observed by a native transcript request.' It also distinguishes the tool from a generic metadata lookup by stating it consumes a transcript unit, which helps differentiate it from the sibling transcript tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear contextual guidance: this is not a free metadata lookup, it consumes a transcript unit, and it never starts generation. However, it does not explicitly name alternatives or state when to prefer get_transcript_status or get_youtube_transcript instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.8
    • First observedget_transcript_status
    • First observedget_youtube_transcript
    • First observedlist_available_languages

TDQS

A4.4/5.0

Scored across 3 tools

Disambiguation4/5

The three tools have largely distinct purposes: fetching content, polling job status, and listing languages. However, get_youtube_transcript and list_available_languages both consume transcript units and populate the cache, creating mild ambiguity about when to use one versus the other.

Naming Consistency5/5

All three tools follow a clean verb_noun snake_case pattern (get_youtube_transcript, get_transcript_status, list_available_languages). Names are predictable and readable.

Tool Count4/5

Three tools for a narrowly scoped transcript-retrieval service is reasonable and each earns its place. It leans slightly thin, with no bulk or batch operation, but the domain doesn't demand more.

Completeness4/5

The surface covers the core lifecycle: retrieve transcript, poll async job status, and list available languages. Minor gaps exist (e.g., no explicit cancel/delete job operation, despite the status tool referencing cancelled jobs), but agents can work around these.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers