Skip to main content
Glama

BulkTranscripts YouTube

Get YouTube video transcript

get_transcript

Fetch the full transcript of one YouTube video as clean text with metadata (title, channel, duration, upload date, language). Accepts a watch URL, youtu.be link, Shorts URL, or bare 11-character video id. Costs 1 credit the first time it is added to this account's library; repeat reads are free. Set include_segments to true only when per-line timestamps are needed (much larger output).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
freshNoBypass the cache and re-extract (costs a credit). Default false.
videoYesYouTube video URL or 11-character video id.
languageNoPreferred caption language code, e.g. 'en' or 'de'. Defaults to 'en', falling back to whatever exists.
include_segmentsNoInclude the timestamped segment list. Default false — the plain text is usually what you want.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYes
textYes
titleYes
cachedNoTrue when served from this account's library (free).
sourceNomanual_caption or auto_caption.
billingYes
channelYes
durationNo
languageNo
segmentsNoOnly when include_segments is true.
video_idYes
paragraphsYesSilence-grouped paragraphs; best for reading and chunking.
word_countNo
upload_dateNo
duration_textNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • changedInput schema / properties / video / description
      Previous value: -"YouTube video URL or 11-character video id (TikTok video URLs also work)."New value: +"YouTube video URL or 11-character video id."
  2. Changed1 schema field changed
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "properties": {
      +    "billing": {
      +      "properties": {
      +        "creditsCharged": {
      +          "description": "Credits this call cost (0 on cache hits).",
      +          "type": "integer"
      +        },
      +        "enabled": {
      +          "type": "boolean"
      +        },
      +        "freeLimit": {
      +          "type": "integer"
      +        },
      +        "granted": {
      +          "type": "integer"
      +        },
      +        "kind": {
      +          "description": "anon (free tier), license (paid) or admin.",
      +          "type": "string"
      +        },
      +        "remaining": {
      +          "description": "Credits left on this account.",
      +          "type": [
      +            "integer",
      +            "null"
      +          ]
      +        },
      +        "unlimited": {
      +          "type": "boolean"
      +        },
      +        "used": {
      +          "type": "integer"
      +        }
      +      },
      +      "required": [
      +        "enabled"
      +      ],
      +      "type": "object"
      +    },
      +    "cached": {
      +      "description": "True when served from this account's library (free).",
      +      "type": "boolean"
      +    },
      +    "channel": {
      +      "type": "string"
      +    },
      +    "duration": {
      +      "type": [
      +        "number",
      +        "null"
      +      ]
      +    },
      +    "duration_text": {
      +      "type": [
      +        "string",
      +        "null"
      +      ]
      +    },
      +    "language": {
      +      "type": [
      +        "string",
      +        "null"
      +      ]
      +    },
      +    "paragraphs": {
      +      "description": "Silence-grouped paragraphs; best for reading and chunking.",
      +      "items": {
      +        "type": "string"
      +      },
      +      "type": "array"
      +    },
      +    "segments": {
      +      "description": "Only when include_segments is true.",
      +      "items": {
      +        "properties": {
      +          "duration": {
      +            "type": "number"
      +          },
      +          "start": {
      +            "type": "number"
      +          },
      +          "text": {
      +            "type": "string"
      +          }
      +        },
      +        "required": [
      +          "text",
      +          "start"
      +        ],
      +        "type": "object"
      +      },
      +      "type": "array"
      +    },
      +    "source": {
      +      "description": "manual_caption or auto_caption.",
      +      "type": [
      +        "string",
      +        "null"
      +      ]
      +    },
      +    "text": {
      +      "type": "string"
      +    },
      +    "title": {
      +      "type": "string"
      +    },
      +    "upload_date": {
      +      "type": [
      +        "string",
      +        "null"
      +      ]
      +    },
      +    "url": {
      +      "type": "string"
      +    },
      +    "video_id": {
      +      "type": "string"
      +    },
      +    "word_count": {
      +      "type": "integer"
      +    }
      +  },
      +  "required": [
      +    "video_id",
      +    "url",
      +    "title",
      +    "channel",
      +    "text",
      +    "paragraphs",
      +    "billing"
      +  ],
      +  "type": "object"
      +}
  3. First observed

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Explains the non-obvious economics that justify readOnlyHint=false: the first fetch costs 1 credit because the video is added to the account's library, while repeat reads are free. It also discloses that include_segments produces a much larger payload and that 'fresh' bypasses the cache. This adds real context beyond the annotations rather than contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, then accepted inputs, then cost and the segments caveat. Every sentence carries information the agent needs; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description still names the metadata fields returned. Input formats, credit cost, caching behavior, and the segments trade-off are all covered for a 4-parameter read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: it enumerates accepted video input formats (watch URL, youtu.be, Shorts, bare 11-char id) and warns that include_segments substantially increases output size. The cost implications of a repeat fetch are also clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch the full transcript of one YouTube video as clean text') and scopes it to a single video, which cleanly separates it from the bulk sibling get_transcripts. The accepted input formats and the returned metadata are spelled out, so an agent knows exactly what it gets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditional guidance for include_segments ('only when per-line timestamps are needed') and implies single-video scope versus the plural get_transcripts sibling. It stops short of explicitly naming that sibling as the alternative for multi-video work, so it is context-rich but not fully routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources