Skip to main content
Glama

BulkTranscripts YouTube

Start a bulk extraction job

start_bulk_extract

Extract every transcript of a YouTube channel, playlist or video list in the background. Returns immediately with a run_id; the server discovers the videos (free), fetches each transcript into this account's library (1 credit per new transcript, cache hits and repeats free, failures never charged) and keeps going after this call returns. Poll get_bulk_status no sooner than its check_again_in_seconds, then page outcomes with list_bulk_results and read any transcript with get_transcript (free once it is in the library). Up to 1,000 videos per job, 2 jobs per account at a time. Prefer this over repeated get_transcripts calls for anything larger than about 20 videos.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesChannel (@handle or URL), playlist URL/id, or a single video URL.
languageNoPreferred caption language code. Default 'en'.
max_videosNoCap on videos to process, default 100.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
errorNo
phaseNo
quotaNoVideos not attempted because credits ran out.
totalNo
cachedNo
failedNo
run_idYes
sourceYes
statusYesrunning, completed or stopped.
billingYes
skippedNoVideos without captions; never charged.
stoppedNo
completedNo
remainingNo
interruptedNo
videos_foundNoNull until a large channel has been listed.
check_again_in_secondsYesDo not poll sooner than this.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "properties": {
      +    "billing": {
      +      "properties": {
      +        "creditsCharged": {
      +          "description": "Credits this call cost (0 on cache hits).",
      +          "type": "integer"
      +        },
      +        "enabled": {
      +          "type": "boolean"
      +        },
      +        "freeLimit": {
      +          "type": "integer"
      +        },
      +        "granted": {
      +          "type": "integer"
      +        },
      +        "kind": {
      +          "description": "anon (free tier), license (paid) or admin.",
      +          "type": "string"
      +        },
      +        "remaining": {
      +          "description": "Credits left on this account.",
      +          "type": [
      +            "integer",
      +            "null"
      +          ]
      +        },
      +        "unlimited": {
      +          "type": "boolean"
      +        },
      +        "used": {
      +          "type": "integer"
      +        }
      +      },
      +      "required": [
      +        "enabled"
      +      ],
      +      "type": "object"
      +    },
      +    "cached": {
      +      "type": "integer"
      +    },
      +    "check_again_in_seconds": {
      +      "description": "Do not poll sooner than this.",
      +      "type": "integer"
      +    },
      +    "completed": {
      +      "type": "integer"
      +    },
      +    "error": {
      +      "type": [
      +        "string",
      +        "null"
      +      ]
      +    },
      +    "failed": {
      +      "type": "integer"
      +    },
      +    "interrupted": {
      +      "type": "integer"
      +    },
      +    "phase": {
      +      "type": "string"
      +    },
      +    "quota": {
      +      "description": "Videos not attempted because credits ran out.",
      +      "type": "integer"
      +    },
      +    "remaining": {
      +      "type": "integer"
      +    },
      +    "run_id": {
      +      "type": "string"
      +    },
      +    "skipped": {
      +      "description": "Videos without captions; never charged.",
      +      "type": "integer"
      +    },
      +    "source": {
      +      "properties": {
      +        "title": {
      +          "type": [
      +            "string",
      +            "null"
      +          ]
      +        },
      +        "type": {
      +          "type": [
      +            "string",
      +            "null"
      +          ]
      +        },
      +        "url": {
      +          "type": "string"
      +        }
      +      },
      +      "required": [
      +        "url"
      +      ],
      +      "type": "object"
      +    },
      +    "status": {
      +      "description": "running, completed or stopped.",
      +      "type": "string"
      +    },
      +    "stopped": {
      +      "properties": {
      +        "code": {
      +          "type": "string"
      +        },
      +        "message": {
      +          "type": "string"
      +        },
      +        "not_attempted": {
      +          "type": "integer"
      +        }
      +      },
      +      "required": [
      +        "code"
      +      ],
      +      "type": "object"
      +    },
      +    "total": {
      +      "type": "integer"
      +    },
      +    "videos_found": {
      +      "description": "Null until a large channel has been listed.",
      +      "type": [
      +        "integer",
      +        "null"
      +      ]
      +    }
      +  },
      +  "required": [
      +    "run_id",
      +    "status",
      +    "source",
      +    "check_again_in_seconds",
      +    "billing"
      +  ],
      +  "type": "object"
      +}
  2. Added

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond annotations: returns immediately with a run_id, work continues server-side after the call, the credit model (1 credit per new transcript, cache hits free, failures never charged), and concurrency/volume limits (1,000 videos, 2 jobs per account). Annotations only cover the readOnly/openWorld/idempotent profile, so this is genuinely additive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: purpose and async behavior come first, then the workflow, then limits and the sibling comparison. Every sentence carries information, though the single-paragraph packing of polling, pricing and quotas is on the heavy side.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the description still surfaces the run_id the agent needs to proceed. For an async, credit-consuming job with concurrency limits, nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents url, language and max_videos including ranges and defaults. The description restates the 1,000-video cap but adds no syntax or format detail beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Extract every transcript of a YouTube channel, playlist or video list') with clear scope (background job, up to 1,000 videos). It explicitly names the sibling it replaces ('repeated get_transcripts calls'), so an agent can distinguish it without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('anything larger than about 20 videos') and a full follow-up workflow: poll get_bulk_status no sooner than check_again_in_seconds, page with list_bulk_results, read with get_transcript. Alternatives and their selection conditions are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources