Skip to main content
Glama
thenavidm

ScrapeCreators MCP Server

by thenavidm

Transcript

youtube_transcript

Retrieve publicly available captions, subtitles, or transcripts from YouTube videos or Shorts, returning timestamped arrays and plain text. Specify a language code to select the caption track.

Instructions

Retrieves publicly available captions, subtitles, or transcripts from a YouTube video or Short. Returns both a timestamped transcript array with start/end times and a plain-text version in transcript_only_text. Supports specifying a language code. Videos of any length are supported when YouTube exposes public captions. This endpoint does not use the two-minute AI transcription fallback. If no matching caption track is available, the transcript fields return null. Potentially consumes paid API credits; requires confirm=true. Read-like POST requests do not publish to social platforms.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesYouTube video or short URL
accountNoNamed private ScrapeCreators account; selects credentials, not a remote account ID.
confirmNoMust be true for the specific approved credit-consuming research call.
languageNoLanguage code, ie 'en', 'es', 'fr' or 'en-US'. Overrides the default track selection unless original_audio=true. If omitted, prefers captions matching the original spoken language when YouTube identifies the original audio. If that metadata is unavailable or ambiguous, prefers an auto-generated caption, otherwise the first caption track. If the requested or identified original language has no matching captions, the transcript will be null and no credits are charged.
cache_max_ageNoIf we have a response in the cache that is this many days old or newer, return the cached response (0 credits, with "cached": true and a "cached_at" timestamp). Otherwise, scrape a live result (1 credit). [See the Caching page for details.](https://docs.scrapecreators.com/caching)
original_audioNoSet to true to return captions only in the original spoken language identified by YouTube. Takes precedence over language. If the original audio cannot be reliably identified or has no matching captions, transcript, transcript_only_text, and language are null and no credits are charged. No extra lookup or credit cost; a returned transcript costs the usual 1 credit. Omit or set to false for the existing default selection.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv2.0.0

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantive behavior beyond annotations: credit consumption, the confirm=true gate, null returns when no caption track matches, and the assurance that read-like POST requests do not publish to social platforms. Annotations already flag readOnlyHint=false and openWorldHint=true, so this context meaningfully lowers agent uncertainty about cost and side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads what is retrieved and the return shape, then adds cost and safety notes in a compact block. Slightly dense but every sentence carries information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by naming the returned fields (timestamped array with start/end times, transcript_only_text) and the null case. Combined with the credit/caching notes in the schema, an agent has enough to call it correctly, though pagination or response envelope details are not addressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so url, language, confirm, cache_max_age, and original_audio semantics are already fully documented in the schema. The description mentions only language-code support generically, adding little beyond what the schema provides; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Retrieves publicly available captions, subtitles, or transcripts from a YouTube video or Short') and distinguishes itself from sibling transcript tools by noting it does not use the two-minute AI transcription fallback. An agent can tell exactly what it gets back and under what source conditions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating context: works for videos of any length when YouTube exposes public captions, honors language selection, and requires confirm=true because it consumes credits. It does not explicitly name when to prefer a sibling (e.g. youtube_video_short_details or a different platform's transcript tool), so routing is left partly to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools