Social Media Video Transcripts
OfficialTurn any MCP client into a transcript and video-discovery tool for YouTube, TikTok, Instagram, X, Facebook, and direct media URLs.
Fetch transcripts (
get_transcript) from a YouTube ID/URL or TikTok, Instagram, X (Twitter), Facebook, or direct media file URLs.AI transcribe audio when no captions exist, by re-calling
get_transcriptwithai_fallback: true(skips captions, async job ~1–3 min, 1 credit on delivery).Search YouTube (
search_videos) by keyword, returning titles, IDs, and URLs (limit 1–50, default 5).List a channel's videos (
list_channel_videos) via handle, channel ID, or URL.List a playlist's videos (
list_playlist_videos) via playlist ID or URL.Check credit balance (
get_credits) for the current API key — never billed.Run locally over stdio via npx/npm/Docker, or use the hosted remote server at
https://transcriptfetch.com/mcp.Requires an API key (
TRANSCRIPTFETCH_API_KEY,tf_live_...); 50 free credits/month, failures never charged.
Allows fetching YouTube transcripts, searching videos, and listing channel and playlist videos.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Social Media Video TranscriptsGet transcript for video dQw4w9WgXcQ"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
TranscriptFetch MCP Server
A Model Context Protocol server that gives any MCP client (Claude Desktop, Cursor, and others) access to the TranscriptFetch API: a production transcript API for YouTube, TikTok and Instagram, with AI transcription when captions are missing. Fetch transcripts, search videos, list channels and playlists, and check your credit balance.
Get an API key
Create a free account at transcriptfetch.com. No card needed.
Open Dashboard, then API keys and create a key. It starts with
tf_live_.Put it in
TRANSCRIPTFETCH_API_KEYin the client configuration below.
Every account gets 50 free credits a month. Failures are free: a fetch that returns no transcript is never charged.
Runs locally over stdio and calls the TranscriptFetch API with your key. Prefer a hosted, remote server? Point your client at https://transcriptfetch.com/mcp instead (OAuth or API key). The hosted server waits inline for short-form AI transcription, so no polling is needed there.
Related MCP server: YouTube Transcript MCP Server
Tools
Tool | What it does |
| Transcript for a video. YouTube, TikTok, Instagram, or a direct media URL. Set |
| Search YouTube by keyword |
| List a YouTube channel's videos (handle, ID, or URL) |
| List a YouTube playlist's videos (ID or URL) |
| Remaining credit balance for the key. Never billed |
Pricing is per successful result: a caption fetch or a video list costs 1 credit, and AI transcription of the audio costs 1 credit per started 5 minutes of audio, charged only on delivery. Failed, blocked and empty results are never charged, which matters on short-form video where many clips have no speech at all.
Install
No install needed. Run it on demand with npx:
TRANSCRIPTFETCH_API_KEY=tf_live_... npx -y transcriptfetch-mcpOr install globally:
npm install -g transcriptfetch-mcpRequires Node 18+.
Run from source
git clone https://github.com/TranscriptFetch/mcp-server
cd mcp-server && npm install && npm run buildThen point your client at the built entrypoint with "command": "node" and
"args": ["/absolute/path/to/mcp-server/dist/index.js"].
Client configuration
Claude Desktop
Add this to claude_desktop_config.json (Settings then Developer then Edit Config):
{
"mcpServers": {
"transcriptfetch": {
"command": "npx",
"args": ["-y", "transcriptfetch-mcp"],
"env": { "TRANSCRIPTFETCH_API_KEY": "tf_live_..." }
}
}
}Cursor
Add the same block under mcpServers in your Cursor MCP settings.
Restart the client, and the five tools appear.
Example
Once connected, ask your assistant naturally:
Get the transcript for https://youtu.be/aircAruvnKk and summarize the key points.
Search YouTube for "how transformers work" and list the top 5 videos.
List the latest videos from @lexfridman and pull the transcript of the newest one.
How many TranscriptFetch credits do I have left?
The assistant picks the matching tool and works from the returned transcript or video list.
Configuration
Env var | Required | Default |
| yes | none |
| no |
|
Docker
The server speaks MCP over stdio, so there is no port to expose. -i is
required: without an attached stdin the transport closes immediately and the
container looks like it crashed.
docker build -t transcriptfetch-mcp .
docker run --rm -i -e TRANSCRIPTFETCH_API_KEY=tf_live_... transcriptfetch-mcpLinks
API docs: https://transcriptfetch.com/docs
MCP docs: https://transcriptfetch.com/docs/mcp
Node SDK: https://github.com/TranscriptFetch/transcript-api-node
Python SDK: https://github.com/TranscriptFetch/transcript-api-python
License
MIT
Available Tools
5 toolsget_creditsAInspect
Check the remaining TranscriptFetch credit balance for the current API key. Free: this call is never billed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It discloses a key behavioral trait: 'this call is never billed.' It also clarifies it applies to 'the current API key.' It does not detail response format or error behavior, but for a zero-parameter balance check, this is reasonable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences: the first states the core purpose, the second adds a valuable cost note. No filler or redundancy. Each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, zero-parameter tool, the description is sufficient. It states what it does and that it is free. It does not describe the return format or edge cases, but no output schema is provided, and the simplicity of the tool makes that less critical. It is complete enough for an agent to select and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema coverage is 100% (empty schema). The description adds no parameter meaning because none exist. For 0-parameter tools, the baseline is 4, and the description does not need to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: checking the remaining TranscriptFetch credit balance. The verb 'check' and resource 'credit balance' are specific, and this tool is distinctly different from sibling tools like get_transcript or search_videos, which fetch video data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage—you would check credits before making billed calls—but does not explicitly state when to use it versus alternatives or any exclusions. The context is clear enough for a simple utility, but it lacks direct guidance on timing or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_transcriptAInspect
Fetch the full transcript for a video. Accepts a YouTube video ID or URL, plus TikTok and Instagram video URLs and direct media file URLs. If no transcript comes back, the result says whether captions definitively do not exist (aiFallback.captionsUnavailable) and whether transcribing the audio would still work (aiFallback.available). When it does, call this tool again with ai_fallback: true.
| Name | Required | Description | Default |
|---|---|---|---|
| video | Yes | Video ID or URL, YouTube (dQw4w9WgXcQ, youtu.be/...), TikTok, Instagram, or a direct media file URL. | |
| ai_fallback | No | Skip captions and transcribe the audio with AI instead. Use this only after a previous call reported aiFallback.available. Charged 1 credit per started 5 minutes of audio (minimum 1), on delivery only; typically ~30 seconds for short videos, longer for long ones. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the burden and does reasonably well: it explains the failure path (captions may definitively not exist) and the fallback behavior, plus the two flag fields returned. It omits auth/rate-limit/billing behavior in the description itself (billing lives in the schema), leaving a small gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and input scope, then the failure/retry contract. Every sentence carries information; nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description helpfully names the returned fallback fields and the retry contract, which is the key knowledge an agent needs. It stops short of describing transcript payload shape or partial-failure cases, but it is adequate for a 2-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3; the description still adds value by naming the accepted URL forms for 'video' and, more importantly, by specifying that ai_fallback should only be set after a prior call reported aiFallback.available — sequencing that the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fetch the full transcript for a video') and immediately scopes the accepted inputs (YouTube ID/URL, TikTok, Instagram, direct media). No sibling tool covers transcripts, so the agent can tell it apart from get_credits/search_videos at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit conditional flow: if no transcript is returned, check aiFallback.captionsUnavailable/available, then call again with ai_fallback: true when available. That is clear when-to-use guidance, though it doesn't state when to prefer this tool over alternatives (e.g., listing a channel first).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_channel_videosAInspect
List recent videos for a YouTube channel. Accepts a channel handle (@name), channel ID (UC...), or URL.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results to return (1-50). Defaults to 5. | |
| channel | Yes | Channel handle, ID, or URL. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of behavioral disclosure. It adds that the tool accepts a channel handle, ID, or URL, which is useful behavioral context. However, it does not mention other behaviors such as result ordering (beyond 'recent'), pagination, rate limits, or error handling. For a read-only listing tool, this is acceptable but not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, front-loaded with the core action ('List recent videos'), and every phrase earns its place. It avoids unnecessary detail, making it easy for an agent to quickly parse the tool's purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple listing tool with full schema coverage and no output schema, the description is adequately complete. It specifies the action, the resource, and accepted input formats. It could optionally mention the default limit or suggest using sibling tools for playlists, but these are not critical for using the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: both parameters (channel and limit) are described in the input schema. The description repeats the channel format info already present in the schema and adds no new meaning for the limit parameter. Since the schema already provides full parameter semantics, a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List' and the resource 'recent videos for a YouTube channel', making the tool's function immediately obvious. It also specifies the three accepted input formats (handle, ID, or URL), which helps differentiate it from sibling tools like list_playlist_videos (playlists vs channels) and search_videos (general search vs channel-specific listing).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: if you need recent videos from a specific channel, this is the tool. However, it does not explicitly mention when to use it over alternatives (e.g., list_playlist_videos for playlists, search_videos for broader search) or provide any exclusion criteria. The guidance is not misleading, but it's only implied, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_playlist_videosAInspect
List the videos in a YouTube playlist. Accepts a playlist ID or URL.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results to return (1-50). Defaults to 5. | |
| playlist | Yes | Playlist ID or URL. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits, but it only states the action and input. It omits details like the default limit (5), pagination, error behavior, or return format, leaving the agent with little insight into the tool's runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is front-loaded with the action and resource. It is concise with no unnecessary words, and every phrase contributes to understanding the tool's purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with two parameters, the description is adequate but incomplete. It does not mention return values (no output schema) or behavioral details, and the absence of annotations increases the burden on the description to provide context, which it only partially fulfills.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides complete descriptions for both parameters (playlist and limit), so the description adds no additional semantic value beyond what the schema already provides. The phrase 'Accepts a playlist ID or URL' merely restates the schema's parameter description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List') and resource ('videos in a YouTube playlist'), using a specific verb that distinguishes it from siblings like list_channel_videos. It also mentions accepting a playlist ID or URL, further clarifying its scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly conveys when to use the tool: when you have a playlist ID or URL and want its videos. It does not explicitly name alternatives or exclusions, but the purpose is clear enough that an agent can infer the appropriate context without ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_videosAInspect
Search YouTube for videos matching a query. Returns titles, IDs, and URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results to return (1-50). Defaults to 5. | |
| query | Yes | Search keywords. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It states that the tool returns titles, IDs, and URLs, which is useful, but it does not disclose any other behavioral traits such as sorting, pagination, rate limits, or whether results are limited by relevance. This is a minimal but not misleading description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: two short sentences that immediately convey purpose and return values. Every word is functional, with no redundancy or unnecessary content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter search tool with no output schema, the description adequately covers the main purpose and return format. However, it lacks guidance on when to use it versus sibling tools, which slightly reduces completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides full descriptions for both parameters (query and limit), covering 100% of the schema. The description adds no additional parameter-level detail, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function with a specific verb ('Search') and resource ('YouTube'), and specifies the scope ('videos matching a query'). It implicitly distinguishes from siblings like list_channel_videos and get_transcript, which focus on specific retrieval rather than general search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is provided on when to use this tool versus the sibling tools. The description does not mention alternatives, exclusions, or scenarios like 'use list_channel_videos to get videos from a specific channel.' The usage context must be inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.2.8- Changed
get_transcript2 fields changed- changed
Input schema / properties / ai_fallback / descriptionPrevious value: -"Skip captions and transcribe the audio with AI instead. Use this only after a previous call reported aiFallback.available, it starts an async job (1 credit on delivery) that takes 1-3 minutes."New value: +"Skip captions and transcribe the audio with AI instead. Use this only after a previous call reported aiFallback.available. Charged 1 credit per started 5 minutes of audio (minimum 1), on delivery only; typically ~30 seconds for short videos, longer for long ones." - changed
Input schema / properties / video / descriptionPrevious value: -"Video ID or URL, YouTube (dQw4w9WgXcQ, youtu.be/...), TikTok, Instagram, X, Facebook, or a direct media file URL."New value: +"Video ID or URL, YouTube (dQw4w9WgXcQ, youtu.be/...), TikTok, Instagram, or a direct media file URL."
5 tool updates
v0.1.0- First observed
get_credits - First observed
get_transcript - First observed
list_channel_videos - First observed
list_playlist_videos - First observed
search_videos
TDQS
Scored across 5 tools
Each tool targets a distinct resource: credit balance, channel video listing, playlist video listing, search results, and transcript fetching. The three discovery tools (search_videos, list_channel_videos, list_playlist_videos) are cleanly separated by their input source, so an agent can easily pick the right one.
All five tools follow a consistent verb_noun snake_case pattern (get_credits, list_channel_videos, list_playlist_videos, search_videos, get_transcript). No deviations or mixed conventions.
Five tools is well-scoped for a transcript-fetching service, covering discovery, retrieval, and account status without redundancy. Every tool earns its place.
Core workflow (discover videos via search/channel/playlist, then fetch transcript) is fully covered, plus a credits check and a documented AI-fallback path. Minor gap: no dedicated video-details/metadata tool, though get_transcript partially compensates by accepting IDs/URLs.
Maintenance
Related MCP Connectors
YouTube transcripts, search, channel/playlist listings and upload tracking for AI agents.
YouTube discovery, transcripts, library search, and monitors with API keys or OAuth.
Transcripts of YouTube videos, playlists and channels with timestamps; SRT or WebVTT too.
Search YouTube, read video metadata, and fetch transcripts with language preferences
Related MCP Servers
- AlicenseAqualityFmaintenanceA Model Context Protocol server that enables retrieval of transcripts from YouTube videos. This server provides direct access to video captions and subtitles through a simple interface.11,107 npm599MIT
- FlicenseBqualityNot gradedmaintenanceEnables extraction and processing of YouTube video transcripts from individual videos, channels, and playlists. Supports transcript search, batch processing, multiple output formats (JSON, text, SRT, VTT), and bulk operations across multiple videos.1134 npm-
- AlicenseAqualityAmaintenanceAn MCP server (stdio + HTTP/SSE) that fetches video transcripts/subtitles via yt-dlp, with pagination for large responses. Supports YouTube, Twitter/X, Instagram, TikTok, Twitch, Vimeo, Facebook, Bilibili, VK, Dailymotion. Whisper fallback — transcribes audio when subtitles are unavailable (local or OpenAI API). Works with Cursor and other MCP host822MIT
- AlicenseAqualityDmaintenanceEnables fetching, searching, and summarizing YouTube video transcripts with multi-language support.4MIT