youtube-transcript-mcp
<div align="center">
# youtube-transcript-mcp
**Turn any YouTube video into a Markdown transcript your LLM can actually afford to read.**
[](https://github.com/pobibibi/youtube-transcript-mcp/actions/workflows/ci.yml)
[](LICENSE)
[](https://nodejs.org)
[](tsconfig.json)
[](https://modelcontextprotocol.io)
<img src="assets/demo.gif" alt="Transcribing a video and then reading only the chapter that answers the question" width="880">
</div>
---
## The problem
Ask an assistant about a 40-minute talk and it has to swallow the whole thing: ten thousand words of context, most of which answer nothing. Pay for it again next time you ask.
**youtube-transcript-mcp** puts the transcript on disk and hands back a receipt.
```
Transcript saved to: transcripts/But what is a neural network.md
Title: But what is a neural network? | Deep learning chapter 1
Channel: 3Blue1Brown · 18:40 · manual en captions
Words: 3,357 Sections: 12
- 0:00 — Introduction example
- 1:07 — Series preview
- 2:42 — What are neurons?
...
```
That summary is about 50 tokens. The model reads the file when it needs the content — and because the index lists every chapter with its offset, it can read *only* the chapter that answers the question. A 3,357-word video costs 50 tokens to know about and a few hundred to answer from.
The transcription itself costs nothing: it happens on your machine, not in the model.
---
## What the Markdown looks like
```markdown
---
title: "But what is a neural network? | Deep learning chapter 1"
channel: "3Blue1Brown"
url: "https://www.youtube.com/watch?v=aircAruvnKk"
video_id: "aircAruvnKk"
duration: "18:40"
published: 2017-10-05
language: "en"
captions: manual
words: 3357
sections: 12
generated: 2026-08-05
---
# But what is a neural network? | Deep learning chapter 1
> **3Blue1Brown** · 18:40 · 2017-10-05 · [Watch on YouTube](...)
>
> Transcribed from manual `en` captions.
## Index
- [0:00](...&t=0) — Introduction example
- [1:07](...&t=67) — Series preview
- [2:42](...&t=162) — What are neurons?
## Transcript
### [0:00](...&t=0) Introduction example
**[0:04]** This is a 3. It's sloppily written and rendered at an extremely low
resolution of 28x28 pixels, but your brain has no trouble recognizing it as a 3...
```
Every design choice in that file exists to make it cheap to navigate:
| Choice | Why |
|---|---|
| YAML front matter | The model can cite the source without opening anything else. |
| Author's chapters as `###` sections | Real semantic boundaries, written by someone who watched the video. |
| No chapters → 5-minute blocks | Still gives the model somewhere to aim. |
| `&t=` links on every heading | One click opens YouTube at that exact second. |
| Timestamps per paragraph | An answer can point at *when* something was said. |
| Paragraphs, not caption lines | Captions break every 3 seconds; prose doesn't. |
---
## Install
Requires **Node.js ≥ 18** and **[yt-dlp](https://github.com/yt-dlp/yt-dlp)**.
```bash
python -m pip install -U yt-dlp
```
```bash
git clone https://github.com/pobibibi/youtube-transcript-mcp.git
cd youtube-transcript-mcp
npm install
npm run build
```
If yt-dlp is not on your `PATH`, set `YTDLP_PATH` to the executable. The server also falls back to `python -m yt_dlp`.
> Keep yt-dlp current. It is the only component that talks to YouTube, and the only one that breaks when YouTube changes something.
### Connect it
**Claude Code**
```bash
claude mcp add youtube-transcript --scope user -- node /absolute/path/to/youtube-transcript-mcp/dist/index.js
```
**Claude Desktop** — in `claude_desktop_config.json`:
```json
{
"mcpServers": {
"youtube-transcript": {
"command": "node",
"args": ["/absolute/path/to/youtube-transcript-mcp/dist/index.js"],
"env": {
"TRANSCRIPTS_DIR": "/where/you/want/the/markdown"
}
}
}
}
```
Any MCP client with stdio transport works the same way.
---
## Use
Just ask:
> Transcribe https://www.youtube.com/watch?v=aircAruvnKk
> What does this video say about activation functions? https://youtu.be/aircAruvnKk
In the second case the model transcribes, reads the index, and opens only the section that answers you.
### `transcribe_video`
| Parameter | Type | Default | Description |
|---|---|---|---|
| `url` | string | — | Video URL or bare ID. `youtube.com/watch`, `youtu.be` and shorts all work. |
| `language` | string | the video's own | Caption language code: `en`, `es`, `pt-BR`… |
| `file_name` | string | video title | Output filename, without extension. |
| `folder` | string | `TRANSCRIPTS_DIR` or `./transcripts` | Where to write the `.md`. |
| `timestamps` | boolean | `true` | Prefix each paragraph with `[mm:ss]`. |
| `include_description` | boolean | `true` | Include the author's description. |
| `block_seconds` | number | `300` | Block size when the video has no chapters. |
| `return_text` | boolean | `false` | Also return the full Markdown. Expensive; rarely what you want. |
### `list_languages`
Reports the video's own language, its duration, its chapter count, and every caption track it carries. Useful before transcribing when you are not sure the language you want exists.
---
## How it works
```
URL
│
├─ yt-dlp --dump-single-json ──────► metadata: title, chapters, caption tracks
│
├─ pickTrack() ───────────────────► which language, manual or automatic
│
├─ yt-dlp --write-subs --sub-format json3
│ └─ parseJson3() ─────────► cues, rollup duplicates removed
│
├─ buildMarkdown()
│ ├─ sections from chapters, or fixed time blocks
│ └─ paragraphs, clipped to section boundaries
│
└─ write .md ─────────────────────► return path + summary
```
Four modules, each with one job: [`ytdlp.ts`](src/ytdlp.ts) talks to YouTube, [`subtitles.ts`](src/subtitles.ts) parses caption formats, [`markdown.ts`](src/markdown.ts) shapes the document, [`transcribe.ts`](src/transcribe.ts) orchestrates. [`index.ts`](src/index.ts) is only the MCP surface.
### Three decisions worth explaining
**json3 over VTT.** YouTube's automatic captions emit every line twice — once as a "rollup" fragment, once complete. In VTT you have to guess which is which; in json3 the duplicate carries an `aAppend` flag and can be dropped exactly. VTT parsing stays as a fallback for the handful of manual tracks yt-dlp cannot hand over as json3.
**No language guessing.** YouTube will machine-translate its machine transcription into ~150 languages on request. Stacking those two produces text the speaker never said. Without an explicit `language`, the server stays on the video's own language and prefers human-written captions when they exist.
**Paragraphs never cross a chapter boundary.** Paragraphs are built *within* each section, not sliced afterwards. Cutting the other way around files a chapter's opening sentences under the previous chapter and leaves the chapter itself looking empty — which is exactly wrong for a document meant to be read one section at a time.
---
## Limitations
- **No captions, no transcript.** This reads subtitles; it does not transcribe audio. Videos with automatic captions disabled cannot be processed.
- **Automatic captions are rough.** Unreliable punctuation, mangled proper nouns and technical terms. The generated file says so in its own header.
- **Live streams** must finish first.
- **HTTP 429.** YouTube rate-limits bursts. The server backs off once and then explains what to do; browser cookies via `YTDLP_COOKIES_FROM_BROWSER` usually settle it.
## Environment variables
| Variable | Purpose |
|---|---|
| `TRANSCRIPTS_DIR` | Default output folder. |
| `YTDLP_PATH` | Path to the yt-dlp executable. |
| `YTDLP_COOKIES_FROM_BROWSER` | Browser to pull cookies from (`chrome`, `firefox`…). Needed for age-restricted videos. |
| `YTDLP_COOKIES_FILE` | Same, from a `cookies.txt` file. |
| `YTDLP_PROXY` | Proxy for all requests. |
---
## Development
```bash
npm test # 21 offline unit tests
npm run dev # hot reload
npm run smoke # end-to-end over stdio, hits the network
node test/smoke.mjs "https://youtu.be/VIDEO_ID"
python assets/make_demo.py # regenerate the demo GIF
```
The unit tests are deliberately offline — parsing, paragraphing and track selection are pure functions, so CI never depends on YouTube being reachable. `npm run smoke` starts the compiled server and speaks JSON-RPC to it, exactly as a client does.
## Notice
This downloads publicly available subtitles, the same way yt-dlp does. Use it for notes, study and research, respecting YouTube's terms of service and the rights of the video's author. Generating a transcript does not make it yours.
## License
MIT © [pobibibi](https://github.com/pobibibi)
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: transcribe_video retrieves and saves a transcript, while list_languages provides metadata about available captions. No overlap or ambiguity exists.
Both tools follow a consistent verb_noun pattern: transcribe_video and list_languages. The naming is clear, predictable, and uniform.
With only two tools, the server feels slightly thin even for its narrow domain. However, the two tools cover the core workflow of checking languages and transcribing, so the count is borderline but defensible.
The server covers the essential transcript retrieval cycle well. A minor gap is that transcribe_video does not return the transcript directly, requiring a follow-up file read, but the workflow is complete and workable.