Skip to main content
Glama
JADRT22
by JADRT22
README.md
# yt-transcript

Give your AI agent **eyes for YouTube**. A lightweight [MCP](https://modelcontextprotocol.io) server + CLI that fetches the **transcript, description, and comments** of any YouTube video — so your agent can read, summarize, and analyze video content.

Works with **opencode, Claude Code, Freebuff, Cursor, and any MCP-compatible client** — or as a plain CLI any agent can call.

## Why this one?

YouTube aggressively blocks scripted access ("confirm you're not a bot", HTTP 429). Most transcript tools break the first time YouTube pushes back. This one ships a **4-layer fallback cascade** baked in:

| Layer | What it does |
|---|---|
| 1. `youtube-transcript-api` | Fast path, no API key needed |
| 2. yt-dlp web client | Baseline fallback |
| 3. yt-dlp android client | Bypasses the "not a bot" check on most videos |
| 4. yt-dlp + browser cookies | Cures hard IP blocks / 429s (uses your Firefox/Chrome cookies) |

Plus:

- **Disk cache** (`~/.cache/yt-transcript/`) — second call for the same video is ~10x faster and consumes zero rate limit
- **Description + comments** — likes, authors, replies; great context for richer summaries
- **Language aware** — pick preferred languages; lists everything available per video
- **Zero config, zero API keys**

## Install

```bash
git clone https://github.com/JADRT22/yt-transcript.git
cd yt-transcript
python3 -m venv .venv
.venv/bin/pip install -e .
```

## MCP setup

Point your client at the server binary:

```
/absolute/path/to/yt-transcript/.venv/bin/yt-transcript-mcp
```

<details>
<summary><b>opencode</b></summary>

In `~/.config/opencode/opencode.json` (global) or `opencode.json` (project):

```json
{
  "mcp": {
    "yt-transcript": {
      "type": "local",
      "command": ["/absolute/path/to/yt-transcript/.venv/bin/yt-transcript-mcp"],
      "enabled": true
    }
  }
}
```
</details>

<details>
<summary><b>Claude Code</b></summary>

```bash
claude mcp add yt-transcript -- /absolute/path/to/yt-transcript/.venv/bin/yt-transcript-mcp
```
</details>

<details>
<summary><b>Freebuff (and other agents with mcp.json)</b></summary>

Create a `mcp.json` in the project root where the agent runs:

```json
{
  "mcpServers": {
    "yt-transcript": {
      "command": "/absolute/path/to/yt-transcript/.venv/bin/yt-transcript-mcp"
    }
  }
}
```
</details>

<details>
<summary><b>No MCP? Use the CLI</b></summary>

Any agent that runs shell commands can read the output directly:

```bash
/absolute/path/to/yt-transcript/.venv/bin/yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID"
```
</details>

## MCP tools

| Tool | Description |
|---|---|
| `get_transcript_tool(url, languages?)` | Full transcript of the video |
| `list_transcripts_tool(url)` | Available captions (languages, auto-generated?) |
| `get_video_info_tool(url)` | Title, channel, duration, upload date |
| `get_video_details_tool(url, max_comments?)` | **Full description + comments** (author, text, likes) |

Example prompts once connected:

```
Summarize this video: https://www.youtube.com/watch?v=...

What are people saying in the comments of this video? [URL]

Compare what the video claims with what the description promises: [URL]
```

## CLI usage

```bash
# Full transcript (default languages: en, es, pt)
.venv/bin/yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID"

# Metadata only
.venv/bin/yt-transcript VIDEO_ID --info

# Description + comments (JSON)
.venv/bin/yt-transcript VIDEO_ID --desc --max-comments 50

# Available captions
.venv/bin/yt-transcript VIDEO_ID --list

# JSON output / other languages
.venv/bin/yt-transcript VIDEO_ID --json
.venv/bin/yt-transcript VIDEO_ID -l de,fr
```

Accepts `youtube.com/watch?v=...`, `youtu.be/...`, `youtube.com/shorts/...`, or the bare 11-character ID.

## Configuration

| Environment variable | Default | Purpose |
|---|---|---|
| `YT_TRANSCRIPT_COOKIES_BROWSER` | `firefox` | Browser for the cookie fallback (`chrome`, `chromium`, `brave`, ...; `none` disables) |
| `YT_TRANSCRIPT_CACHE` | `~/.cache/yt-transcript` | Cache directory (`none` disables) |

To force a fresh fetch for one video, delete its file (SHA-256-named) from the cache directory.

## How it works

1. `extract_video_id` normalizes any YouTube URL format into a video ID.
2. Transcripts: captions API first; on any failure, yt-dlp downloads the `.vtt` through the client cascade and converts it to plain text (each attempt only counts if the caption file is actually written — yt-dlp's simulation mode silently skips it, and `--no-simulate` guards that).
3. Metadata/details: `yt-dlp --print` / `--dump-json --write-comments` through the same cascade.
4. Everything is cached on disk; agents can retry cheaply.

## Limitations

- Videos with **no captions at all** can't be transcripted (no Whisper/ASR — this tool stays lightweight on purpose).
- Browser-cookie access requires the target browser installed locally; in logged-in sessions YouTube sometimes serves empty format lists to yt-dlp, which is exactly why cookies are the **last** layer (and always paired with `--ignore-no-formats-error`).
- Cache is content-blind: if a video's description/comments change, clear the cache entry.

## License

[MIT](LICENSE) © 2026 JADRT22

TDQS

B3.4/5.0

Scored across 4 tools

Disambiguation3/5

get_video_details_tool and get_video_info_tool overlap heavily, both returning title, channel and duration; the distinction (details adds description/comments, info adds upload date) is only implicit. get_transcript_tool and list_transcripts_tool are clearly separable, but the two metadata tools risk misselection.

Naming Consistency5/5

All four tools follow a consistent verb_noun_tool pattern (get_video_details_tool, get_transcript_tool, list_transcripts_tool, get_video_info_tool) with uniform snake_case. Naming is highly predictable.

Tool Count4/5

Four tools are well-scoped for a YouTube transcript/metadata server, each targeting a distinct operation. Slightly thin, and the count is partly inflated by two near-duplicate metadata tools that could be merged.

Completeness4/5

The surface covers listing available transcripts, fetching a transcript, retrieving metadata, and gathering description/comments, which spans the core domain lifecycle. No update/delete is needed for a read-only extractor, so coverage is solid with only minor gaps (e.g. no bulk or search operation).

Maintenance

ActivityMaintained
ResponsivenessNo issues