web-summaries-mcp
by extempore66
README.md
# web-summaries-mcp
A Streamable HTTP MCP (Model Context Protocol) server that fetches YouTube transcripts and lets an LLM semantically search within them.
## Tools
- **`get_transcript`** — fetches the full transcript of a YouTube video, given a video ID or URL (`watch`, `youtu.be`, `embed`, `shorts` links all work).
- **`search_transcript`** — semantically searches a transcript by meaning. Chunks the transcript, embeds each chunk and the query with a local sentence-embedding model (`Xenova/bge-base-en-v1.5` via `@huggingface/transformers`), and ranks chunks by cosine similarity (dot product on normalized vectors). Returns the top-k matching passages with timestamps.
Transcripts are fetched via YouTube's InnerTube API (the same one the YouTube Android app uses) and parsed from the caption track XML.
## Setup
```bash
npm install
npm start
```
The server listens on `http://localhost:8080/mcp` (Streamable HTTP transport, stateless per the 2026-07-28 MCP spec — every request is self-contained, no session handshake).
## Project layout
- `mcp_server.js` — Express + `@modelcontextprotocol/server`/`@modelcontextprotocol/node` (MCP SDK v2) server, tool definitions, embedding/search logic. See `MIGRATION.md` for details on the v1→v2 migration.
- `transcript_core.js` — InnerTube fetch + transcript XML parsing, shared by the server and the CLI script.
- `fetch_transcript.js` — standalone CLI for downloading a transcript to a file.
- `htmlroot/` — static assets served alongside the MCP endpoint.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues