pipecat-docs-mcp
pipecat-docs-mcp
Search 317 pages of Pipecat docs + 1,028 GitHub Issues from Claude Desktop or Cursor — grounded answers, no hallucinated APIs.
What it does
Indexes 8,817 chunks from the Pipecat docs and GitHub Issues, then exposes them through 4 MCP tools. Claude can answer questions about configuration, code examples, architecture concepts, and provider comparisons — with source-cited results pulled directly from the official docs and community issue threads.
Retrieval pipeline: BM25 + FAISS dense vectors → Reciprocal Rank Fusion → Cross-encoder reranker
Metric | Value |
Pages indexed | 317 |
GitHub Issues indexed | 1,028 |
Total chunks | 8,817 |
Mean Reciprocal Rank | 0.863 |
Recall@5 | 100% |
Avg query latency | 35ms |
Tools
Tool | Use it when... |
| General questions, config options, error messages |
| You need a runnable Python pipeline |
| "What is a Frame / Pipeline / VAD / Transport?" |
| Choosing between STT, TTS, or transport providers |
Quickstart
1. Clone and install
git clone https://github.com/DTufail/pipecat-docs-mcp.git
cd pipecat-docs-mcp
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt2. Build the search index
data/chunks.jsonl is included. This builds the BM25 and FAISS indexes (~2 min):
python indexer.py3. (Optional) Index GitHub Issues
Adds 1,028 issues from pipecat-ai/pipecat for error-driven retrieval:
export GITHUB_TOKEN="your_token_here" # optional but recommended (5k req/hr vs 60)
python github_indexer.py
python indexer.py # rebuild indexes with issues merged4. Connect to Claude Desktop
Add to ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"pipecat-docs": {
"command": "/absolute/path/to/.venv/bin/python3",
"args": ["/absolute/path/to/pipecat-docs-mcp/server.py"]
}
}
}Use the Python binary from your virtual environment, not system Python. Restart Claude Desktop after saving.
5. Connect to Cursor (HTTP)
fastmcp run server.py:mcp --transport http --port 8000Point your MCP client at http://localhost:8000/mcp.
Project structure
pipecat-docs-mcp/
├── server.py # FastMCP server — 4 MCP tools
├── retrieval.py # Hybrid search (BM25 + FAISS + RRF + reranker)
├── indexer.py # Builds BM25, FAISS, and metadata indexes
├── github_indexer.py # Fetches GitHub Issues → data/github_issues.jsonl
├── config.py # Paths and model config
├── test_server.py # In-process tests for all 4 tools
├── test_retrieval.py # Retrieval eval — MRR, Recall@5, latency (17 queries)
├── requirements.txt
├── pipecat-scraper/
│ ├── scraper.py # Sitemap-driven scraper (317 pages, resumable)
│ ├── utils.py # HTML parsing and chunking
│ ├── audit.py # Per-page extraction audit
│ └── config.py
└── data/
├── chunks.jsonl # 6,936 scraped doc chunks (included)
├── chunk_metadata.json # Title, section, URL, content type per chunk
├── bm25_index.pkl # Generated by indexer.py (gitignored)
├── faiss_index.bin # Generated by indexer.py (gitignored)
├── embeddings.npy # Generated by indexer.py (gitignored)
└── github_issues.jsonl # Generated by github_indexer.py (gitignored)Re-scraping the docs
To refresh against a newer Pipecat release:
cd pipecat-scraper
python scraper.py # scrapes docs.pipecat.ai → data/chunks.jsonl
cd ..
python indexer.py # rebuilds indexesRoadmap
Auto-refresh on new Pipecat releases via GitHub webhook
Expanded coverage — LiveKit and Vapi docs
License
MIT