Skip to main content
Glama
DTufail

pipecat-docs-mcp

by DTufail
README.md
# pipecat-docs-mcp

> Search 317 pages of [Pipecat](https://docs.pipecat.ai) docs + 1,028 GitHub Issues from Claude Desktop or Cursor — grounded answers, no hallucinated APIs.

[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)
[![Python 3.10+](https://img.shields.io/badge/Python-3.10%2B-blue)](https://python.org)
[![FastMCP](https://img.shields.io/badge/FastMCP-2.0%2B-blueviolet)](https://github.com/jlowin/fastmcp)

---

## What it does

Indexes 8,817 chunks from the Pipecat docs and GitHub Issues, then exposes them through 4 MCP tools. Claude can answer questions about configuration, code examples, architecture concepts, and provider comparisons — with source-cited results pulled directly from the official docs and community issue threads.

**Retrieval pipeline:** BM25 + FAISS dense vectors → Reciprocal Rank Fusion → Cross-encoder reranker

| Metric | Value |
|---|---|
| Pages indexed | 317 |
| GitHub Issues indexed | 1,028 |
| Total chunks | 8,817 |
| Mean Reciprocal Rank | 0.863 |
| Recall@5 | 100% |
| Avg query latency | 35ms |

---

## Tools

| Tool | Use it when... |
|---|---|
| `search_pipecat_docs` | General questions, config options, error messages |
| `get_example_code` | You need a runnable Python pipeline |
| `explain_concept` | "What is a Frame / Pipeline / VAD / Transport?" |
| `compare_services` | Choosing between STT, TTS, or transport providers |

---

## Quickstart

### 1. Clone and install

```bash
git clone https://github.com/DTufail/pipecat-docs-mcp.git
cd pipecat-docs-mcp
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
```

### 2. Build the search index

`data/chunks.jsonl` is included. This builds the BM25 and FAISS indexes (~2 min):

```bash
python indexer.py
```

### 3. (Optional) Index GitHub Issues

Adds 1,028 issues from `pipecat-ai/pipecat` for error-driven retrieval:

```bash
export GITHUB_TOKEN="your_token_here"   # optional but recommended (5k req/hr vs 60)
python github_indexer.py
python indexer.py                        # rebuild indexes with issues merged
```

### 4. Connect to Claude Desktop

Add to `~/Library/Application Support/Claude/claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "pipecat-docs": {
      "command": "/absolute/path/to/.venv/bin/python3",
      "args": ["/absolute/path/to/pipecat-docs-mcp/server.py"]
    }
  }
}
```

Use the Python binary from your virtual environment, not system Python. Restart Claude Desktop after saving.

### 5. Connect to Cursor (HTTP)

```bash
fastmcp run server.py:mcp --transport http --port 8000
```

Point your MCP client at `http://localhost:8000/mcp`.

---

## Project structure

```
pipecat-docs-mcp/
├── server.py              # FastMCP server — 4 MCP tools
├── retrieval.py           # Hybrid search (BM25 + FAISS + RRF + reranker)
├── indexer.py             # Builds BM25, FAISS, and metadata indexes
├── github_indexer.py      # Fetches GitHub Issues → data/github_issues.jsonl
├── config.py              # Paths and model config
├── test_server.py         # In-process tests for all 4 tools
├── test_retrieval.py      # Retrieval eval — MRR, Recall@5, latency (17 queries)
├── requirements.txt
├── pipecat-scraper/
│   ├── scraper.py         # Sitemap-driven scraper (317 pages, resumable)
│   ├── utils.py           # HTML parsing and chunking
│   ├── audit.py           # Per-page extraction audit
│   └── config.py
└── data/
    ├── chunks.jsonl        # 6,936 scraped doc chunks (included)
    ├── chunk_metadata.json # Title, section, URL, content type per chunk
    ├── bm25_index.pkl      # Generated by indexer.py (gitignored)
    ├── faiss_index.bin     # Generated by indexer.py (gitignored)
    ├── embeddings.npy      # Generated by indexer.py (gitignored)
    └── github_issues.jsonl # Generated by github_indexer.py (gitignored)
```

---

## Re-scraping the docs

To refresh against a newer Pipecat release:

```bash
cd pipecat-scraper
python scraper.py      # scrapes docs.pipecat.ai → data/chunks.jsonl
cd ..
python indexer.py      # rebuilds indexes
```

---

## Roadmap

- Auto-refresh on new Pipecat releases via GitHub webhook
- Expanded coverage — LiveKit and Vapi docs

---

## License

MIT