Docs MCP
by Ruoth1111
README.md
# Docs MCP — AI-Powered Documentation Search
An MCP (Model Context Protocol) server that searches official documentation for popular libraries, scrapes the relevant pages, and returns clean, LLM-ready text. Connect it to Claude Desktop, Cursor, or any MCP-compatible client so your AI assistant can look up up-to-date docs on demand.
## What it does
The project follows a three-step pipeline:
1. **Search the web** — Uses the [Serper API](https://serper.dev/) to run a Google search scoped to a library's official documentation site.
2. **Fetch and clean pages** — Downloads the top result URLs and strips HTML noise using Groq LLM (`openai/gpt-oss-20b`).
3. **Return structured context** — Sends back cleaned text with `SOURCE:` links so the calling agent can cite where the information came from.
The MCP server exposes a single tool, `get_docs`, that AI clients can call when they need documentation context.
## Supported libraries
| Library key | Documentation site |
|---------------|---------------------------------------|
| `langchain` | python.langchain.com/docs |
| `lama-index` | docs.llamaindex.ai/en/stable |
| `openai` | platform.openai.com/docs |
| `uv` | docs.astral.sh/uv |
## Prerequisites
- **Python 3.12+**
- **[uv](https://docs.astral.sh/uv/)** — Python package and project manager
- **Serper API key** — [serper.dev](https://serper.dev/)
- **Groq API key** — [console.groq.com](https://console.groq.com/)
## Project structure
```
web-scraping/
├── mcp_server.py # MCP server (FastMCP) — exposes the get_docs tool
├── client.py # Example MCP client that calls get_docs and summarizes results
├── utils.py # LLM helpers (Groq) and HTML cleaning (trafilatura)
├── test.py # Quick sanity check for the Groq API connection
├── pyproject.toml # Project dependencies (managed by uv)
└── .env # API keys (create this locally — do not commit)
```
## Setup
### 1. Clone and enter the project
```bash
cd /path/to/web-scraping
```
### 2. Install dependencies with uv
```bash
uv sync
```
This creates a `.venv` virtual environment and installs all packages listed in `pyproject.toml`.
### 3. Create a `.env` file
Create a `.env` file in the project root with your API keys:
```env
SERPER_API_KEY=your_serper_api_key_here
GROQ_API_KEY=your_groq_api_key_here
```
Both keys are required. The server loads them via `python-dotenv`.
### 4. Verify your Groq connection (optional)
```bash
uv run test.py
```
You should see `API Key Found: True` followed by a short LLM response.
## Usage
### Run the MCP server directly
```bash
uv run mcp_server.py
```
The server communicates over **stdio** (standard input/output), which is how MCP clients connect to it. Running it standalone will appear to hang — that is expected; it is waiting for an MCP client.
### Run the example client
From the project root:
```bash
uv run client.py
```
This starts the MCP server as a subprocess, calls `get_docs` with a sample query (`"How to publish a package with uv on gitlab"`), and prints a human-readable answer generated by Groq.
## Connect to Claude Desktop
Add the server to your Claude Desktop MCP config.
**macOS config path:**
```
~/Library/Application Support/Claude/claude_desktop_config.json
```
**Recommended configuration** — use `--directory` so the server works even when Claude Desktop does not honor the `cwd` field:
```json
{
"mcpServers": {
"docs-mcp": {
"command": "/opt/homebrew/bin/uv",
"args": [
"run",
"--directory",
"/Users/YOUR_USERNAME/Desktop/web-scraping",
"mcp_server.py"
],
"env": {
"SERPER_API_KEY": "your_serper_api_key_here",
"GROQ_API_KEY": "your_groq_api_key_here"
}
}
}
}
```
Replace `/Users/YOUR_USERNAME/Desktop/web-scraping` with the absolute path to this project, and update the `uv` path if yours differs (`which uv`).
**Alternative** — run the venv Python directly (most reliable if `cwd` is ignored):
```json
{
"mcpServers": {
"docs-mcp": {
"command": "/Users/YOUR_USERNAME/Desktop/web-scraping/.venv/bin/python",
"args": [
"/Users/YOUR_USERNAME/Desktop/web-scraping/mcp_server.py"
],
"env": {
"SERPER_API_KEY": "your_serper_api_key_here",
"GROQ_API_KEY": "your_groq_api_key_here"
}
}
}
}
```
After saving the config, **fully quit** Claude Desktop (Cmd+Q) and reopen it. The `docs-mcp` server should appear as connected.
## MCP tool reference
### `get_docs`
Search official documentation for a library and return cleaned text.
**Parameters:**
| Parameter | Type | Description |
|------------|--------|--------------------------------------------------|
| `query` | string | What to search for (e.g. `"Publish a package"`) |
| `library` | string | One of: `langchain`, `lama-index`, `openai`, `uv`|
**Example call (from an MCP client):**
```json
{
"query": "How to use async with LangChain",
"library": "langchain"
}
```
**Returns:** Cleaned documentation text with `SOURCE: <url>` headers for each fetched page.
## How it works internally
```
User / AI client
│
▼
get_docs(query, library)
│
├─► Serper API ──► Google search (site-scoped to official docs)
│
├─► httpx ──► Fetch top result URLs
│
└─► Groq LLM ──► Clean HTML from each page (4000-char chunks)
│
▼
Structured text + source links
```
- **`mcp_server.py`** — Defines the FastMCP server and the `get_docs` tool.
- **`utils.py`** — `get_response_from_llm()` calls Groq; `clean_html_to_txt()` uses trafilatura (available as a fallback utility).
- **`client.py`** — Demonstrates programmatic MCP usage with the official Python MCP SDK.
## Troubleshooting
### `Failed to spawn: mcp_server.py — No such file or directory`
Claude Desktop launched `uv` from the wrong working directory. Fix by using `--directory` with an absolute project path (see config above), or point directly at `.venv/bin/python`.
### `GROQ_API_KEY not found`
Ensure `.env` exists in the project root with a valid key, or pass the key via the `env` block in your MCP config.
### Server connects then immediately disconnects
Check Claude Desktop MCP logs for Python import errors. Run `uv sync` to reinstall dependencies, then test locally with `uv run mcp_server.py`.
### `Library X not supported by this tool`
The `library` parameter must exactly match one of the supported keys: `langchain`, `lama-index`, `openai`, or `uv`.
## Dependencies
Managed in `pyproject.toml`:
- [fastmcp](https://github.com/jlowin/fastmcp) — MCP server framework
- [mcp](https://github.com/modelcontextprotocol/python-sdk) — MCP Python SDK
- [httpx](https://www.python-httpx.org/) — Async HTTP client
- [groq](https://github.com/groq/groq-python) — Groq LLM API client
- [python-dotenv](https://github.com/theskumar/python-dotenv) — Load `.env` files
- [trafilatura](https://trafilatura.readthedocs.io/) — HTML text extraction
## License
MIT License
TDQS
A3.6/5.0
Scored across 1 tool
Disambiguation5/5
Only one tool exists, so there is no potential for confusion between tools.
Naming Consistency5/5
With a single tool, naming consistency is inherently maintained; the name 'get_docs' follows a clear verb_noun pattern.
Tool Count3/5
One tool is minimal but can be acceptable for a narrowly scoped server dedicated solely to documentation search; however, it feels thin for the claimed support of multiple libraries.
Completeness3/5
The server covers the primary search operation but lacks tools for listing supported libraries, retrieving specific document content, or handling version-specific queries, which are notable gaps.
Maintenance
ActivityInactive
ResponsivenessNo issues