Web Search MCP Server
by hundare
README.md
# Web Search MCP Server
A **production-ready Model Context Protocol (MCP) Server** that acts as a universal web search and content retrieval tool for AI agents.
## Features
- š **Universal search** ā any topic, any language query
- š **Multi-provider** ā Tavily, Brave, Bing, SerpAPI, Google CSE (pluggable)
- š **Content extraction** ā HTTP + BeautifulSoup (primary), Playwright (fallback for SPAs)
- ā” **Async-first** ā parallel page fetching, connection pooling
- šļø **Caching** ā in-memory TTL cache to save API quota
- š **Retry logic** ā exponential back-off via tenacity
- š **Structured JSON** ā Pydantic v2 models, MCP-compliant output
- šŖµ **Structured logging** ā JSON or text format
---
## Project Structure
```
web_search_mcp/
ā
āāā server.py ā FastMCP server + tool registration
āāā tools.py ā Tool orchestration (search ā extract ā rank)
āāā search.py ā Pluggable search providers
āāā extractor.py ā HTML content extraction (BS4 + Playwright)
āāā browser.py ā Playwright browser manager
āāā models.py ā Pydantic data models
āāā config.py ā Settings (pydantic-settings + .env)
āāā logger.py ā Structured logging
āāā utils.py ā Shared helpers
āāā requirements.txt
āāā .env ā Configuration (fill in your API keys)
āāā README.md
```
---
## Quick Start
### 1. Prerequisites
- Python 3.11 or higher
- pip
### 2. Install Dependencies
```bash
pip install -r requirements.txt
```
### 3. Install Playwright Browser
```bash
playwright install chromium
```
> This downloads the Chromium binary (~130 MB). Required for JavaScript-heavy page extraction.
### 4. Configure API Keys
Edit `.env` and add at least one search provider key:
```env
SEARCH_PROVIDER=tavily
TAVILY_API_KEY=your_key_here
```
**Getting a free Tavily key** (recommended):
1. Visit [app.tavily.com](https://app.tavily.com)
2. Sign up for a free account
3. Copy your API key ā paste into `.env`
### 5. Run the Server
```bash
python server.py
```
The server starts in **STDIO mode** (default), ready to connect with any MCP client.
---
## MCP Tools
### `web_search`
Search the web and return ranked snippets (no page visits).
**Input:**
```json
{
"query": "Latest AI trends in healthcare",
"max_results": 10
}
```
**Output:**
```json
{
"query": "Latest AI trends in healthcare",
"total_results": 10,
"search_provider": "tavily",
"results": [
{
"title": "AI in Healthcare 2025",
"url": "https://example.com/ai-health",
"domain": "example.com",
"snippet": "Short summary of the article...",
"content": "Same as snippet for web_search",
"published_date": "2025-06-15",
"relevance_score": 0.92
}
],
"cached": false,
"execution_time_ms": 312.5
}
```
---
### `webpage_content`
Extract full readable content from a specific URL.
**Input:**
```json
{
"url": "https://example.com/article",
"use_browser": false
}
```
Set `use_browser: true` to force Playwright rendering for JavaScript-heavy pages.
---
### `search_and_extract`
End-to-end: search ā visit pages ā extract content ā rank results.
**Input:**
```json
{
"query": "Latest UK visa requirements 2025",
"max_results": 5,
"use_browser_fallback": true
}
```
Returns full page content for each result including title, author, publish date, and extracted text.
---
## Search Providers
| Provider | Env Key | Free Tier | Notes |
|---|---|---|---|
| **Tavily** ā | `TAVILY_API_KEY` | 1,000/month | Best snippets, recommended |
| **Brave** | `BRAVE_API_KEY` | 2,000/month | Privacy-focused |
| **Bing** | `BING_API_KEY` | 1,000/month | Azure Cognitive Services |
| **SerpAPI** | `SERPAPI_API_KEY` | 100/month | Proxies Google |
| **Google CSE** | `GOOGLE_CSE_API_KEY` + `GOOGLE_CSE_ID` | 100/day | Custom Search Engine |
Switch provider by changing `SEARCH_PROVIDER` in `.env`.
---
## Configuration Reference
| Setting | Default | Description |
|---|---|---|
| `SEARCH_PROVIDER` | `tavily` | Active search backend |
| `MAX_RESULTS` | `10` | Default result count |
| `REQUEST_TIMEOUT` | `30` | HTTP timeout (seconds) |
| `CONCURRENCY_LIMIT` | `5` | Parallel page extractions |
| `CACHE_TTL` | `300` | Cache time-to-live (seconds, 0 = disabled) |
| `CACHE_MAX_SIZE` | `256` | Max cache entries |
| `PLAYWRIGHT_HEADLESS` | `true` | Headless browser mode |
| `PLAYWRIGHT_TIMEOUT` | `30000` | Browser nav timeout (ms) |
| `RETRY_ATTEMPTS` | `3` | Max HTTP retry attempts |
| `LOG_LEVEL` | `INFO` | Logging verbosity |
| `LOG_FORMAT` | `json` | `json` or `text` |
---
## Connecting with MCP Clients
### Claude Desktop
Add to `claude_desktop_config.json`:
```json
{
"mcpServers": {
"web-search": {
"command": "python",
"args": ["C:/path/to/websearchMcp/server.py"],
"env": {
"TAVILY_API_KEY": "your_key_here"
}
}
}
}
```
### Custom MCP Client (Python)
```python
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
server_params = StdioServerParameters(
command="python",
args=["server.py"],
)
async with stdio_client(server_params) as (read, write):
async with ClientSession(read, write) as session:
await session.initialize()
result = await session.call_tool(
"search_and_extract",
{"query": "Latest AI trends", "max_results": 3}
)
print(result)
```
---
## Architecture
```
User Query
ā
ā¼
FastMCP Server (server.py)
ā validates input (Pydantic)
ā¼
Tool Orchestrator (tools.py)
ā checks cache ā calls provider
ā¼
Search Provider (search.py)
ā Tavily / Brave / Bing / SerpAPI / Google
ā¼
Raw Search Results
ā
ā¼
Content Extractor (extractor.py)
ā HTTP + BS4 ā Playwright fallback
ā¼
Cleaned & Ranked Results
ā
ā¼
Structured JSON Response
```
---
## Error Handling
All tools return structured error JSON on failure:
```json
{
"error": "No API key configured for provider 'tavily'",
"error_type": "RuntimeError",
"tool": "web_search",
"query": "AI trends",
"timestamp": "2025-06-30T18:00:00Z"
}
```
---
## Performance Tips
- Use `web_search` when you only need snippets (faster, uses less quota).
- Use `search_and_extract` for deep research requiring full article content.
- Increase `CONCURRENCY_LIMIT` for faster parallel extraction (be mindful of rate limits).
- Increase `CACHE_TTL` to reduce repeated API calls for the same queries.
- Set `PLAYWRIGHT_HEADLESS=true` (default) in production.
---
## Deploying to Render
1. Push the repo to GitHub (`.env` is git-ignored ā API keys are safe)
2. Go to **render.com ā New ā Blueprint** ā connect repo
3. Render detects `render.yaml` automatically
4. Set `TAVILY_API_KEY` in Render dashboard ā Environment Variables
5. Your SSE endpoint: `https://your-app.onrender.com/sse`
---
## Deploying to Azure Container Apps
**Prerequisites:**
- [Azure CLI](https://learn.microsoft.com/en-us/cli/azure/install-azure-cli): `winget install Microsoft.AzureCLI`
- [Docker Desktop](https://www.docker.com/products/docker-desktop/)
- Azure account with active subscription
**One-command deploy:**
```powershell
# 1. Login to Azure
az login
# 2. Run the deployment script (reads TAVILY_API_KEY from .env automatically)
.\deploy-azure.ps1
```
The script will:
- Create a Resource Group + Azure Container Registry
- Build and push the Docker image via **ACR Tasks** (builds in Azure cloud ā no local build needed)
- Create a Container Apps Environment
- Deploy the MCP server with your Tavily key stored as a **secret** (never in plain text)
- Print your live SSE endpoint URL
**Custom options:**
```powershell
.\deploy-azure.ps1 `
-ResourceGroup "my-rg" `
-Location "westeurope" `
-AppName "my-mcp-server" `
-Cpu "2.0" `
-Memory "4.0Gi"
```
**Add to your no-code platform after deploy:**
| Field | Value |
|---|---|
| **Transport** | `Server-Sent Events (SSE)` |
| **URL** | `https://<your-app>.<region>.azurecontainerapps.io/sse` |
**Health check:** `https://<your-app>.<region>.azurecontainerapps.io/health`
**Update after code changes:**
```powershell
# Just re-run the deploy script ā it rebuilds and redeploys
.\deploy-azure.ps1
```
---
## License
MIT
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues