Skip to main content
Glama
adi-santoso

free-web-search-mcp

by adi-santoso
README.md
<div align="center">

# free-web-search-mcp

A free web search MCP server powered by real browsers.

[Features](#features) · [Install](#install) · [Configure](#configure) · [Tools](#tools) · [Contributing](#contributing) · [License](#license)

</div>

---

## Features

- **MCP server** exposing `web_search`, `fetch_page`, and `setup_camoufox` tools
- **Multiple browser backends selectable at runtime**:
  - `camoufox` — stealth Firefox (anti-detect), default when installed
  - `chrome` — Playwright Chromium via installed Google Chrome
  - `msedge` — Playwright Chromium via installed Microsoft Edge
  - `chromium` — Playwright bundled Chromium
  - `firefox` — Playwright Firefox
  - `webkit` — Playwright WebKit
- **Auto-setup Camoufox** — pip install + binary download on demand
- Multiple search engines: Google (default), Bing, DuckDuckGo
- Page content extraction (plain text) for deep-reading a result URL
- Humanized typing, configurable proxy, locale, headless mode
- Block detection with actionable hints (e.g. "retry with `browser='camoufox'`")

## Install

```bash
# Default install (Playwright path)
pip install -e .

# With Camoufox (recommended for stealth)
pip install -e ".[camoufox]"
camoufox fetch

# Playwright browsers (only download the ones you'll use)
playwright install chromium        # for backend=chromium
playwright install firefox         # for backend=firefox
playwright install webkit          # for backend=webkit
# chrome / msedge backends use the already-installed system browser, no download needed.
```

> Camoufox missing on first run? It is installed and downloaded automatically — no manual steps required.

## Configure

```bash
cp .env.example .env
# edit .env
```

| Env              | Default     | Description                                                        |
| ---------------- | ----------- | ------------------------------------------------------------------ |
| `SEARCH_ENGINE`  | `google`    | Default search engine: `google` \| `bing` \| `duckduckgo`         |
| `MAX_RESULTS`    | `10`        | Max results returned per search                                    |
| `DEFAULT_BROWSER`| auto        | `camoufox` if installed, else `chrome`. Override per-call via param |
| `HEADLESS`       | `true`      | Run browser headless                                               |
| `HUMANIZE`       | `true`      | Humanized typing                                                   |
| `LOCALE`         | `en-US`     | Browser locale                                                     |
| `PROXY_URL`      | _empty_     | Optional proxy `http://user:pass@host:port`                        |
| `MCP_LOG_LEVEL`  | `INFO`      | `DEBUG` \| `INFO` \| `WARNING` \| `ERROR`                          |

## Run

```bash
# Direct
python -m free_web_search_mcp

# Via script entry point
free-web-search-mcp
```

## Integrate

### Claude Desktop — `claude_desktop_config.json`

```json
{
  "mcpServers": {
    "free-web-search": {
      "command": "python",
      "args": ["-m", "free_web_search_mcp"],
      "cwd": "<path-to-this-project>"
    }
  }
}
```

### opencode — `opencode.json` (global or project)

```json
{
  "mcp": {
    "free-web-search": {
      "type": "local",
      "command": ["python", "-m", "free_web_search_mcp"],
      "enabled": true,
      "environment": {
        "HEADLESS": "true",
        "HUMANIZE": "true",
        "DEFAULT_BROWSER": "camoufox"
      }
    }
  }
}
```

## Tools

### `web_search(query, max_results=10, engine="google", browser="")`

Search the web and return a list of `{position, title, url, snippet}`.

| Parameter     | Type  | Default   | Description                                                                  |
| ------------- | ----- | --------- | ---------------------------------------------------------------------------- |
| `query`       | str   | _required_| Search query string                                                          |
| `max_results` | int   | `10`      | 1-20                                                                         |
| `engine`      | str   | `google`  | `google` \| `bing` \| `duckduckgo`                                           |
| `browser`     | str   | auto      | `camoufox` \| `chrome` \| `msedge` \| `chromium` \| `firefox` \| `webkit`   |

### `fetch_page(url, max_chars=20000, browser="")`

Fetch a URL and extract its main content as plain text.

| Parameter  | Type  | Default    | Description                  |
| ---------- | ----- | ---------- | ---------------------------- |
| `url`      | str   | _required_ | Page URL                     |
| `max_chars`| int   | `20000`    | 1000-50000                   |
| `browser`  | str   | auto       | Same options as `web_search` |

### `setup_camoufox(force=False)`

Ensure the Camoufox stealth browser is installed and its binary is downloaded. Runs `pip install camoufox` and `python -m camoufox fetch` automatically if missing. No-op if already installed.

| Parameter | Type | Default | Description                              |
| --------- | ---- | ------- | ---------------------------------------- |
| `force`   | bool | `False` | Re-download the binary even if present   |

## Tips: avoid captcha & blocks

Search engines flag requests by **IP reputation**, not just browser fingerprint. Camoufox handles the fingerprint side, but the IP matters more than people expect.

**Run locally when possible.** A residential home IP is clean and almost never gets captchas. A VPS/datacenter IP is frequently flagged — even with Camoufox, Google may serve reCAPTCHA. Prefer running this server on your local machine (or home server) and pointing your MCP client at it.

**If you must run on a VPS:**
- Set `HUMANIZE=true` and keep requests spaced out (no rapid bursts).
- Use a **residential proxy** via `PROXY_URL` (`http://user:pass@host:port`). This is the single most effective fix for captcha on datacenter IPs.
- Fall back to `engine="bing"` or `engine="duckduckgo"` — they are generally more permissive than Google.
- Rotate locale / user-agent across calls.

**Block detection is built in.** When a search hits a captcha/consent page, the tool returns a clear error and suggests retrying with `browser="camoufox"` — so the caller can switch backend automatically.

## Contributing

Pull requests are welcome. Please open them against a feature branch (not `main`):

```bash
git checkout -b feature/your-feature
git push -u origin feature/your-feature
```

Then open a PR targeting `main`. Keep commits focused and titled in the imperative mood (e.g. "Add support for Yandex"). Run `ruff check .` before submitting.

## License

MIT — see [LICENSE](LICENSE).

TDQS

A4.1/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a distinct purpose: setup_camoufox handles installation, web_search performs queries, and fetch_page retrieves page content. No overlap in functionality.

Naming Consistency4/5

All names use snake_case, but web_search is a compound noun rather than a clear verb_noun construction like the other two. Minor deviation from the otherwise consistent pattern.

Tool Count5/5

Three tools is a well-scoped set for a web search server, covering setup, searching, and fetching without unnecessary bloat.

Completeness5/5

The domain of web search is fully covered: setup ensures the browser is ready, web_search finds results, and fetch_page reads full content. No obvious gaps for basic search workflows.

Maintenance

ActivitySlowing
ResponsivenessNo issues