Skip to main content
Glama
aidvizhhub

camoufox-research

by aidvizhhub
README.md
# Camoufox Research

**An MCP server that gives your AI agent a real browser** — search the web, read
JS/SPA pages, click, fill forms, crawl sites, extract tables, save files, watch
for changes. Runs on the anti-detect [Camoufox](https://github.com/daijro/camoufox)
(Firefox), so pages see a normal browser instead of a headless bot.

> По-русски: MCP-сервер для веб-ресёрча. Даёт агенту живой браузер — поиск,
> чтение тяжёлых страниц, клики, формы, сбор данных. Ставится одной командой.

![30-second demo](images/demo.gif)

[![Python](https://img.shields.io/badge/Python-3.10%20%7C%203.11%20%7C%203.12%20%7C%203.13-blue)](https://www.python.org/downloads/)
[![CI](https://img.shields.io/github/actions/workflow/status/aidvizhhub/camoufox-research/ci.yml?label=CI)](https://github.com/aidvizhhub/camoufox-research/actions/workflows/ci.yml)
[![MCP](https://img.shields.io/badge/MCP-2026--07--28%20%C2%B7%20SDK%202.2.0-blue)](https://modelcontextprotocol.io)
[![Camoufox](https://img.shields.io/badge/Camoufox-0.5.6-orange)](https://github.com/daijro/camoufox)
[![Version](https://img.shields.io/github/v/release/aidvizhhub/camoufox-research?sort=semver)](https://github.com/aidvizhhub/camoufox-research/releases)
[![Tools](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/aidvizhhub/camoufox-research/main/metrics/tools-badge.json)](https://github.com/aidvizhhub/camoufox-research)
[![License](https://img.shields.io/badge/license-MIT-blue)](LICENSE)
[![Showcase](https://img.shields.io/badge/showcase-site-blue?logo=github)](https://aidvizhhub.github.io/camoufox-research/)

**46 tools** total, **18** in the default profile — a profile is just the set of
tool groups you switch on (`--caps`); fewer tools make the agent choose better.

```
AI agent ──MCP──▶ Camoufox Research ──▶ Camoufox (anti-detect Firefox) ──▶ web
```

## Why

Most MCP browser servers hand an agent isolated actions. This one is a complete
research toolkit in a single server — the agent searches, reads, interacts and
exports without you writing browser code.

- 🔎 **Search** — live web search that keeps working when a source is blocked,
  plus deep `research` that gathers 10+ sources across many distinct sites in one call
- 🌐 **Browse** — JS/SPA text, live sessions with tabs, clicks, forms, uploads
- 👁️ **See pages** — Set-of-Mark screenshots and a compact snapshot tree with `ref`s
- 📊 **Extract** — CSS/XPath fields, tables → CSV, PDF/DOCX/XLSX, export JSON/MD

## 30-second demo

> *"Find all pricing pages on this site, extract the prices and save them to CSV."*

```
Agent
 ├─ map_site     discover every /pricing page
 ├─ crawl        read them (cached)
 ├─ extract      {"plan": "css:.plan", "price": "css:.price"}
 └─ export       format=csv  →  prices.csv
```

No browser-automation code — just a sentence to your agent.

## Install (one command)

Requirements: **Python 3.10–3.13** and **git** (Windows: also PowerShell 7). The
installer creates a venv, installs the package **from the clone**, downloads the
browser once, registers the MCP server and checks the handshake. There is no
PyPI package — the code comes from this repo or Docker.

| OS | One command (from the clone root) |
|---|---|
| **Linux** (Fedora/Ubuntu/Debian/Arch/…) | `bash scripts/install.sh` |
| **macOS** | `bash scripts/install.sh` (same script; no system packages needed) |
| **Windows** (native, PowerShell 7) | `pwsh -NoProfile -File scripts\install.ps1` |
| **Docker** (any OS with Docker) | `bash scripts/install/run_in_docker.sh --build --run` |

### No clone, one command (`uv`) or the prebuilt image

`uv` installs the package straight from this repo — no clone, no PyPI:

```bash
uv tool install git+https://github.com/aidvizhhub/camoufox-research
camoufox-research                 # stdio MCP server; browser downloads on first run
```

One-shot without installing (`uvx`) — handy for an MCP client:

```bash
uvx --from git+https://github.com/aidvizhhub/camoufox-research camoufox-research
```

The image is prebuilt in GHCR (browser already inside, nothing to fetch):

```bash
docker run -i --rm -v camoufox-data:/data ghcr.io/aidvizhhub/camoufox-research:latest
```

`server.json` carries the MCP Registry metadata and points at that image. There
is **no PyPI package** by design — git, `uv`, or Docker.

Preview first — it changes nothing:

```bash
bash scripts/install.sh --dry-run                    # plan for THIS machine
pwsh -NoProfile -File scripts\install.ps1 -WhatIf    # Windows plan
```

No clone yet? The same installer pulls the repo to `~/camoufox-research`:

```bash
curl -fsSL https://raw.githubusercontent.com/aidvizhhub/camoufox-research/main/scripts/install.sh | bash
```

Windows: `irm …/install.ps1 | iex`. Full matrix, flags and troubleshooting —
[docs/install-crossplatform.md](docs/install-crossplatform.md); native Windows
details — [docs/install-windows.md](docs/install-windows.md).

Check after install:

```bash
~/.venvs/camoufox-research/bin/python scripts/checks/mcp_probe.py   # caps → handshake → tools count
opencode mcp list               # opencode: camoufox connected
claude mcp list                 # Claude Code; Claude Desktop / Cursor — status in their UI
```

Update, or remove (cache and browser stay):

```bash
bash scripts/install/update_mcp.sh           # pull → pip → reconnect
bash scripts/install.sh --reinstall  # reinstall the package from the clone
bash scripts/install.sh --uninstall  # venv + MCP entry; cache/browser kept
```

By hand, if you already have a clone and want full control:

```bash
python3 -m venv ~/.venvs/camoufox-research
~/.venvs/camoufox-research/bin/pip install .
~/.venvs/camoufox-research/bin/python -m camoufox fetch   # browser, once
```

`pip install .` installs **the code on disk** — not from PyPI. Optional extras:
`pip install '.[geoip]'` adds IP-based geolocation, locale and timezone for
proxied runs (only active with a proxy and `CAMOUFOX_GEOIP=1`). The common env
vars (`CAMOUFOX_*`: cache dir, timeouts, caps, report dir) are documented in
`configs/example.env`; the advanced switches (stealth, SSRF,
budgets, LLM planner, transports) live in the code — list them all with
`grep -rhoE 'CAMOUFOX_[A-Z0-9_]+' camoufox_research/ | sort -u`.

## Connect to MCP

The installer writes the client config for you. To do it by hand, the canon is
**opencode v2**: the config is `~/.config/opencode/opencode.jsonc` and servers
live under `mcp.servers.<name>` — *not* `mcp.<name>`, and `disabled` — *not*
`enabled` (the old v1 shape is silently ignored by v2).

OpenCode (`~/.config/opencode/opencode.jsonc`): copy-ready section —
[mcp/config/opencode.jsonc.example](mcp/config/opencode.jsonc.example).

Claude Desktop and Cursor use their own shape (`mcpServers` + `command` string +
`env`); copy-ready files are in `mcp/config/`
(`claude_desktop_config.json.example`, `cursor_mcp.json.example`). Trying it from
sources still means installing once (the MCP SDK lives in the venv); then run the
console script from that venv:
`~/.venvs/camoufox-research/bin/camoufox-research`. Don't launch `mcp/server.py`
from the repo root — the local `mcp/` folder shadows the SDK and the import
fails.

Installed with `uv` (no clone)? Use as `command` (same opencode v2 shape):
`["uvx", "--from", "git+https://github.com/aidvizhhub/camoufox-research", "camoufox-research"]`.

Check — CLI `opencode mcp list` → `✓ camoufox connected`; inside a session the
`/mcps` command shows the same status. A one-page walkthrough (path, paste,
troubleshooting) is in
[mcp/config/opencode-connect.md](mcp/config/opencode-connect.md).

## Tools (46)

Tools are grouped, and the groups double as `--caps` profiles (see below).

| Group | Tools | Count |
|---|---|---|
| `research` | `web_search`, `research`, `paper_search` | 3 |
| `browser` | `fetch_page`, `batch_fetch`, `extract_links`, `extract`, `table_extract`, `crawl`, `map_site`, `sitemap`, `rss`, `read_document`, `check_links`, `export`, `page_diff` | 13 |
| `session` | `session_*` (tabs, clicks, forms, keys, JS, network, files), `set_proxy`, `profile_save/load` | 26 |
| `vision` | `snapshot` (refs), `screenshot` (`som=True`) | 2 |
| always on | `ping`, `stats` (audit, secrets masked) | 2 |

A few things worth knowing:

- Search runs a chain of independent engines: one that is blocked or returns
  nothing cools down and the query moves on; results are cached per engine+query,
  so a blocked source never takes `web_search` down. For engines that need a
  proxy there is a pool (`CAMOUFOX_SEARCH_PROXY` / `CAMOUFOX_SEARCH_PROXIES`).
- `fetch_page` reuses a 24 h cache — a repeat costs no network. `delta=True` returns a
  marker when nothing changed (saves answer tokens) but still does a fresh fetch —
  use it for "what changed", not for cheap repeats.
- `snapshot` returns a ~2–5 KB YAML tree of interactive elements with a `ref`
  on each — click by `ref`, no fragile selectors.
- `research` runs a deep hunt in one call: query expansion, quality ranking
  (docs/GitHub/arXiv first), domain dedup, and honest `partial` when the goal
  isn't reached. For a single fact, use `web_search`.
- `export` writes JSON / CSV / Markdown to disk.

## Tool profiles (fewer tools, better choice)

46 tools in one prompt degrade an agent's tool choice (industry rule of thumb:
past ~40 it drops). Pick the groups you need — the same idea as Playwright MCP's
`--caps`:

```bash
camoufox-research --caps research,browser     # or env CAMOUFOX_CAPS
```

| Profile | Tools | What's inside |
|---|---|---|
| `research,browser` (default) | 18 | search + reading/extraction |
| `research,browser,session` | 44 | + live tab, forms, network, files |
| `research,browser,vision` | 20 | + `snapshot` (refs) and `screenshot` (PNG) |
| `all` | 46 | the full registry |

`ping`/`stats` are always available. `CAMOUFOX_TOOLS_ONLY` / `CAMOUFOX_TOOL_HIDE`
still apply on top. Invalid group → warning, valid groups still load.

## MCP resources & prompts

- **Resources:** `camoufox://stats`, `camoufox://cache`, `camoufox://session`,
  `camoufox://info`, `camoufox://health`, `camoufox://search`
- **Prompts:** `research_plan`, `extract_schema`, `monitor_page`

## Transports

`stdio` (default), `streamable-http` (stateless; the remote option), `sse`
(legacy, kept for the 2026-07-28 spec's 12-month window):

```bash
camoufox-research --transport http --port 8833
CAMOUFOX_PORT=8833 camoufox-research --transport http
```

## Troubleshooting

`Unknown tool` usually means the client is still talking to an old server — a
restart after a code change fixes it. A one-command, read-only diagnosis:

```bash
~/.venvs/camoufox-research/bin/python scripts/checks/mcp_probe.py            # human-readable
~/.venvs/camoufox-research/bin/python scripts/checks/mcp_probe.py --json     # machine-readable
```

It shows python / repo → caps → protocol handshake → **how many tools
`tools/list` returns** → package version → search-watchdog pulse. A small
`tools=N` or a failed handshake means stale code in the venv → reinstall from
the clone (`bash scripts/install.sh --reinstall`) and reconnect (API
disconnect/connect, not kill).

## Honest numbers

There is no single head-to-head here. The three numbers below come from one
third-party benchmark run; ours come from separate runs on a smaller sample.
They sit next to each other for context — not as the same measurement.

**Reference (not our run):** [fastCRW `diagnose_3way.py`](https://fastcrw.com/blog/truth-recall-explained-web-scrapers)
scoring (`phrases > 20 chars`, `recall >= 0.3` = found) on the public dataset
`firecrawl/scrape-content-dataset-v1` — measured 2026-05-08 over the full 819 URLs:

| Tool | Truth-recall |
|---|---|
| fastCRW | 63.7% (522/819) |
| Crawl4AI | 60.0% (491/819) |
| Firecrawl | 56.0% (459/819) |

**Our own number is deliberately kept out of that table.** Our script runs a
30-URL prefix of the same dataset, so the comparison is indirect — different
sample size and a different date. At n=30 the 95% confidence interval is roughly
±18 p.p., wider than the gap between all the tools above. Two runs on 2026-09-29
gave **36.7%** (11/30) and **40.0%** (12/30), and a re-run within 24h measures the
page cache rather than the web. Read it as a smoke number for the extractor, not
as a leaderboard position. Reproduce with
`python scripts/dev/bench_truth_recall.py --sample N`.

## Documentation

| Doc | What's there |
|---|---|
| [docs/agent-usage.md](docs/agent-usage.md) | guide for the agent: profiles, call order, limits, pitfalls |
| [docs/RESEARCH-PLAYBOOK.md](docs/RESEARCH-PLAYBOOK.md) | recipes for fast / full / deep research, measured chars and seconds |
| [docs/VERIFICATION.md](docs/VERIFICATION.md) | what was verified live and by what, plus an honest "what was NOT verified" |
| [docs/install-crossplatform.md](docs/install-crossplatform.md) | install matrix: Linux / macOS / Windows / Docker |
| [docs/install-windows.md](docs/install-windows.md) | native Windows (PowerShell 7) install notes |
| [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) | layers, modules, state, extension points, debts |
| [docs/landmines.md](docs/landmines.md) | verified landmines ("what not to step on") — index in [EXPERIENCE.md](EXPERIENCE.md) |
| [SECURITY.md](SECURITY.md) | how to report a vulnerability (private advisory), scope, supported versions |

## Development

See [CONTRIBUTING.md](CONTRIBUTING.md): layout, how to add a tool, the live-smoke
ritual and the unit tests (`pip install -e ".[dev]"` → `pytest -q`). CI runs
`pytest`, `py_compile` + import + an MCP stdio smoke on Python 3.10–3.13, plus
the lint/filesize/version gates; full browser flows are checked locally.

## License

MIT — see [LICENSE](LICENSE).

Dependencies: [camoufox](https://github.com/daijro/camoufox) (see its repo),
[mcp](https://github.com/modelcontextprotocol/python-sdk) (MIT),
[trafilatura](https://trafilatura.readthedocs.io/) (GPL-3.0; **обязательна** —
read_document/table_extract и article_only извлекают ею, без неё чтение
деградирует до текста всего body).

TDQS

A3.8/5.0

Scored across 18 tools

Disambiguation4/5

Most tools target clearly distinct stages of web research: search, fetch, crawl, extract, export, and monitoring. A few adjacent tools exist (fetch_page vs batch_fetch vs crawl; map_site vs sitemap vs extract_links; extract vs table_extract), but their descriptions explicitly clarify WHEN and NOT WHEN to use each.

Naming Consistency4/5

All tool names use consistent snake_case English, with no camelCase or mixed conventions. The set is mostly verb_noun or noun_noun, though a few single nouns/verbs (rss, sitemap, stats, ping, crawl, export) are minor deviations from a strict verb_noun pattern.

Tool Count4/5

18 tools is slightly above the ideal 3-15 range, but most earn their place by covering distinct research/scraping workflows. The set does not feel severely bloated, though a few specialized tools could arguably be consolidated.

Completeness3/5

Core research, fetching, crawling, extraction, document reading, and monitoring are well covered. However, several descriptions reference missing tools such as session_start, session_download, and snapshot, creating dead ends for interactive browsing, downloads, and selector discovery workflows.

Maintenance

ActivityMaintained
ResponsivenessNo issues