camoufox-research
# Camoufox Research
**An MCP server that gives your AI agent a real browser** — search the web, read
JS/SPA pages, click, fill forms, crawl sites, extract tables, save files, watch
for changes. Runs on the anti-detect [Camoufox](https://github.com/daijro/camoufox)
(Firefox), so pages see a normal browser instead of a headless bot.
> По-русски: MCP-сервер для веб-ресёрча. Даёт агенту живой браузер — поиск,
> чтение тяжёлых страниц, клики, формы, сбор данных. Ставится одной командой.

[](https://www.python.org/downloads/)
[](https://github.com/aidvizhhub/camoufox-research/actions/workflows/ci.yml)
[](https://modelcontextprotocol.io)
[](https://github.com/daijro/camoufox)
[](https://github.com/aidvizhhub/camoufox-research/releases)
[](https://github.com/aidvizhhub/camoufox-research)
[](LICENSE)
[](https://aidvizhhub.github.io/camoufox-research/)
**46 tools** total, **18** in the default profile — a profile is just the set of
tool groups you switch on (`--caps`); fewer tools make the agent choose better.
```
AI agent ──MCP──▶ Camoufox Research ──▶ Camoufox (anti-detect Firefox) ──▶ web
```
## Why
Most MCP browser servers hand an agent isolated actions. This one is a complete
research toolkit in a single server — the agent searches, reads, interacts and
exports without you writing browser code.
- 🔎 **Search** — live web search that keeps working when a source is blocked,
plus deep `research` that gathers 10+ sources across many distinct sites in one call
- 🌐 **Browse** — JS/SPA text, live sessions with tabs, clicks, forms, uploads
- 👁️ **See pages** — Set-of-Mark screenshots and a compact snapshot tree with `ref`s
- 📊 **Extract** — CSS/XPath fields, tables → CSV, PDF/DOCX/XLSX, export JSON/MD
## 30-second demo
> *"Find all pricing pages on this site, extract the prices and save them to CSV."*
```
Agent
├─ map_site discover every /pricing page
├─ crawl read them (cached)
├─ extract {"plan": "css:.plan", "price": "css:.price"}
└─ export format=csv → prices.csv
```
No browser-automation code — just a sentence to your agent.
## Install (one command)
Requirements: **Python 3.10–3.13** and **git** (Windows: also PowerShell 7). The
installer creates a venv, installs the package **from the clone**, downloads the
browser once, registers the MCP server and checks the handshake. There is no
PyPI package — the code comes from this repo or Docker.
| OS | One command (from the clone root) |
|---|---|
| **Linux** (Fedora/Ubuntu/Debian/Arch/…) | `bash scripts/install.sh` |
| **macOS** | `bash scripts/install.sh` (same script; no system packages needed) |
| **Windows** (native, PowerShell 7) | `pwsh -NoProfile -File scripts\install.ps1` |
| **Docker** (any OS with Docker) | `bash scripts/install/run_in_docker.sh --build --run` |
### No clone, one command (`uv`) or the prebuilt image
`uv` installs the package straight from this repo — no clone, no PyPI:
```bash
uv tool install git+https://github.com/aidvizhhub/camoufox-research
camoufox-research # stdio MCP server; browser downloads on first run
```
One-shot without installing (`uvx`) — handy for an MCP client:
```bash
uvx --from git+https://github.com/aidvizhhub/camoufox-research camoufox-research
```
The image is prebuilt in GHCR (browser already inside, nothing to fetch):
```bash
docker run -i --rm -v camoufox-data:/data ghcr.io/aidvizhhub/camoufox-research:latest
```
`server.json` carries the MCP Registry metadata and points at that image. There
is **no PyPI package** by design — git, `uv`, or Docker.
Preview first — it changes nothing:
```bash
bash scripts/install.sh --dry-run # plan for THIS machine
pwsh -NoProfile -File scripts\install.ps1 -WhatIf # Windows plan
```
No clone yet? The same installer pulls the repo to `~/camoufox-research`:
```bash
curl -fsSL https://raw.githubusercontent.com/aidvizhhub/camoufox-research/main/scripts/install.sh | bash
```
Windows: `irm …/install.ps1 | iex`. Full matrix, flags and troubleshooting —
[docs/install-crossplatform.md](docs/install-crossplatform.md); native Windows
details — [docs/install-windows.md](docs/install-windows.md).
Check after install:
```bash
~/.venvs/camoufox-research/bin/python scripts/checks/mcp_probe.py # caps → handshake → tools count
opencode mcp list # opencode: camoufox connected
claude mcp list # Claude Code; Claude Desktop / Cursor — status in their UI
```
Update, or remove (cache and browser stay):
```bash
bash scripts/install/update_mcp.sh # pull → pip → reconnect
bash scripts/install.sh --reinstall # reinstall the package from the clone
bash scripts/install.sh --uninstall # venv + MCP entry; cache/browser kept
```
By hand, if you already have a clone and want full control:
```bash
python3 -m venv ~/.venvs/camoufox-research
~/.venvs/camoufox-research/bin/pip install .
~/.venvs/camoufox-research/bin/python -m camoufox fetch # browser, once
```
`pip install .` installs **the code on disk** — not from PyPI. Optional extras:
`pip install '.[geoip]'` adds IP-based geolocation, locale and timezone for
proxied runs (only active with a proxy and `CAMOUFOX_GEOIP=1`). The common env
vars (`CAMOUFOX_*`: cache dir, timeouts, caps, report dir) are documented in
`configs/example.env`; the advanced switches (stealth, SSRF,
budgets, LLM planner, transports) live in the code — list them all with
`grep -rhoE 'CAMOUFOX_[A-Z0-9_]+' camoufox_research/ | sort -u`.
## Connect to MCP
The installer writes the client config for you. To do it by hand, the canon is
**opencode v2**: the config is `~/.config/opencode/opencode.jsonc` and servers
live under `mcp.servers.<name>` — *not* `mcp.<name>`, and `disabled` — *not*
`enabled` (the old v1 shape is silently ignored by v2).
OpenCode (`~/.config/opencode/opencode.jsonc`): copy-ready section —
[mcp/config/opencode.jsonc.example](mcp/config/opencode.jsonc.example).
Claude Desktop and Cursor use their own shape (`mcpServers` + `command` string +
`env`); copy-ready files are in `mcp/config/`
(`claude_desktop_config.json.example`, `cursor_mcp.json.example`). Trying it from
sources still means installing once (the MCP SDK lives in the venv); then run the
console script from that venv:
`~/.venvs/camoufox-research/bin/camoufox-research`. Don't launch `mcp/server.py`
from the repo root — the local `mcp/` folder shadows the SDK and the import
fails.
Installed with `uv` (no clone)? Use as `command` (same opencode v2 shape):
`["uvx", "--from", "git+https://github.com/aidvizhhub/camoufox-research", "camoufox-research"]`.
Check — CLI `opencode mcp list` → `✓ camoufox connected`; inside a session the
`/mcps` command shows the same status. A one-page walkthrough (path, paste,
troubleshooting) is in
[mcp/config/opencode-connect.md](mcp/config/opencode-connect.md).
## Tools (46)
Tools are grouped, and the groups double as `--caps` profiles (see below).
| Group | Tools | Count |
|---|---|---|
| `research` | `web_search`, `research`, `paper_search` | 3 |
| `browser` | `fetch_page`, `batch_fetch`, `extract_links`, `extract`, `table_extract`, `crawl`, `map_site`, `sitemap`, `rss`, `read_document`, `check_links`, `export`, `page_diff` | 13 |
| `session` | `session_*` (tabs, clicks, forms, keys, JS, network, files), `set_proxy`, `profile_save/load` | 26 |
| `vision` | `snapshot` (refs), `screenshot` (`som=True`) | 2 |
| always on | `ping`, `stats` (audit, secrets masked) | 2 |
A few things worth knowing:
- Search runs a chain of independent engines: one that is blocked or returns
nothing cools down and the query moves on; results are cached per engine+query,
so a blocked source never takes `web_search` down. For engines that need a
proxy there is a pool (`CAMOUFOX_SEARCH_PROXY` / `CAMOUFOX_SEARCH_PROXIES`).
- `fetch_page` reuses a 24 h cache — a repeat costs no network. `delta=True` returns a
marker when nothing changed (saves answer tokens) but still does a fresh fetch —
use it for "what changed", not for cheap repeats.
- `snapshot` returns a ~2–5 KB YAML tree of interactive elements with a `ref`
on each — click by `ref`, no fragile selectors.
- `research` runs a deep hunt in one call: query expansion, quality ranking
(docs/GitHub/arXiv first), domain dedup, and honest `partial` when the goal
isn't reached. For a single fact, use `web_search`.
- `export` writes JSON / CSV / Markdown to disk.
## Tool profiles (fewer tools, better choice)
46 tools in one prompt degrade an agent's tool choice (industry rule of thumb:
past ~40 it drops). Pick the groups you need — the same idea as Playwright MCP's
`--caps`:
```bash
camoufox-research --caps research,browser # or env CAMOUFOX_CAPS
```
| Profile | Tools | What's inside |
|---|---|---|
| `research,browser` (default) | 18 | search + reading/extraction |
| `research,browser,session` | 44 | + live tab, forms, network, files |
| `research,browser,vision` | 20 | + `snapshot` (refs) and `screenshot` (PNG) |
| `all` | 46 | the full registry |
`ping`/`stats` are always available. `CAMOUFOX_TOOLS_ONLY` / `CAMOUFOX_TOOL_HIDE`
still apply on top. Invalid group → warning, valid groups still load.
## MCP resources & prompts
- **Resources:** `camoufox://stats`, `camoufox://cache`, `camoufox://session`,
`camoufox://info`, `camoufox://health`, `camoufox://search`
- **Prompts:** `research_plan`, `extract_schema`, `monitor_page`
## Transports
`stdio` (default), `streamable-http` (stateless; the remote option), `sse`
(legacy, kept for the 2026-07-28 spec's 12-month window):
```bash
camoufox-research --transport http --port 8833
CAMOUFOX_PORT=8833 camoufox-research --transport http
```
## Troubleshooting
`Unknown tool` usually means the client is still talking to an old server — a
restart after a code change fixes it. A one-command, read-only diagnosis:
```bash
~/.venvs/camoufox-research/bin/python scripts/checks/mcp_probe.py # human-readable
~/.venvs/camoufox-research/bin/python scripts/checks/mcp_probe.py --json # machine-readable
```
It shows python / repo → caps → protocol handshake → **how many tools
`tools/list` returns** → package version → search-watchdog pulse. A small
`tools=N` or a failed handshake means stale code in the venv → reinstall from
the clone (`bash scripts/install.sh --reinstall`) and reconnect (API
disconnect/connect, not kill).
## Honest numbers
There is no single head-to-head here. The three numbers below come from one
third-party benchmark run; ours come from separate runs on a smaller sample.
They sit next to each other for context — not as the same measurement.
**Reference (not our run):** [fastCRW `diagnose_3way.py`](https://fastcrw.com/blog/truth-recall-explained-web-scrapers)
scoring (`phrases > 20 chars`, `recall >= 0.3` = found) on the public dataset
`firecrawl/scrape-content-dataset-v1` — measured 2026-05-08 over the full 819 URLs:
| Tool | Truth-recall |
|---|---|
| fastCRW | 63.7% (522/819) |
| Crawl4AI | 60.0% (491/819) |
| Firecrawl | 56.0% (459/819) |
**Our own number is deliberately kept out of that table.** Our script runs a
30-URL prefix of the same dataset, so the comparison is indirect — different
sample size and a different date. At n=30 the 95% confidence interval is roughly
±18 p.p., wider than the gap between all the tools above. Two runs on 2026-09-29
gave **36.7%** (11/30) and **40.0%** (12/30), and a re-run within 24h measures the
page cache rather than the web. Read it as a smoke number for the extractor, not
as a leaderboard position. Reproduce with
`python scripts/dev/bench_truth_recall.py --sample N`.
## Documentation
| Doc | What's there |
|---|---|
| [docs/agent-usage.md](docs/agent-usage.md) | guide for the agent: profiles, call order, limits, pitfalls |
| [docs/RESEARCH-PLAYBOOK.md](docs/RESEARCH-PLAYBOOK.md) | recipes for fast / full / deep research, measured chars and seconds |
| [docs/VERIFICATION.md](docs/VERIFICATION.md) | what was verified live and by what, plus an honest "what was NOT verified" |
| [docs/install-crossplatform.md](docs/install-crossplatform.md) | install matrix: Linux / macOS / Windows / Docker |
| [docs/install-windows.md](docs/install-windows.md) | native Windows (PowerShell 7) install notes |
| [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) | layers, modules, state, extension points, debts |
| [docs/landmines.md](docs/landmines.md) | verified landmines ("what not to step on") — index in [EXPERIENCE.md](EXPERIENCE.md) |
| [SECURITY.md](SECURITY.md) | how to report a vulnerability (private advisory), scope, supported versions |
## Development
See [CONTRIBUTING.md](CONTRIBUTING.md): layout, how to add a tool, the live-smoke
ritual and the unit tests (`pip install -e ".[dev]"` → `pytest -q`). CI runs
`pytest`, `py_compile` + import + an MCP stdio smoke on Python 3.10–3.13, plus
the lint/filesize/version gates; full browser flows are checked locally.
## License
MIT — see [LICENSE](LICENSE).
Dependencies: [camoufox](https://github.com/daijro/camoufox) (see its repo),
[mcp](https://github.com/modelcontextprotocol/python-sdk) (MIT),
[trafilatura](https://trafilatura.readthedocs.io/) (GPL-3.0; **обязательна** —
read_document/table_extract и article_only извлекают ею, без неё чтение
деградирует до текста всего body).
TDQS
Scored across 18 tools
Most tools target clearly distinct stages of web research: search, fetch, crawl, extract, export, and monitoring. A few adjacent tools exist (fetch_page vs batch_fetch vs crawl; map_site vs sitemap vs extract_links; extract vs table_extract), but their descriptions explicitly clarify WHEN and NOT WHEN to use each.
All tool names use consistent snake_case English, with no camelCase or mixed conventions. The set is mostly verb_noun or noun_noun, though a few single nouns/verbs (rss, sitemap, stats, ping, crawl, export) are minor deviations from a strict verb_noun pattern.
18 tools is slightly above the ideal 3-15 range, but most earn their place by covering distinct research/scraping workflows. The set does not feel severely bloated, though a few specialized tools could arguably be consolidated.
Core research, fetching, crawling, extraction, document reading, and monitoring are well covered. However, several descriptions reference missing tools such as session_start, session_download, and snapshot, creating dead ends for interactive browsing, downloads, and selector discovery workflows.